<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Big Data Pipeline Discovery through Process Mining: Challenges and Research Directions⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simone Agostinelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dario Benvenuti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca De Luzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrea Marrella</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Sapienza Universiat ́ di Roma</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Big Data pipelines are essential for leveraging Dark Data, i.e., data collected but not used and turned into value. However, tapping their potential requires going beyond the current approaches and frameworks for managing their life-cycle. In this paper, we present the challenges associated to the achievement of the Pipeline Discovery task, which aims to learn the structure of a Big Data pipeline by extracting, processing and interpreting huge amounts of event data produced by several data sources. Then, we discuss how traditional Process Mining solutions can be potentially employed and customized to overcome such challenges, outlining a research agenda for future work in this area.</p>
      </abstract>
      <kwd-group>
        <kwd>Big Data Pipeline • Pipeline Discovery • Process Mining</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The main objective of the project is to develop a software ecosystem
consisting of new languages, methods and tools for supporting data pipelines on
heterogeneous resources. Six life-cycle phases will be covered: (1) pipeline
discovery, (2) pipeline definition, (3) pipeline simulation, (4) resource provisioning,
(5) pipeline deployment and (6) pipeline adaptation.</p>
      <p>In this paper, we focus on the phase of pipeline discovery, whose target is to
provide robust techniques to learn the structure of data pipelines by extracting,
processing and interpreting huge amounts of event data produced by several
data sources. To achieve this ambitious yet unexplored research goal, the idea
is to employ (and potentially customize) existing Process Mining solutions to
the discovery and analytics of data pipelines. In this paper, after presenting
in Section 2 the background on data pipelines and the challenges to properly
conduct the discovery task, in Section 3 we discuss how the application of existing
process mining solutions can be exploited to tackle and overcome the identified
challenges, towards the denfiition of novel approaches for pipeline discovery.
Finally, in Section 4, we conclude the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background and Challenges on Pipeline Discovery</title>
      <p>
        The literature on Big Data processing and analytics has often neglected the
research on pipeline discovery, working with the assumption that the anatomy
of data pipelines is already known at the outset, before running any Big Data
processing feature. A couple of relevant approaches exists that aims at studying
the structure of data pipelines. In [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], a framework that reveals key layers and
components to design data pipelines for manufacturing systems is presented. In
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], the authors derive a set of data and system requirements for implementing
equipment maintenance applications in industrial environments, and propose an
information system model that provides a scalable and resilient data pipeline for
integrating, processing and analysing industrial data.
      </p>
      <p>
        However, to date, there is no explicit research study that investigates the issue
of pipeline discovery. Consequently, even if the concept of “Big data pipeline” can
be traced back to 2012 (cf. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]), the literature lacks a shared understanding of
what a data pipeline is and how it can be defined. For instance, in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] the authors
refer to a data pipeline as the “path through which Big Data is transmitted,
stored, processed and analyzed”. In [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a data pipeline is defined as “a complex
chain of interconnected activities from data generation through data reception,
where the output of one activity becomes the input of the next one”. Similarly, in
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], a data pipeline is “a set of data processing elements connected in series, often
executed in time-sliced fashion, where the output of one element is the input of
the next one”. Then, in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], data pipelines are described as a “mechanism to
decompose complex analyses of large data sets into a series of simpler tasks, with
independently tuned components for each task”.
      </p>
      <p>The above definitions confirm that there is no uniefid specification of the
concept of data pipeline; nonetheless, some common features that are inherently
related to it can be identiefid:
– A data pipeline consists of chains of processing elements that manipulate
and interact with data sets;
– The outcome of a processing element of a data pipeline will be the input of
the next element in the pipeline;
– Each processing element of a data pipeline interacts with data sets considered
as “big”, i.e., with at least one of the Vs dimensions that is verified to hold.</p>
      <p>With this knowledge at hand, we performed many rounds of interviews with
the five business case partners involved in the DataCloud project (cf. also Section
4), which were useful not only to confirm the validity of the three above features
that characterize a data pipeline, but also to identify four major challenges to
be tackled towards the development of a robust pipeline discovery approach:
C1 Event Data Extraction: The challenge is to analyze and turn torrents of
rough data stored in several data sources or exchanged within the underlying
Cloud Computing infrastructures into valuable event data that reveal the
events that concretely happened into day-to-day operations.</p>
      <p>C2 Event Log Generation: Event data may contain interleaved information
related to the enactment of diferent data pipelines, or of multiple instances
of the same data pipeline. Moreover, the possibility exists that many events
must be filtered out by the analysis, since they do not refer to processing
elements that manipulate data (e.g., events that track the sending or
receiving of notifications). Therefore, the generation of event logs from the set
of event data is strongly required to (later) learn the structure of a data
pipeline. Each entry of the generated event logs should possess at least the
following characteristics: (i) a case identifier that maps each event to a case,
(ii) a timestamp that records when the event happened, (iii) the
processing element associated to the event, and (iv) the set of data processed or
manipulated during the event enactment.</p>
      <p>C3 Pipeline Structure Learning: This challenge is about the analysis and the
interpretation of event logs to learn the pipelines’ structure and to extract
valuable insights related to their performance and compliance.</p>
      <p>C4 Dark Data Analysis: This is the hardest part, because it requires to know
what to look for and where to look within the event data, without deploying
intrusive agents that manipulate the systems and networks of an
organization. Identifying Dark Data through the analysis of data pipelines would
enable to unlock their semantics and understand if some of them provide
insights and, finally, a certain business value.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Pipeline Discovery through Process Mining</title>
      <p>
        Even if the specification of a shared denfiition of a data pipeline is still a research
challenge, from the previous section it is evident that many similarities exist
between the concepts of “data pipeline” and “business process”. With the main
diference that any element of a data pipeline is thought to manipulate some (big)
data set. Conversely, business processes include activities that do not necessarily
interact with any kind of data. In fact, in the Business Process Management
(BPM) field, data flow is usually not considered as a first-class citizen [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Nonetheless, the discovery of data pipelines resembles the discovery of
business processes [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], as both require an event log as a starting point to enact the
discovery task. For this reason, in the range of the DataCloud project, we
investigate how the (customized) use of Process Mining solutions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] may support
the development of novel techniques to achieve the pipeline discovery task.
      </p>
      <p>Process mining is a family of data analysis techniques that enable decision
makers to discover process models from data (process discovery ), compare
expected and actual behaviours (conformance checking ), and enrich models with
information retrieved from data (process enhancement ). Process mining focuses
on the real execution of processes, as reflected by the footprint of reality logged
(in the form of explicit event logs) by the software systems of an organization.</p>
      <p>
        Within DataCloud, we are investigating and elaborating the following
research solutions that are inspired from the process mining literature in order to
tackle the challenges presented in Section 2:
S1 Human-in-the-Loop Methodology for Extracting Event Data. The
literature on process mining provides a number of semi-automated methods
to support organizations in extracting event data from data sources, such
as PM2 and L∗ [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. However, often their application is hampered by the
considerable preparation efort that needs to be conducted by human experts
at diferent stages of the extraction procedure [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This issue is even more
severe in presence of heterogeneous data sources that store huge amount of
data, like in the case of data pipelines. Furthermore – to date – there is no
deep understanding of how human experts should be involved in the
process of event data extraction. Within DataCloud, we aim to tackle C1 by
enhancing the existing event data extraction methods through the
identiifcation and specification of the manual activities that the human experts
need to perform in the context of event data extraction. This will include,
for example, activities like the assessment of the quality of the available data
sources, the detection of those data elements that relate to events, etc.
S2 Pre-processing, Clustering and Filtering techniques. To tackle C2,
i.e., to reduce the overall dataset complexity by extrapolating only its
relevant fragments for an efective event log generation, we will work on the
realization of four techniques: (i) Segmentation pre-processes the event data to
identify the events that belong to the same pipeline (i.e., case); (ii)
Aggregating events reduces complexity and improves the structure of discovery results
by merging multiple events into larger ones; (iii) Clustering partitions the
event log to discover simpler models for each partition of a complex pipeline;
(iv) Filtering removes potential outliers from the log. Concerning (i), (ii) and
(iii), the literature on process mining provides many solutions that can be
potentially customized and re-used in the context of data pipelines [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. On
the other hand, segmentation is a rather unexplored topic in process mining,
since it is assumed that any event in a log is always associated to a known
case. In practice, the majority of information systems do not record case
identifiers explicitly. To mitigate this issue, we will leverage our previous
works on segmentation performed in the Robotic Process Automation field
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to semi-automatically detect the diferent pipeline cases from a log.
S3 Pipeline Discovery algorithm. To learn the structure of a data pipeline,
we aim to leverage existing process discovery algorithms from process
mining [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], enhancing them with other techniques coming from diferent areas,
ranging from data mining to automated planning in Artificial Intelligence,
like already experienced in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Our solution is to realize a pipeline
discovery algorithm that enables not only to eficiently build the sequence flow
of the discovered pipelines, but also learning all data flows and event-based
conditions that ruled their execution.
      </p>
      <p>
        S4 Conformance Checking technique for Dark Data analysis. To tackle
C4, we aim to customize existing conformance checking techniques [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to
replay the streams of event data filtered out during the event log generation
phase (and stored into a dedicated Dark database) over the structure of
the discovered data pipelines. The target is to understand if some of the
discarded data can be potentially exploited to improve the quality or the
business value of the identified data pipelines. Of course, the definition of
specific threshold values to quantify if a dark data should be restored in an
event log must be investigated and specified as well.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Concluding Remarks</title>
      <p>The expected impact of the DataCloud project is to lower the technological entry
barriers for the incorporation of Big Data pipelines in organizations’ workflows
and make them accessible to a wider set of stakeholders regardless of the
hardware infrastructure. In this context, we discussed the key considerations around
the concept of pipeline discovery and suggested a number of research challenges
and potential ways to tackle them employing process mining solutions, to serve
as a research agenda for the future.</p>
      <p>All the proposed solutions for pipeline discovery will be validated through
a strong selection of complementary business cases ofered by four SMEs and
a large company targeting higher mobile business revenues in smart marketing
campaigns, reduced live streaming production costs of sport events, trustworthy
eHealth patient data management, and reduced time to production and better
analytics in Industry 4.0 manufacturing.</p>
      <p>Acknowledgments. This work has been supported by the Horizon 2020 project
DataCloud (Grant number 101016835).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>van der Aalst</surname>
          </string-name>
          , W.: Data Science in Action, pp.
          <fpage>3</fpage>
          -
          <lpage>23</lpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Agostinelli</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marrella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mecella</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Automated Segmentation of User Interface Logs</article-title>
          . In: Robotic Process Automation, pp.
          <fpage>201</fpage>
          -
          <lpage>222</lpage>
          . De Gruyter (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Augusto</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Conforti</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dumas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosa</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maggi</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marrella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mecella</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>Automated Discovery of Process Models from Event Logs: Review and Benchmark. IEEE Trans. on Know. and Data Eng</source>
          .
          <volume>31</volume>
          (
          <issue>4</issue>
          ) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Barika</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garg</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zomaya</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moorsel</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ranjan</surname>
          </string-name>
          , R.:
          <article-title>Orchestrating big data analysis workflows in the cloud: Research challenges, survey, and future directions</article-title>
          .
          <source>ACM Comput. Surv</source>
          .
          <volume>52</volume>
          (
          <issue>5</issue>
          ) (
          <year>Sep 2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Carmona</surname>
            , J., van Dongen,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weidlich</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Conformance checking. Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chakrabarty</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joshi</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          :
          <article-title>Dark Data: People to People Recovery</article-title>
          .
          <source>In: ICT Analysis and Applications</source>
          , pp.
          <fpage>247</fpage>
          -
          <lpage>254</lpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Diba</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batoulis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weidlich</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weske</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Extraction, correlation, and abstraction of event data for process mining</article-title>
          .
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source>
          <volume>10</volume>
          (
          <issue>3</issue>
          ),
          <year>e1346</year>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Dumas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>La</given-names>
            <surname>Rosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Mendling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Reijers</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.A.</surname>
          </string-name>
          , et al.:
          <source>Fundamentals of Business Process Management</source>
          , vol.
          <volume>1</volume>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. van Eck,
          <string-name>
            <given-names>M.L.</given-names>
            ,
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ,
            <surname>Leemans</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.J.J</surname>
          </string-name>
          ., van der Aalst,
          <string-name>
            <surname>W.M.P.:</surname>
          </string-name>
          <article-title>Pm2: A process mining project methodology</article-title>
          . In: Zdravkovic,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kirikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Johannesson</surname>
          </string-name>
          , P. (eds.)
          <source>Advanced Information Systems Engineering</source>
          . pp.
          <fpage>297</fpage>
          -
          <lpage>313</lpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gimpel</surname>
          </string-name>
          , G.:
          <article-title>Bringing dark data into the light: Illuminating existing IoT data lost within your organization</article-title>
          .
          <source>Business Horizons</source>
          <volume>63</volume>
          (
          <issue>4</issue>
          ),
          <fpage>519</fpage>
          -
          <lpage>530</lpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Gressling</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Data Science in Chemistry: Artificial Intelligence, Big Data, Chemometrics and Quantum Computing with Jupyter</article-title>
          . De Gruyter (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mannhardt</surname>
          </string-name>
          , F.,
          <string-name>
            <surname>de Leoni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reijers</surname>
          </string-name>
          , H.A.,
          <string-name>
            <surname>van der Aalst</surname>
            ,
            <given-names>W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toussaint</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Guided process discovery-a pattern-based approach</article-title>
          .
          <source>Information Systems</source>
          <volume>76</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Marrella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Lesper´ance, Y.:
          <article-title>Synthesizing a library of process templates through partial-order planning algorithms</article-title>
          .
          <source>In: Enterprise, Business-Process and Information Systems Modeling</source>
          , pp.
          <fpage>277</fpage>
          -
          <lpage>291</lpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Munappy</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olsson</surname>
            ,
            <given-names>H.H.</given-names>
          </string-name>
          :
          <article-title>Data pipeline management in practice: Challenges and opportunities</article-title>
          .
          <source>In: International Conference on Product-Focused Software Process Improvement</source>
          . pp.
          <fpage>168</fpage>
          -
          <lpage>184</lpage>
          . Springer (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Oleghe</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salonitis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>A framework for designing data pipelines for manufacturing systems</article-title>
          .
          <source>Procedia CIRP 93</source>
          ,
          <fpage>724</fpage>
          -
          <lpage>729</lpage>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>O</given-names>
            <surname>'Donovan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Leahy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Bruton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>O'Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.T.</surname>
          </string-name>
          :
          <article-title>An industrial big data pipeline for data-driven analytics maintenance applications in large-scale smart manufacturing facilities</article-title>
          .
          <source>Journal of Big Data</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>26</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Plale</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kouper</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>The centrality of data: data lifecycle and data pipelines. In: Data analytics for intelligent transportation systems</article-title>
          , pp.
          <fpage>91</fpage>
          -
          <lpage>111</lpage>
          . Elsevier (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Rabl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jacobsen</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          :
          <article-title>Big data generation</article-title>
          . In: Rabl,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Poess</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Baru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Jacobsen</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.A</surname>
          </string-name>
          . (eds.)
          <article-title>Specifying Big Data Benchmarks</article-title>
          . pp.
          <fpage>20</fpage>
          -
          <lpage>27</lpage>
          . Springer (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Raman</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Swaminathan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gehrke</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joachims</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Beyond myopic inference in big data pipelines</article-title>
          .
          <source>In: Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          . pp.
          <fpage>86</fpage>
          -
          <lpage>94</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Van Eck</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leemans</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Der Aalst</surname>
            ,
            <given-names>W.M.:</given-names>
          </string-name>
          <article-title>PM2: a process mining project methodology</article-title>
          .
          <source>In: International Conference on Advanced Information Systems Engineering</source>
          . pp.
          <fpage>297</fpage>
          -
          <lpage>313</lpage>
          . Springer (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>