<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DataOps - Towards a Definition</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julian Ereth</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Stuttgart</institution>
          ,
          <addr-line>Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Organizations seek to streamline their data and analytics structures in order to meet increasingly demanding business requirements. This can be difficult due to complex and fast-moving data landscapes. DataOps promises a remedy by combining an integrated and process-oriented perspective on data with automation and methods from agile software engineering, like DevOps, to improve quality, speed, and collaboration and promote a culture of continuous improvement. The goal of this on-going research is to elaborate DataOps as a new discipline. For this, it explores the body of knowledge and presents a working definition of DataOps as well as an initial research framework based on an explorative literature review and eight interviews with industry experts.</p>
      </abstract>
      <kwd-group>
        <kwd>DataOps</kwd>
        <kwd>Analytics</kwd>
        <kwd>Agile</kwd>
        <kwd>DevOps</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Data is a key asset to compete in today’s business. Data-driven decision making
significantly increases business success [1] and data is essential for many business processes
or even entire business models [1, 2]. Consequently, companies seek to streamline their
data and analytics processes in order to make them more efficient, provide data faster
and in superior quality, and ensure a stable and reliable operation in general. This,
however, can be hard due to fractured data landscapes with heterogeneous tools and
technologies, a broad scope with various stakeholders, rapidly changing business
requirements, and a general lack of standards [3, 4]. To revise these insufficient and inefficient
structures, there is a need to shape enterprise data and analytics in a way that enables a
stable operation and increases speed, quality, and overall productivity. Discussions on
how to achieve this, often includes topics like agile methods, data governance concepts,
or the use of automation. Here, many see similarities to challenges in software
engineering, where DevOps and continuous integration were introduced to provide
highquality software at an every-increasing pace [5]. However, data analytics is different to
software engineering, and consequently the new term DataOps emerged [6].
The goal of this ongoing research is to academically elaborate DataOps as a new
discipline. This paper explores the body of knowledge and presents a working definition of
DataOps and an initial research framework. We firstly depict related topics and our
methodical approach. Then, we discuss first results from a qualitative exploration and
derive a working definition and an initial research framework. The paper finally
concludes with an outlook for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>The field of DataOps goes hand in hand with a continuous professionalization of data
and analytics processes in companies. It combines ideas from information systems
research with different areas like agile and lean thinking and modern software
engineering [7]. For one thing, DataOps is taking its cue from DevOps which is “an
organizational approach that stresses empathy and cross-functional collaboration within and
between teams” [5] in order to accelerate delivery of changes and increase quality of
software [5, 8]. For this, DevOps promotes integrated and highly-automated engineering
pipelines to bridge the gap between development and IT operations. This can be
challenging, since development and operations are historically separate tasks with
conflicting goals. Software engineers need to quickly respond to changing requirements and
seek to rapidly deploy new features. The operations team, in contrast, is interested in
providing stable and reliable services and infrastructure and therefore avoids risks and
works as predictably as possible [8]. This is why DevOps requires a mind-shift and
strives for change in company culture [5, 8]. As DevOps is a rather novel concept,
related research is mainly focusing on exploring ways of adaption [9, 10] and its
business value [10, 11]. Next to DevOps, other organizational and technological approaches
from software engineering like, behavior-driven development [12] or scrum [13], as
well as the general philosophy of the agile manifesto [14] appear in DataOps
discussions.</p>
      <p>From an information systems research perspective, DataOps fits into the continuous
stream of work about agility in business intelligence (BI) and data science. Here, agility
is seen as “the ability to react to unforeseen or volatile requirements regarding the
functionality or the content of a BI solution” [15]. This encompasses the transformation of
processes with agile methods like Scrum or Kanban, as well as the adaption of new
technologies and architectural concepts to provide flexibility and increase value [3, 16,
17]. There is also an overlap with research about BI and analytics maturity models that
tries to assess the level of development of organizational capabilities and resources in
organizations [18]. Here, the associated goals are usually to reveal necessary steps to
advance BI and analytics in order to increase productivity and business value [19].
Findings in this area shows that mature analytics solutions involve a high-level of
collaboration, strategic alignment, and an enterprise-wide integration [20, 21], what
matches the goals of DataOps.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <p>In order to examine the field of DataOps, we followed an exploratory research design
outlined in Fig. 1. First, we conducted an exploratory literature review for an initial
exploration of the body of knowledge and to provide a better general understanding of
terms and concepts [22].</p>
      <p>Exploratory literature
review</p>
      <p>
        Interviews with
industry experts
For this, we followed Levy and Ellis [23] and queried four scientific databases, namely
SpringerLink, ACM, AISeL and IEEE, as they cover numerous important journals and
conferences in the information systems domain. In the search we used the keyword
DataOps as well as a combination of the terms DevOps and Data. We manually
screened the results and filtered out articles that seemed to have no relevancy, e.g. focus
on another domain or too specific topics. This process resulted in 783 documents of
which 243 were selected relevant and were further investigated by screening the
abstract and full text. It turned out that only 6 documents contained DataOps in its title or
abstract. The temporal distribution of the publications confirmed the novelty and
relevancy of the topic, as t
        <xref ref-type="bibr" rid="ref8">he first document is dated at 2012</xref>
        and since then the number of
publications increased continuously (c.f. Fig. 2). In addition to the academic literature,
we reviewed practitioner literatures, like blog posts or white papers of industry analysts,
in order to investigate the understanding of related terms and concepts in practice and
increase relevance for practitioners [24].
      </p>
      <p>300
200
100
0
18
2012
14
2013
60
2014
210</p>
      <p>241
76</p>
      <p>157
2015
2016
2017
2018
For a qualitative exploration, we then interviewed selected industry experts to gain
deeper insights about the scope of DataOps. To select relevant experts, we searched for
companies that offer DataOps services (vendors) or use DataOps in practice (users).
We conducted eight interviews with international experts from different companies
(c.f. Table 1). Most companies were vendors of which five can be denoted as startups
(&lt; 10 years or &lt; 50 employees) and three can be assigned to the enterprise level. The
interviews were conducted as semi-structured discussions of 30 – 60 minutes by phone.
The experts were asked about their understanding of DataOps, how their services fit
into this area and about goals and principles related to DataOps. The interviews were
transcribed and coded to identify more general components of DataOps. For this, we
first labeled statements (e.g. “we allow unit-tests in databases” as technical tests) and
then grouped similar labels (e.g. technical tests and acceptance tests as testing).</p>
      <p>Company
Company C1
Company C2
Company C3
Company C4
Company C5
Company C6
Company C7
Company C8
The results of the literature review and the insights of the interviews were then
triangulated to (i) formulate an initial working definition for the term DataOps and (ii) to derive
a research framework for further work. Moreover, the derived working definition was
discussed with two experts of the preceding interviews to test its relevance for practice.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>The exploration of the domain shows, that there is no general accepted definition of
DataOps yet. The term itself was introduced by Andy Palmer who described DataOps
as a discipline that “addresses the needs of data professionals on the modern internet
and inside the modern enterprise” [6]. Since then, there are different understandings of
the term within various scopes, like DataOps is “a hub for collecting and distributing
data” [25], “the function within an organization that controls the data journey from
source to value” [26], or “a new way of managing data that promotes communication
between, and integration of, formerly siloed data, teams, and systems” [7]. Similarly,
the conducted interviews also included various understandings like “all activities
between the data and operation teams” or “an integrated perspective over the entire data
lifecycle”.</p>
      <p>When describing these vague understandings, most authors emphasize continuous
improvement and a culture of collaboration and trust as the underlying goal of DataOps
[7, 27, 28] and elaborate their definitions with a set of goals, principles, and
components. Moreover, many mentioned the empowerment of citizen-users and an end-to-end
thinking as core objectives next to enhanced speed and quality [6, 28, 29]. When it
comes to implementation, some use the term data pipeline to describe a
process-oriented structure where data is transferred through multiple stages (e.g. extracted,
transformed, and visualized) [26, 28]. This concept is often used to support orchestration
and automation in complex scenarios. In this context, some advocate the idea that every
artifact (e.g. data models or visualizations) can be represented as code (analytics as
code) [27]. From an agile perspective, DataOps adopts the strive for short-cycles and
incremental change. One interview summarized the DataOps with the goal to “become
as fast as a startup, while being as robust as a manufactory”. In addition to that, the
examined literature and the interviews revealed data-driven improvements, reuse of
artifacts, and testing and monitoring as key principles of DataOps [7, 27, 30].
The results indicate, that DataOps is rather a collection of various practices and
technologies, than a particular method or tool. Based on our initial findings, we suggest the
following working definition of the term DataOps:
DataOps is a set of practices, processes and technologies that combines an integrated
and process-oriented perspective on data with automation and methods from agile
software engineering to improve quality, speed, and collaboration and promote a culture
of continuous improvement.</p>
      <p>This definition is intended to be a starting point for a common understanding of the
term DataOps and an early limitation of its scope. It does not claim to be final and
should be subject of discussion in future work.</p>
      <p>Next to the working definition, we propose the research framework in Fig. 3 to further
explore the DataOps space. This framework differentiates between the exploration of
DataOps as a discipline, which includes methods, technologies and concrete
implementations, and the investigation of the business value of DataOps.</p>
      <p>Environment
Technology</p>
      <p>Organization
DataOps as a Discipline</p>
      <p>Dynamic</p>
      <p>Capabilities</p>
      <p>Business Value
DataOps as a discipline can be further broken down by the means of the
TechnologyOrganization-Environment (TOE) Framework [31] that helps to explain the relation
between technology, organization and external factors in the diffusion of innovations.
This framework brings the necessary flexibility to cover the wide scope of DataOps
that spans from technological advancements, like automated data testing and
continuous deployments, to organizational initiatives, like end-to-end collaboration throughout
various business functions. Moreover, the TOE-framework has already been
successfully used in similar areas, e.g. cloud computing [32] or big data [33]. Regarding the
examination of the business value of DataOps, the concept of dynamic capabilities [34,
35] seems to be a valid candidate for a theoretical foundation, as it focuses “the ability
to integrate, build, and reconfigure internal and external competencies to address
rapidly-changing environments” [34] and it was already used in prior work to explain the
business value of data and analytics [1, 36, 37] or DevOps [38]. This research
framework is intended as an initial structure for future research. The prosed methods and
theoretical foundations need to be validated and refined in upcoming work.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Outlook</title>
      <p>This research illustrated the broad character of DataOps and showed that it is not a
particular method or tool, but rather a collection of principles and a way of doing things
on a cultural, organizational, and technological level. With its exploratory character,
the research contributes to the body of knowledge by academically defining DataOps
and arranging it in the field of information system research. Moreover, the research can
help organizations to understand and adopt DataOps in practice.</p>
      <p>DataOps is a rather novel term and there is only little experience and many research
challenges. The proposed framework provides a starting point for future work that
should, first, elaborate DataOps as a discipline, e.g. by developing blueprints for roles,
processes and governance, and further explore related technologies, and, second,
investigate the business value proposition of DataOps. As the diffusion of DataOps is still
low, we suggest to conduct in-depth case studies to compare traditional approaches with
DataOps-like implementations in order to gain further insights and refine this work.
1.
2.
3.
4.
5.
6.
7.
8.
9.
10.
11.
12.
13.
14.
15.
16.
17.
18.
19.
20.
21.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Seddon</surname>
            ,
            <given-names>P.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Constantinidis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dod</surname>
          </string-name>
          , H.:
          <article-title>How does business analytics contribute to business value?</article-title>
          .
          <source>Information Systems Journal</source>
          <volume>4</volume>
          (
          <issue>3</issue>
          ),
          <fpage>3380</fpage>
          -
          <lpage>3396</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>International Journal of Production Economics</source>
          <volume>165</volume>
          ,
          <fpage>234</fpage>
          -
          <lpage>246</lpage>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Baars</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ereth</surname>
          </string-name>
          , J.:
          <article-title>From Data Warehouses to Analytical Atoms - The Internet of Things as a Centrifugal Force in Business Intelligence and Analytics</article-title>
          .
          <source>In 24th European Conference on Information Systems (ECIS) Istanbul</source>
          , Turkey (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>LaValle</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lesser</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shockley</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hopkins</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kruschwitz</surname>
          </string-name>
          , N.:
          <article-title>Big data, analytics and the path from insights to value</article-title>
          .
          <source>MIT sloan management review 52(2)</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>32</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Dyck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Penners</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lichter</surname>
          </string-name>
          , H..
          <article-title>Towards Definitions for Release Engineering and DevOps</article-title>
          . In 2015 IEEE/ACM 3rd International Workshop on Release Engineering, Florence, p.
          <fpage>3</fpage>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          : From DevOps to DataOps, https://www.tamr.
          <article-title>com/from-devops-todataops-by-andy-palmer/</article-title>
          ,
          <source>last accessed</source>
          <year>2018</year>
          /04/21 (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Thusoo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sen</surname>
            <given-names>Sarma</given-names>
          </string-name>
          , J.:
          <article-title>Creating a Data-Driven Enterprise with DataOps - Insights from Facebook, Uber</article-title>
          , LinkedIn, Twitter, and
          <string-name>
            <surname>eBay. O'Reilly Media Inc</surname>
          </string-name>
          ., Sebastopol, CA (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Hüttermann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>DevOps for developers</article-title>
          .
          <source>Apress</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Wiedemann</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>IT Governance Mechanisms for DevOps Teams - How Incumbent Companies Achieve Competitive Advantages</article-title>
          .
          <source>In Proceedings of the 51st Hawaii International Conference on System Sciences</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>König</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steffens</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Towards a Quality Model for DevOps</article-title>
          .
          <source>In Continuous Software Engineering &amp; Full-scale Software Engineering</source>
          , p.
          <volume>37</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Mishra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garbajosa</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bosch</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abrahamsson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Future directions in Agile research: Alignment and divergence between research and practice</article-title>
          .
          <source>Journal of Software: Evolution and Process</source>
          <volume>29</volume>
          (
          <issue>6</issue>
          ), p.
          <source>e1884</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Wynne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hellesoy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tooke</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The cucumber book: behaviour-driven development for testers and developers</article-title>
          .
          <source>Pragmatic Bookshelf</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Schwaber</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beedle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Agile software development with Scrum</article-title>
          . Prentice Hall, Upper Saddle River (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , et al.:
          <article-title>Manifesto for agile software development</article-title>
          . http://agilemanifesto.org/,
          <source>last accessed</source>
          <year>2018</year>
          /07/29 (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Zimmer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baars</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kemper</surname>
          </string-name>
          , H.-G.:
          <article-title>The impact of agility requirements on business intelligence architectures</article-title>
          .
          <source>In: Proceedings of the 45th Hawaii International Conference on System Science (HICSS)</source>
          , pp.
          <fpage>4189</fpage>
          -
          <lpage>4198</lpage>
          . IEEE (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>Information Systems Management</source>
          <volume>32</volume>
          (
          <issue>3</issue>
          ),
          <fpage>177</fpage>
          -
          <lpage>191</lpage>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>A review and future direction of agile, business intelligence, analytics and data science</article-title>
          .
          <source>International Journal of Information Management</source>
          <volume>36</volume>
          (
          <issue>5</issue>
          ),
          <fpage>700</fpage>
          -
          <lpage>710</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Knabke</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Olbrich</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Capabilities To Achieve Business Intelligence Agility - Research Model And Tentative
          <string-name>
            <surname>Results</surname>
          </string-name>
          (
          <article-title>Research-in-progress)</article-title>
          .
          <source>In: Proceedings of the 20th Pacific Asia Conference on Information Systems (PACIS)</source>
          , p.
          <volume>35</volume>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Forsgren</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Humble</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The role of continuous delivery in IT and organizational performance</article-title>
          .
          <source>In: Proceedings of the Western Decision Sciences Institute (WDSI)</source>
          ,
          <source>Las Vegas</source>
          ,
          <string-name>
            <surname>NV</surname>
          </string-name>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>