<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Knowledge-Based Approach for Structuring Cyclic Workflows</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rafael Brandão</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vitor Lourenço</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcelo Machado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leonardo Azevedo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcelo Cardoso</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renan Souza</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guilherme Lima</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renato Cerqueira</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcio Moreno</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM Research</institution>
          ,
          <addr-line>Rio de Janeiro, RJ</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper showcases the Cycle Orchestrator, a microservices infrastructure designed to structure and manage workflows related to heterogeneous data, through a knowledge-based perspective. It aims at leveraging reasoning, explainability and collaboration among users over experiments that comprise workflow executions. We briefly discuss about design and implementation aspects to support the lifecycle of workflows (i.e., modeling, configuration, execution, provenance tracking and querying), exploring a holistic representation called Hyperknowledge that is amenable to be consumed and reasoned upon.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge-based Workflow Orchestration</kwd>
        <kwd>Workflow Management Systems</kwd>
        <kwd>Hyperknowledge Representation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Processing massive amounts of data through different techniques while enabling
stakeholders to collaborate and consume experiments’ results is a crosscutting challenge
tackled by many industries and technical domains. In the natural resources domain,
particularly in the oil and gas (O&amp;G) industry, a motivating use case is seismic data
interpretation, which is key in exploration processes that analyze geological structures
in the subsurface. Experts, supported by specialized tools and domain knowledge,
identify patterns and correlate geological factors by exploring different data sources. This
practice aims at detecting geological structures, enhancing information, correcting
potential inconsistencies in the data acquisition process, and so on. An increasing number
of works in the literature have been proposed to apply Machine Learning (ML)
workflows to support aspects of such processing. To systematically model complex data
processing pipelines, such as the ML workflows of this motivating use case, while
promoting collaboration and knowledge curation, a holistic perspective is required.</p>
      <p>In this sense, we conceptualized and developed the Cycle Orchestrator, a
knowledge-based workflow management system (WfMS) to support and operationalize
the whole lifecycle of ML and general-purpose workflows. Including specification,
setup, execution and provenance data management of such workflows. In the
motivating use case, it was conceived primarily to support O&amp;G exploration use cases that
apply cyclic ML workflows. That is, streams of ML tasks that can yield improved
results through a chain of execution iterations. These workflows are associated to
particular types of data sources, e.g. pre-stack and post-stack seismic data. The initially
considered use cases comprised unsupervised ML pipelines that train new models and
reuse pre-trained models and weights against new datasets for improving the quality by
cyclic evolution. In this context, orchestration involves the definition of what model
and version should be applied to analyze specific data sources in exploration processes.
2</p>
      <p>
        Knowledge-based Workflow Modeling
The Cycle Orchestrator draws on the Hyperknowledge [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] conceptual model for
relating knowledge specifications aligned through a domain ontology to segments of
multimodal content. In this model, the relation between nodes is done through links and
connectors. Nodes’ fragments can be referenced to with the anchor mechanism, making
the structuring of data segments explicit. In addition, the use of nested contexts as a
structuring mechanism promotes the overall organization of the knowledge base,
allowing the clustering of information through different perspectives or dimensions.
      </p>
      <p>The modeling process considers different moments of the workflow lifecycle,
namely specification, setup, execution and analysis. In the workflow specification, the
tasks are defined with their input data types, expected output data types and execution
ordering. Then, in the setup step references to actual data are connected to the tasks’
specifications. After these definitions, the execution stage may take place, followed by
a fourth moment where users inspect workflows’ output, querying and reasoning over
experiments’ results.</p>
      <p>The specification and configuration of workflows rely on the Hyperknowledge
Specification Language (HSL), a JSON-based definition scheme. Following the principles
of a glue language, it does not constrain any modeling aspects, nor define supported
data formats or specific content structuring. The language allows any modeling design
that makes sense for interested parts (users and systems) who produce and consume
information from the knowledge base. Assuming of course there is a previously agreed
ontology with well-defined terms and relations. The Cycle Orchestrator’s basic
ontology defines handy entities and connectors for specifying and configuring tasks,
execution strategies (sequential, cyclic, parallel, map-reduce like), ordering and other aspects.</p>
      <p>To illustrate modeling decisions, consider the following “hello world” example to
model a basic Python function that takes two integer numbers passed as parameters
factor_a and factor_b and writes the resulting multiplication value in a text file. In the
specification step, the basics of the task are specified, i.e. its input data types (two
numbers) and output data types (a file). No execution ordering is needed, since we have a
single task. For each input parameter identifier of the executable Python function, it
should be an anchor in the media node with the same identifier.</p>
      <p>Listing 1 depicts the HSL modeling with the specification for this executable node’s
input and output to entities representing their compatible data types. Afterwards,
Listing 2 shows the setup of the task. Following the same rationale, it connects the entities
representing actual data sources, numbers (10 and 25 as input parameters) and file
(mult_result.txt as output data) to the task specification.</p>
      <p>["context", "BasicOntology", {"subconceptOf":"Ontology"}, [
["concept", "Number", {"subconceptOf": "Attribute"}],
["concept", "SystemPath", {"subconceptOf": "Attribute"}]]
],
],
["node", "MultiplicationTask", {"instanceOf": "DataTransformation", "mimeType":
"application/wf_task", "uri": "multiply.py"}, [
["property", {"package": "my_package", "function": "multiply"}],
["anchor", "factor_a", {"type": "data", "description": "Parameter A to be multiplied"}],
["anchor", "factor_b", {"type": "data", "description": "Parameter B to be multiplied"}],
],
["link", null, {"connector": "consumes"}, [
["bind", {"task": "MultiplicationTask#factor_a"}],
["bind", {"data": "Number"}]]
],
["link", null, {"connector": "consumes"}, [
["bind", {"task": "MultiplicationTask#factor_b"}],
["bind", {"data": "Number"}]]
],
["link", null, {"connector": "produces"}, [
["bind", {"task": "MultiplicationTask"}],
["bind", {"data": "SystemPath"}]]</p>
      <p>Listing 1. HSL example for the specification of a multiplication task.
],
["link", null, {"connector": "consumes"}, [
["bind", {"task": "MultiplicationTask#factor_b"}],
["bind", {"data": " number_25"}]]
],
["link", null, {"connector": "produces"}, [
["bind", {"task": "MultiplicationTask"}],
["bind", {"data": " result_file"}]]</p>
      <p>Listing. 2. HSL example for the setup of a multiplication task.
3</p>
    </sec>
    <sec id="sec-2">
      <title>System implementation</title>
      <p>
        Information in the Cycle Orchestrator is represented in the Hyperknowledge Base, a
hybrid storage solution that uses a direct hyperlinked knowledge graph to maintain all
information about workflow execution plans and provenance data stored in the
knowledge base. The proposed modeling adheres to the MLWfM ontology [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to
structure basic aspects of ML and the PROV-ML [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as provenance data model.
      </p>
      <p>Cycle Orchestrator</p>
      <p>API</p>
      <p>Orchestrator
no Specification Controler
iiitfccape lrde SpePcaifrisceartion
loSw naH
fr
k
o
W</p>
      <p>Setup Parser
Workflow
Representation</p>
      <p>Workflow</p>
      <p>DAGs</p>
      <p>
        Users interact with the system through a REST API and a web UI for curating and
querying information, named Knowledge Explorer System (KES) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The REST API
has three endpoint sets for workflow specification, execution, and lineage retrieval. The
specification endpoints provide basic operations for workflow plans. HSL files with
workflow definitions are parsed, producing both a Hyperknowledge representation and
a directed acyclic graph (DAG) data structure. The Execution endpoint interfaces with
the execution engine’s API (Apache Airflow1) to maintain and command workflows.
The execution handler captures provenance data, structuring according to the
provenance data model that can be queried through the Lineage endpoint. Figure 1 shows
the architectural overview of the system.
      </p>
      <p>Workflow Specification API</p>
      <p>Execution Control API</p>
      <p>Lineage API
Workflow Builder</p>
      <p>Orchestrator Execution Provenance Colector</p>
      <p>Controler Service</p>
      <p>Orchestrator</p>
      <p>Lineage Controler
itxcenuo lre E(AxepcaucthioenAEirnflgoiwn)e
lfrkooEwW danH OauntpduLtoDgasta
Hyperknowledge Base</p>
      <p>Provenance</p>
      <p>Data
e
g
a
iLne lrde
fkow naH
lr
o
W</p>
      <p>Provenance
Manager Service
Knowledge Explorer</p>
      <p>System
The main goal of this demo is to showcase the Cycle Orchestrator’s approach in action,
presenting how workflows modeled with HSL definitions can be posted to the provided
API with a command-line tool, and inspected through the KES dashboard UI. For that
end, we explore a simple example for a parallel matrix multiplication workflow.
Short video: https://ibm.box.com/s/byg4wb6beo632f0d68tqg5f5jh0o8k4f
Long video: https://ibm.box.com/s/g3kupho1pn3gv3zupl2xcf1samqdfs98</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.:
          <article-title>Managing Machine Learning Workflow Components</article-title>
          .
          <source>In: 14th IEEE Conference on Semantic Computing, ICSC</source>
          . pp.
          <fpage>25</fpage>
          -
          <lpage>30</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          et al.:
          <article-title>Extending Hypermedia Conceptual Models to Support Hyperknowledge Specifications</article-title>
          .
          <source>Int. J. Semantic Computing</source>
          .
          <volume>11</volume>
          ,
          <issue>01</issue>
          ,
          <fpage>43</fpage>
          -
          <lpage>64</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>M.F.</given-names>
          </string-name>
          et al.:
          <article-title>KES: The Knowledge Explorer System</article-title>
          . In: 2018 International Semantic Web Conference (P&amp;D/Industry/BlueSky), ISWC. (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Souza</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          et al.:
          <article-title>Provenance Data in the Machine Learning Lifecycle in Computational Science and Engineering</article-title>
          . In: 2019 IEEE/
          <article-title>ACM Workflows in Support of Large-Scale Science</article-title>
          , WORKS. pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>