<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Naples, Italy
$ iacopo.colonnelli@unito.it (I. Colonnelli)
 https://alpha.di.unito.it/iacopo-colonnelli/ (I. Colonnelli)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Workflow Models for Heterogeneous Distributed Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Iacopo Colonnelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Università degli Studi di Torino, Department of Computer Science</institution>
          ,
          <addr-line>Corso Svizzera 185, 10149, Torino</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This article introduces a novel hybrid workflow abstraction that injects topology awareness directly into the definition of a distributed workflow model. In particular, the article briefly discusses the advantages brought by this approach to the design and orchestration of large-scale data-oriented workflows, the current level of support from state-of-the-art workflow systems, and some future research directions. When considering data-oriented workflows, all the aspects of data management become crucial for performance optimisation, privacy preservation and security. The data locality principle, i.e., moving computation close to the data, inspired the foundational algorithms [1] and data structures [2] of modern Big Data analysis frameworks and became an unwaivable requirement of federated learning approaches [3]. On the other hand, there are scenarios in which it is worth, or even unavoidable, to transfer data between diferent modules of a complex application. The modular nature of modern applications and the heterogeneity in contemporary hardware resources and their features, further exacerbated by the end-to-end co-design approach [4], require Workflow Management Systems (WMSs) to support a large ecosystem of execution environments (from HPC to cloud, to the Edge), optimisation policies (performance vs. energy eficiency) and computational models (from classical to quantum). For these reasons, modern workflow models and tools need to be topology-aware, allowing an explicit mapping of workflow steps onto (families of) processing elements. This mapping can be either manual, driven by the combined experience of domain experts and computer scientists, or (semi-)automatic, using advanced learning algorithms to infer the best-suited execution environment for each step. A hybrid workflow can be defined as a workflow whose steps can span multiple, heterogeneous, and independent computing infrastructures [5]. Each of these aspects has significant implications. Support for multiple infrastructures implies that each step must potentially target a diferent deployment location in charge of executing it. Locations can be heterogeneous, exposing diferent methods and protocols for authentication, communication, resource allocation and job</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Scientific Workflows</kwd>
        <kwd>HPC</kwd>
        <kwd>Cloud Computing</kwd>
        <kwd>Distributed Computing</kwd>
        <kwd>Hybrid Workflows</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>execution. Plus, they can be independent of each other, meaning that direct communications
and data transfers among them may not be allowed. A suitable model for hybrid workflows
must then be composite, enclosing a specification of the workflow dependencies, a topology of
the involved locations, and a mapping relation between steps and locations.</p>
    </sec>
    <sec id="sec-2">
      <title>2. State of the art and future directions</title>
      <p>
        Grid-native WMSs [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
        ] typically support distributed workflows out of the box, providing
automatic scheduling and data transfer management across multiple execution locations. However,
all the orchestration aspects are delegated to external, grid-specific libraries and frameworks,
limiting the spectrum of supported execution environments.
      </p>
      <p>
        Recently, a new class of topology-aware WMSs is starting to be designed and implemented,
bringing advantages in performance and costs of workflow executions on top of heterogeneous
distributed environments. StreamFlow [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] augments the Common Workflow Language (CWL)
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] open standard with a topology of deployment locations and relies on a set of connectors to
support several execution environments, from HPC queue managers to container orchestrators.
DagOnStar [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] allows users to model hybrid workflows as pure Python scripts, scheduling each
task on an HPC facility, a cloud VM, or a software container. Jupyter Workflow [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] transforms
a sequential computational notebook into a hybrid workflow by treating each cell as a workflow
step, semi-automatically extracting inter-cell data dependencies from the code, and mapping
each cell into one or more execution locations. Mashup [13] automatically maps each workflow
step onto the best-suited location, choosing between Cloud VMs and serverless platforms.
      </p>
      <p>
        Hybrid workflows proved themselves flexible enough to eficiently model and orchestrate
large-scale applications from a diverse set of domains, including bioinformatics [
        <xref ref-type="bibr" rid="ref9">9, 14</xref>
        ],
largescale scientific simulations [
        <xref ref-type="bibr" rid="ref12">15, 12</xref>
        ], and deep learning [16, 17], on top of hybrid cloud-HPC
environments. Nevertheless, the syntax and semantics used to model and execute distributed
workflows are still product-specific, hindering the portability and reusability of both workflow
models and orchestration strategies. Hybrid workflow models [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] represent a first step toward
a vendor-agnostic way to incorporate topology awareness directly in the workflow definition,
and further research eforts are ongoing to distil a formal representation of hybrid workflows,
enabling optimisation strategies with theoretical correctness and consistency guarantees [18].
Another promising research direction involves relying on topology information to improve
the overall workflow execution plan, e.g., developing location-aware scheduling algorithms or
transparently injecting streaming capabilities in file-based workflows [19].
      </p>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by the ACROSS project, “HPC Big Data Artificial
Intelligence cross-stack platform toward exascale,” and the EUPEX project, “European Pilot
for Exascale,” which have received funding from the EuroHPC JU under grant agreements No.
955648 and 101033975, respectively. Plus, it has been partially supported by the ICSC – Centro
Nazionale di Ricerca in High Performance Computing, BigData and Quantum Computing,
funded by European Union – NextGenerationEU.
10.007.
[13] R. B. Roy, T. Patel, V. Gadepally, D. Tiwari, Mashup: making serverless computing
useful for HPC workflows via hybrid execution, in: PPoPP ’22: 27th ACM SIGPLAN
Symposium on Principles and Practice of Parallel Programming, ACM, 2022, pp. 46–60.
doi:10.1145/3503221.3508407.
[14] A. Mulone, S. Awad, D. Chiarugi, M. Aldinucci, Porting the variant calling pipeline
for NGS data in cloud-hpc environment, in: 47th IEEE Annual Computers, Software,
and Applications Conference, COMPSAC 2023, IEEE, Torino, Italy, 2023, pp. 1858–1863.
doi:10.1109/COMPSAC57700.2023.00288.
[15] R. Montella, D. Di Luccio, S. Kosta, DagOn*: Executing direct acyclic graphs as parallel jobs
on anything, in: IEEE/ACM Workshop on Workflows in Support of Large-Scale Science,
WORKS@SC 2018, IEEE, 2018, pp. 64–73. doi:10.1109/WORKS.2018.00012.
[16] I. Colonnelli, B. Cantalupo, R. Esposito, M. Pennisi, C. Spampinato, M. Aldinucci, HPC
application cloudification: The StreamFlow toolkit, in: 12th Workshop on Parallel Programming
and Run-Time Management Techniques for Many-core Architectures and 10th Workshop
on Design Tools and Architectures for Multicore Embedded Computing Platforms,
PARMADITAM 2021, volume 88 of OASIcs, Schloss Dagstuhl - Leibniz-Zentrum für Informatik,
Budapest, Hungary, 2021, pp. 5:1–5:13. doi:10.4230/OASIcs.PARMA-DITAM.2021.5.
[17] I. Colonnelli, B. Casella, G. Mittone, Y. Arfat, B. Cantalupo, R. Esposito, A. R. Martinelli,
D. Medić, M. Aldinucci, Federated learning meets HPC and cloud, in: Astrophysics
and Space Science Proceedings, volume 60, Springer, Catania, Italy, 2023, pp. 193–199.
doi:10.1007/978-3-031-34167-0_39.
[18] D. Medić, M. Aldinucci, Towards formal model for location aware workflows, in: 47th
IEEE Annual Computers, Software, and Applications Conference, COMPSAC 2023, IEEE,
Torino, Italy, 2023, pp. 1864–1869. doi:10.1109/COMPSAC57700.2023.00289.
[19] A. R. Martinelli, M. Torquati, I. Colonnelli, B. Cantalupo, M. Aldinucci, CAPIO: a
middleware for transparent I/O streaming in data-intensive workflows, in: 30th IEEE International
Conference on High Performance Computing, Data, and Analytics, HiPC 2023, IEEE, Goa,
India, 2023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. Ghemawat,</surname>
          </string-name>
          <article-title>MapReduce: Simplified data processing on large clusters</article-title>
          ,
          <source>in: 6th Symposium on Operating System Design and Implementation</source>
          (OSDI
          <year>2004</year>
          ), USENIX Association, San Francisco, California, USA,
          <year>2004</year>
          , pp.
          <fpage>137</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaharia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chowdhury</surname>
          </string-name>
          ,
          <string-name>
            <surname>T. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Dave</surname>
            , J. Ma,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>McCauly</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          <string-name>
            <surname>Franklin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Shenker</surname>
            ,
            <given-names>I. Stoica</given-names>
          </string-name>
          ,
          <article-title>Resilient distributed datasets: A fault-tolerant abstraction for in-memory cluster computing</article-title>
          ,
          <source>in: Proceedings of the 9th USENIX Symposium on Networked Systems Design and Implementation</source>
          ,
          <string-name>
            <surname>NSDI</surname>
          </string-name>
          <year>2012</year>
          , San Jose, CA, USA, April
          <volume>25</volume>
          -
          <issue>27</issue>
          ,
          <year>2012</year>
          ,
          <string-name>
            <given-names>USENIX</given-names>
            <surname>Association</surname>
          </string-name>
          ,
          <year>2012</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>McMahan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ramage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hampson</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. A. y Arcas</surname>
          </string-name>
          ,
          <article-title>Communication-eficient learning of deep networks from decentralized data</article-title>
          ,
          <source>in: Proceedings of the 20th International Conference on Artificial Intelligence and Statistics</source>
          , AISTATS
          <year>2017</year>
          ,
          <volume>20</volume>
          -
          <fpage>22</fpage>
          April 2017,
          <string-name>
            <given-names>Fort</given-names>
            <surname>Lauderdale</surname>
          </string-name>
          , FL, USA,
          <year>2017</year>
          , pp.
          <fpage>1273</fpage>
          -
          <lpage>1282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gannon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Dongarra</surname>
          </string-name>
          ,
          <article-title>Reinventing high performance computing: Challenges and opportunities</article-title>
          ,
          <source>CoRR abs/2203</source>
          .02544 (
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .48550/arXiv.2203. 02544. arXiv:
          <volume>2203</volume>
          .
          <fpage>02544</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>I. Colonnelli</surname>
          </string-name>
          ,
          <article-title>Workflow models for heterogeneous distributed systems</article-title>
          ,
          <source>Ph.D. thesis, Università degli Studi di Torino</source>
          ,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .5281/zenodo.7135483.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>I. J.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , M. S. Shields,
          <string-name>
            <surname>I. Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Harrison</surname>
          </string-name>
          ,
          <article-title>The Triana workflow environment: Architecture and applications</article-title>
          , in: Workflows for e-Science,
          <source>Scientific Workflows for Grids</source>
          , Springer,
          <year>2007</year>
          , pp.
          <fpage>320</fpage>
          -
          <lpage>339</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>84628</fpage>
          -757-2\_
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Fahringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Prodan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nadeem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nerieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Podlipnig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Siddiqui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. L.</given-names>
            <surname>Truong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Villazón</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Wieczorek, ASKALON: A development and grid computing environment for scientific workflows</article-title>
          , in: Workflows for e-Science,
          <source>Scientific Workflows for Grids</source>
          , Springer,
          <year>2007</year>
          , pp.
          <fpage>450</fpage>
          -
          <lpage>471</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>84628</fpage>
          -757-2\ _
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Deelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Vahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rynge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mayani</surname>
          </string-name>
          , R. F. da
          <string-name>
            <surname>Silva</surname>
            , G. Papadimitriou,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Livny</surname>
          </string-name>
          ,
          <article-title>The evolution of the Pegasus workflow management software</article-title>
          ,
          <source>Computing in Science and Engineering</source>
          <volume>21</volume>
          (
          <year>2019</year>
          )
          <fpage>22</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .1109/
          <string-name>
            <surname>MCSE</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <volume>2919690</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>I.</given-names>
            <surname>Colonnelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Cantalupo</surname>
          </string-name>
          , I. Merelli, M. Aldinucci,
          <article-title>StreamFlow: cross-breeding cloud with HPC</article-title>
          ,
          <source>IEEE Transactions on Emerging Topics in Computing</source>
          <volume>9</volume>
          (
          <year>2021</year>
          )
          <fpage>1723</fpage>
          -
          <lpage>1737</lpage>
          . doi:
          <volume>10</volume>
          .1109/TETC.
          <year>2020</year>
          .
          <volume>3019202</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Crusoe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Abeln</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Iosup</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Amstutz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tijanic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ménager</surname>
          </string-name>
          , S. SoilandReyes,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Goble</surname>
          </string-name>
          ,
          <article-title>Methods included: Standardizing computational reuse and portability with the Common Workflow Language, Communication of the ACM (</article-title>
          <year>2022</year>
          ). doi:
          <volume>10</volume>
          .1145/ 3486897.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>D. D.</surname>
            Sánchez-Gallegos,
            <given-names>D. D.</given-names>
          </string-name>
          <string-name>
            <surname>Luccio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kosta</surname>
            ,
            <given-names>J. L. G.</given-names>
          </string-name>
          <string-name>
            <surname>Compeán</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Montella</surname>
          </string-name>
          ,
          <article-title>An eficient pattern-based approach for workflow supporting large-scale science: The DagOnStar experience</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>122</volume>
          (
          <year>2021</year>
          )
          <fpage>187</fpage>
          -
          <lpage>203</lpage>
          . doi:
          <volume>10</volume>
          .1016/J. FUTURE.
          <year>2021</year>
          .
          <volume>03</volume>
          .017.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>I.</given-names>
            <surname>Colonnelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aldinucci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Cantalupo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Padovani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rabellino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Spampinato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Morelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Di</given-names>
            <surname>Carlo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Magini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cavazzoni</surname>
          </string-name>
          ,
          <article-title>Distributed workflows with Jupyter</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>128</volume>
          (
          <year>2022</year>
          )
          <fpage>282</fpage>
          -
          <lpage>298</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.future.
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>