<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Adaptation of Legacy Fortran Applications to Cloud Computing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eugene Tulika</string-name>
          <email>eugene.tulika@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anatoliy Doroshenko</string-name>
          <email>doroshenkoanatoliy2@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kostiantyn Zhereb</string-name>
          <email>zhereb@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Software Systems of National Academy of Sciences of Ukraine</institution>
          ,
          <addr-line>Glushkov prosp. 40, 03187 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Key Terms. ConcurrentComputation</institution>
          ,
          <addr-line>DataGrid, HighPerformanceComputing, FormalMethod, SoftwareSystem</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>21</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>We propose an approach to the semi-automatic transformation of legacy Fortran applications for execution on cloud computing platforms. An architecture is proposed based on web-services choreography, which allows unlimited scalability of the system and reduces overhead on message passing. The approach is tested on an example program from the quantum chemistry field.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>virtualization</kwd>
        <kwd>cloud computing</kwd>
        <kwd>scalable parallelism</kwd>
        <kwd>web-services choreography</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Fortran language exists since the 1950s and during this time positioned itself as one of
the best tools for scientific research. Language has huge support from the software
industry, new libraries and compilers are released that allow using modern
technologies and standards. Examples are: compiler from Portland Group [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for Fortran
supporting GPGPU and CUDA; High Performance Fortran Forum [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] creates compilers
and standards for high performance Fortran; OpenMP and MPI for parallel and
distributed computation allow to use Fortran for modern cluster computing; Intel
popularizes Fortran [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in order to support proprietary Intel compilers and programs;
Coarray Fortran language fork became a standard in Fortran 2003 and allows to
support computation on distributed arrays [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]; there are many Fortran libraries [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for
numerical analysis. Besides, conservative policy for backward compatibility makes
code written on old standards compatible with the new compilers. All this makes
Fortran an attractive platform for scientific research. However, Fortran suffers from
outdated standards that make it hard to write an efficient code for computations with
distributed memory. Another problem is a great amount of legacy code which was
written without taking into account distributed architectures.
      </p>
      <p>Cloud computing became popular because of the need to reduce the cost of
computations. The difference between cloud computing and other parallel approaches is
usage of scalability (ability to support more resources) vs. performance (running fast</p>
      <p>- 112
on given resources). Scalable systems can contain components which individually
perform poorly. Cloud computing is focused on reducing the cost and
time/performance balance using cheap components, and therefore allows using more
resources.</p>
      <p>This paper describes work in progress aimed at parallelizing Fortran scientific
applications for the cloud platform. We propose a methodology for semi-automatic
parallelization of legacy applications. The paper also describes an architecture for
running distributed applications on the cloud.</p>
      <p>
        Related parallelization approaches for porting legacy applications to cloud
platform have been described, such as Pydron [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and Bio-Cirrus [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The contribution of
this paper is 1) using rewriting rules technique to automate transformation steps and
2) using service choreography instead of orchestration to reduce overhead.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Description of the Studied Application and Problem</title>
    </sec>
    <sec id="sec-3">
      <title>Statement</title>
      <p>In this paper, we discuss our approach of transforming legacy scientific applications
based on an example Fortran application from the quantum chemistry field. Some
properties we discuss are specific to this application, such as time distribution of
subroutines (Section 2) or the structure of control flow graph (Section 3). However, our
approach can be applied in the more general case.</p>
      <p>
        The example application calculates the geometry of electron orbitals [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The
application was initially optimized for single-core performance, without any parallelism.
The sequential processing time has quadratic increase depending on the input size.
      </p>
      <p>The first step of our approach requires identifying the most promising subroutines
for parallelization. The example program consists of data input, computation
subroutines, logging intermediate results into the file and data output. Profiling the
application before optimization shows the time distribution for the main computational steps:
 Data input from files and initialization – approximately 1% of time;
 Subroutine hcore calculates integrals for every atom – 60% of time;
 Subroutine iterc – optimization of the geometry of molecule, 30% of the time.</p>
      <p>Therefore, the main focus of parallelization would be spent in subroutines hcore
and iterc. Time performance of algorithms in both subroutines has a linear
dependency on the input size. Each of two subroutines is applied independently for every
atom in the input, which allows parallelization. However, iterc depends on the
results of hcore, because the calculation of the Fock operator requires integrals
computed in hcore.</p>
      <p>
        The generic question of Fortran parallelization for distributed architecture is widely
studied [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. As a part of this paper, Asynchronous Network technique with the
usage of web-service choreography is applied to the parallel algorithms. This provides
a robust framework for the development of applications which can be distributed
across cloud infrastructure without additional efforts, while usage of MPI for the
highly scalable applications requires manual control over data placement and
interprocess communication.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Equivalent Code Transformation for Scalable Parallelization</title>
      <p>On the next step, we try to find and parallelize independent loops that allow unlimited
scalability. Structurally, the application can be modeled with the following graph:
I -&gt; A ([a1..n ]) -&gt; B ([b1..n ]) -&gt; O
(1)</p>
      <p>Graph vertices I, A, B, O stand for the sequential steps of computation: I – data
input and allocation of the memory, A – calculation of the integrals for every atom
(hcore), B – calculation of the Fock operator for every atom (iterc), O – final
steps and data output. Graph edges correspond to the control flow between parts of
the program, a1..n and b1..n stand for input data of subroutines. The goal of this step is
to transform program to the parallel form:</p>
      <p>I -&gt; A1..n ([a1..n ]) -&gt; B1..n ([b1..n ]) -&gt; O
(2)</p>
      <p>We transform sequential loop A, into n parallel processes A1..n([a1..n ]). Each process
Ai will perform calculations on a single segment of data ai. The same transformation
will be applied to B as well. After manual analysis of the data dependencies, it was
identified that there are no dependencies between iterations of the loop in code
fragment A, and processes A1..n can be invoked in parallel, same applicable to the
processes B1..n . However, B had a dependency on the results of A and in order to address this
dependency barrier synchronization will be used between parallel processes Ai and Bi.</p>
      <p>After selection of the model for parallel computation, we should ensure that code
fragments invoked in parallel do not have any side effects. This is done by replacing
subroutines with pure functions (introduced in Fortran 95 standard). The following
changes in code are needed:
1. Create FUNCTION instead of SUBROUTINE
2. Remove IMPLICIT statements – all local variables of the function should be
explicitly declared.
3. Remove COMMON BLOCKs (global variables) – all corresponding global
variables should be passed as an inputs/outputs of the function, and reads/writes to
global variables performed by the calling code.
4. Remove read and write operations – all such operations should be invoked before
and after the call to the pure function.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Automatic Code Transformation Using TermWare</title>
      <p>
        TermWare [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is a tool for automatic code transformation which can be applied to the
task of transformation to pure functions. TermWare allows describing transformations
in a declarative way which simplifies its development and makes them reusable. In
the earlier work [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] TermWare was used to build high-level algebraic models of
Fortran code and perform transformations on them. This paper uses a similar
approach but focuses on the transition of the subroutines to the pure functions.
      </p>
      <p>As an example of transformation we use a set of rules to transform IMPLICIT
statements to explicit declarations of variables:
1. _MarkPure(Subroutine($name,$params,$return,$body))
-&gt; Function($name,$params,$return,_MkImp($body))
2. _MkImp([$x:$y]) -&gt; [_MkImp($x):_MkImp($y)]
3. _MkImp(NIL) -&gt; NIL
4. _MkImp(Declare($var,$type,$val) -&gt;</p>
      <p>Declare($var,$type,$val) [check($var, $type)]
5. [_MkImp(Assign($var,$expr)): $y] [isUnchecked
($var)] -&gt; [Declare_MARK($var, $type) : [Assign
($var,$expr) : $y ]] [inferType ($var,$type)]
6. [$x:[Declare_MARK($var,$type) : $y]] -&gt;</p>
      <p>[Declare_MARK($var,$type):[$x:$y]]
7. Function ($name,$params,$return,
[Declare_ MARK($var,$type):$y]) -&gt; Function
($name,$params,$return,[Declare($var,$type):$y])
Rule 1 triggers transformation, marking the body of the function with the marker
term _MkImp. Rules 2 and 3 walk through the body of the function and expand the
_MkImp marker to all operations. Rule 4 memorizes the variables which have explicit
declaration using the method check($var, $type) from the facts DB. Rule 5
finds variables without explicit declaration using method isUnchecked($var).
For these variables it determines the type with the method
inferType($var,$type) and adds declaration marked with
Declare_MARK($var,$type). To determine the variable type, the method checks
the variable name against IMPLICIT statements in Fortran code, as well as default
convention that declares variables starting with 'i'-'n' as INTEGER and all others
as REAL. Rule 6 moves this declaration to the beginning of the function, and rule 7
removes the mark. As a result, term Declare($var,$type) is generated, which
later is transformed to the declaration of the variable in the code.</p>
      <p>As an example of rules application, consider a simple procedure of square matrix
multiplication (Table 1).</p>
      <p>
        In the initial code, some variables are used with no declaration. After
transformation, all the variables are declared. Also, note that the syntax for
SUBROUTINE is different from the FUNCTION. But such changes should not be
described as additional rules. A simple substitution of the term “Subroutine” with the
term “Function” is sufficient. During the code generation phase, all necessary changes
are added automatically, which is one of the advantages of the TermWare and
highlevel algebraic models [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Transition to Distributed Application Executed on Cloud</title>
      <p>
        Previously discussed transformation steps are not specific to any given parallel
platform. Starting from this section, we discuss additional steps needed for the cloud
platform. In order to transform application to distributed architecture, its source code has
to be transformed to support network calls. Functions Ai and Bi are converted into
web-services – separate programs with HTTP interface which can be invoked
remotely. The body of the program is transformed into transaction script which invokes
remote web-services and aggregates results. In order to use HTTP calls from the
Fortran, libcurl library is used with the C interface [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Transformations of the
functions to the separate web-services is done with the Java wrapper. Data for invocation
of the remote services is composed in the transaction script. The script collects all
input data of the program and sends messages to the remote services.
      </p>
      <p>
        Cloud platform operation system, such as CloudStack, provides APIs to do scaling
– provisioning of the nodes with the predefined configuration on demand. This
capability is used by transaction script: after reading the input data and extracting the
number of atoms, it makes a call to API of the cloud operating system and requires a
startup of the necessary amount of nodes of needed type. In the simplest case,
processing of N atoms will require 2N nodes, one for each Ai and Bi. However, the
number of nodes, time they are actively working and the size of the used memory affects
the cost of computation. For optimization of the cost the optimal parameters of
configuration should be chosen, so that cost is minimized. In [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] a method of
performance optimization of the service-oriented program was proposed based on load
estimation. A similar approach can be used for minimization of the cost.
      </p>
      <p>
        Approach when transaction script calls remote web-services in a service-oriented
architecture is called Orchestration. Usage of Orchestration has disadvantages. Paper
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] shows that usage of separate transaction script increases the amount of calls
between processes in most patterns of message passing in distributed systems, which
increases overhead on data processing.
6
      </p>
    </sec>
    <sec id="sec-7">
      <title>Transition to Choreography</title>
      <p>
        Our goal is to reduce message passing overhead and eliminate a single point of failure
represented by transaction script, by using choreography. From the perspective of
distributed systems modeling, Choreography could be represented as an asynchronous
network [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The asynchronous network consists of the set of processes that
communicate with each other. Communication can include: direct message exchange
between nodes; broadcast when a node sends a message to each node including itself;
multicast when message is sent to the subset of nodes. During the transition to
choreography, transaction script is eliminated and its responsibilities are distributed
between services. Unlike Orchestration or MPI where transaction script is waiting for
the results of web-services execution, services in Choreography don't know about
each other and send results of execution to the communication channel. During the
execution, the service takes into account the state of the process and type of inbound
message in order to determine its position relatively to other services. Having
understood its position and taking into account communication protocol, service decides on
what type of message should be sent back to the channel when the result is ready. The
protocol has the following format: “If current process role is Ai and message M1 is
received as input, procedure F should be invoked and message M2 should be passed to
the processes A2..m”. This format is identical to the description of finite state machine.
      </p>
      <p>Communication Protocol can replace the part of transaction script responsible for
service invocation and results processing. Part of responsibilities related to data input
is transferred to web-service itself. Code fragment I performing initial data processing
is implemented by the new service I1 which plays the role of starting point in the
machine description. Invocation of tail fragment O which outputs the result should be
implemented by a separate service O1, which will be executed after synchronizing all
the processes B1..n and will do the post-processing of the data and data output.</p>
      <p>In order to follow the protocol, every web-service should have a controlling
mechanism which executes state machine by the description of the protocol. Controlling
mechanism is written in Java. Choreography execution is initiated by an initial event,
e.g. the file with input data has been uploaded to the cloud data volume. Process I1 in
the waiting state receives this event from the file system and invokes Fortran code
responsible for initial data processing. After the event is processed, process I1 sends
broadcast message with the results of invocation to all the processes of the
asynchronous network. In the simplest case, this message is received by all the
processes A1..n , B1..n, I1, and O1. The processes B1..n, O1, and I1 ignore this message because
Bi and O1 processes do not have enough information yet to start, and process I1
already finished its part of the work. It means that only processes A1..n start execution.</p>
      <p>
        An important aspect of choreography is the implementation of barrier
synchronization. In [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] global and local synchronizers are described that can be used for
synchronization of the processes of the asynchronous network. If global synchronizer is
used, then the separate process should be responsible for controlling the conditions of
the barrier. The local synchronizer is controlling just its neighbors. In this case
message "process reached barrier" is gradually distributed through the network which
allows reducing the amount of interactions between services.
      </p>
      <p>In order to achieve better resource utilization using Choreography, following
considerations should be taken into account. If certain web-service, such as I1, is expected
to consume fewer resources than others, it should run on the smaller node (with less
RAM and CPUs). If some web-services are never expected to run simultaneously,
such as Ai and Bi, they should re-use the same node. Ideally, it should be the same
service with a same controlling mechanism which can play different roles in the
protocol. This will guarantee that minimal amount of nodes are started at any point of
process execution. Choreography allows reducing the number of messages passed
between transaction script and web-services. It also allows better resource utilization
having non-blocking requests between web-services. Besides, it provides integration
framework with the generic rule sets defined which reduce the amount of code needed
to be written to convert the application into distributed system.
7</p>
    </sec>
    <sec id="sec-8">
      <title>Testing of the Approach</title>
      <p>In order to verify the proposed approach, we used a simplified task with the same
model of the data dependency – Gaussian elimination. Testing was performed on the
Amazon cloud platform and compares two different configurations of the system.
Both configurations have the same amount of processors – 8, and the same amount of
RAM – 32Gb. The first configuration consists from one server AWS m4.2xlarge (26
ECUs, 8 vCPUs, 2.4 GHz, Intel Xeon E5-2676v3, 32 GiB memory, EBS only) and is
used to run the sequential program. Second configuration consists of four servers
AWS m4.large (6.5 ECUs, 2 vCPUs, 2.4 GHz, Intel Xeon E5-2676v3, 8 GiB
memory, EBS only) and it is used to run the service-oriented application (with
choreography). We have measured the time spent on the processing of the square matrices
of different sizes. Comparison of execution time is in Fig. 1. For smaller matrix size,
the sequential program runs faster because of overhead in a parallel program.
However, for larger matrices the execution time of sequential program grows faster, and it
becomes less efficient.
The paper describes the work in progress of scaling legacy Fortran code using cloud
platforms. Proposed architecture uses choreography of web-services which allows
unlimited scalability and reduces overhead on message passing. Scaling exercise is
performed on application from quantum chemistry field for calculation of atoms
orbitals. One of the main results of the paper is a methodology for adjustment of the
legacy source code to the cloud infrastructure, including transition steps to distributed
scalable architecture. Our future research directions include automating additional
transformation steps using TermWare framework, applying our approach to different
applications, as well as testing different cloud configurations to find the most efficient
ways of parallelizing legacy applications.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>PGI</given-names>
            <surname>Compilers</surname>
          </string-name>
          &amp; Tools, http://www.pgroup.com/products/pvf.htm
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>High</given-names>
            <surname>Performance</surname>
          </string-name>
          , http://hpff.rice.edu
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <article-title>Fortran is more popular than ever; Intel makes it fast</article-title>
          , https://software.intel.com/enus/blogs/2011/09/24/fortran-is
          <article-title>-more-popular-than-ever-intel-makes-it-fast</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>4. Coarrays in the next Fortran Standard</article-title>
          , ftp://ftp.nag.co.uk/sc22wg5/
          <fpage>N1751</fpage>
          - N1800/N1787.pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Netlib</given-names>
            <surname>Repository</surname>
          </string-name>
          , http://netlib.org
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>S. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alonso</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Csillaghy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Pydron: Semi-automatic parallelization for multi-core and the cloud</article-title>
          .
          <source>In: 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI 14)</source>
          , pp.
          <fpage>645</fpage>
          -
          <lpage>659</lpage>
          .(
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Karlsson</surname>
            ,
            <given-names>T. J. M.</given-names>
          </string-name>
          , et al.:
          <article-title>Bio-Cirrus: a framework for running legacy bioinformatics applications with cloud computing resources</article-title>
          .
          <source>In Advances in Computational Intelligence</source>
          , pp.
          <fpage>200</fpage>
          -
          <lpage>207</lpage>
          . Springer Berlin Heidelberg. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khavryuchenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iegorov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suslova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Modeling for quantum chemistry computations (in Russian)</article-title>
          .
          <source>Upravliayushchie sistemy i mashiny N 5</source>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>87</lpage>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Doroshenko</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shevchenko</surname>
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A Rewriting Framework for Rule-Based Programming Dynamic Applications</article-title>
          .
          <source>Fundamenta Informaticae</source>
          <volume>72</volume>
          (
          <issue>1-3</issue>
          ), pp.
          <fpage>95</fpage>
          -
          <lpage>108</lpage>
          . (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tulika</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhereb</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Fortran Programs Parallelization Using Rewriting Rules Technique (in Ukrainian)</article-title>
          .
          <source>Problems in Programming N 2-3</source>
          , pp.
          <fpage>388</fpage>
          -
          <lpage>397</lpage>
          .
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <article-title>Libcurl - the multiprotocol file transfer library</article-title>
          , http://curl.haxx.se/libcurl
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Tulika</surname>
          </string-name>
          , E.:
          <article-title>Performance Optimization in SOA Using Load Estimation and Load Balancing (in Ukrainian)</article-title>
          .
          <source>Problems in Programming, N 2-3</source>
          , pp.
          <fpage>193</fpage>
          -
          <lpage>201</lpage>
          .
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Barker</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weissman</surname>
            ,
            <given-names>J.B. and Van</given-names>
          </string-name>
          <string-name>
            <surname>Hemert</surname>
            ,
            <given-names>J.I..</given-names>
          </string-name>
          :
          <article-title>Reducing data transfer in serviceoriented architectures: The circulate approach</article-title>
          .
          <source>Services Computing, IEEE Transactions on Services Computing, N</source>
          <volume>5</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>437</fpage>
          -
          <lpage>449</lpage>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Peltz</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Web services orchestration and choreography</article-title>
          .
          <source>Computer</source>
          , (
          <volume>10</volume>
          ), pp.
          <fpage>46</fpage>
          -
          <lpage>52</lpage>
          . (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lynch</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Distributed algorithms</article-title>
          . Morgan Kaufmann, San Francisco (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Golub</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <source>Ortega J.: Scientific Computing: An Introduction with Parallel Computing</source>
          ,
          <fpage>215</fpage>
          -
          <lpage>236</lpage>
          . Academic Press (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Schelter</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boden</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schenck</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alexandrov</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markl</surname>
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Distributed matrix factorization with MapReduce using a series of broadcast-joins</article-title>
          .
          <source>In Proceedings of the 7th ACM conference on Recommender systems (RecSys '13)</source>
          . ACM, New York, NY, USA, pp.
          <fpage>281</fpage>
          -
          <lpage>284</lpage>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Barker</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walton</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robertson</surname>
            <given-names>D.</given-names>
          </string-name>
          : Choreographing Web Services.
          <source>IEEE Transactions on Services Computing, N</source>
          <volume>2</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>152</fpage>
          -
          <lpage>166</lpage>
          . IEEE Computer Society. (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>