<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Generation of OpenCL Programs Based on Algebra-Algorithmic Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anatoliy Doroshenko</string-name>
          <email>doroshenkoanatoliy2@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksii Beketov</string-name>
          <email>beketov.oleksii@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mykola Bondarenko</string-name>
          <email>bondarenko_mykola@yahoo.com.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olena Yatsenko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Software Systems of National Academy of Sciences of Ukraine</institution>
          ,
          <addr-line>Glushkov prosp. 40, 03187 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Taras Shevchenko National University of Kyiv</institution>
          ,
          <addr-line>Glushkov prosp. 4d, 03680 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The paper proposes the further development of algebra-algorithmic design and synthesis tools towards the development of OpenCL programs. The method for semi-automatic parallelization of cyclic operators is proposed. The particular feature of the approach consists in using high-level algebraalgorithmic program specifications (schemes) and rewriting rules technique. The developed tools provide the construction of parallel algorithm schemes by superposition of predefined language constructs of Glushkov's system of algorithmic algebra, which are considered as reusable components. An algorithm scheme is a basis for the generation of corresponding source code in a target programming language. The approach is illustrated with an example of developing an OpenCL interpolation program used in a numerical weather forecasting. The results of the experiment consisting in executing the generated OpenCL program on a graphics processing unit are given.</p>
      </abstract>
      <kwd-group>
        <kwd />
        <kwd>Algorithmic algebra</kwd>
        <kwd>automated algorithm design</kwd>
        <kwd>CUDA</kwd>
        <kwd>heterogeneous platform</kwd>
        <kwd>OpenCL</kwd>
        <kwd>parallel computation</kwd>
        <kwd>software synthesis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Further progress in improving the quality of parallel software development is
associated with using heterogeneous architectures of parallel computing systems. One of the
facilities for programming heterogeneous parallel systems is OpenCL (Open
Computing Language) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], a framework for developing parallel software that executes across
platforms consisting of central processing units (CPUs), graphics processing units
(GPUs), digital signal processors (DSPs), field-programmable gate arrays (FPGAs)
and other processors or hardware accelerators. Unlike Nvidia CUDA [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which
implements GPGPU (General-Purpose computing on GPU) only for Nvidia accelerators
and requires the use of a specific compiler, OpenCL is the specification supported by
various hardware developers and only requires to set the path to the OpenCL library
at compilation. It should be noted that programming heterogeneous platforms is a
complex task and therefore there is a necessity of developing the tools for automated
software design and parallelization of existing sequential programs for such
platforms.
      </p>
      <p>
        In the previous works, we had been developing a theory, methodology and tools
for automated program design, based on the algebra of algorithmics [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. The
algorithmics formalizes the knowledge about subject domains with the help of algebraic
facilities and deals with problems of formalization, substantiation of correctness and
transformation of algorithms. The formal facilities for development of parallel
programs for multicore CPUs [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Nvidia GPUs (using CUDA) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] were developed.
They were based on Glushkov’s system of algorithmic algebra (SAA) [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] and term
rewriting technique [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. On the basis of the developed theory and methodology, the
integrated toolkit for designing and synthesis of programs (IDS) was developed [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ].
      </p>
      <p>In this paper, we propose the further development of our algebra-algorithmic
methodology and tools in the direction of formalized and automated design of
OpenCL programs. A method and a software tool intended for semi-automatic
parallelization of cyclic operators are described. The approach is illustrated with an
example of designing a parallel interpolation algorithm, which is the part of the numerical
weather forecasting program. The results of the experiment consisting in executing
the generated OpenCL program on a GPU are given.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Algebra-Algorithmic Software Design Tools</title>
      <p>
        The approach to designing parallel programs being proposed is based on Glushkov’s
system of algorithmic algebra [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], intended for formal representation of algorithmic
knowledge in the form of high-level specifications. SAA is the two-sorted algebra
GA   {Pr, Op}; GA  , where Pr and Op are the sets of logic conditions
(predicates) and operators defined on an information set; GA is the signature consisting of
logic operations (disjunction, conjunction, negation) and operator constructs, in
particular:
 serial execution of operators: “operator 1”; “operator 2” ;
 branching: IF ‘condition’ THEN “operator 1” ELSE “operator 2” END IF ;
 for loop: FOR (counter FROM start TO fin) “operator” END OF LOOP ;
 asynchronous execution of n operators: PARALLEL(i  0,..., n – 1)(“operator i”) ;
 synchronizer, which delays the computation until the value of the specified
condition is true: WAIT ‘condition’ .
      </p>
      <p>
        The algorithms, represented in SAA, are called SAA schemes. Automated design
of algorithm schemes and generation of corresponding programs is provided by the
developed IDS toolkit [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. The design process is represented by a tree of an
algorithm. The user chooses SAA constructs from the list and adds them to an algorithm
tree. On each step of the design process, the system allows a user to select only those
operations, the insertion of which into a scheme does not break its syntactical
correctness. The algorithm tree is then used for the automatic generation of SAA scheme text
and programming code in one of the target languages (C, C++, Java). The mapping of
SAA operators to text in a programming language is defined in a form of code
templates in the database of IDS. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], additional facilities for designing GPU programs
using the CUDA platform were developed.
      </p>
      <p>In this work, we introduce new SAA operations intended for high-level design of
OpenCL programs and propose the method for parallelization of nested loops in
sequential programs.</p>
      <p>
        OpenCL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] combines the application programming interface (API) and a variant
of C language for programming and simultaneous use of various parallel computing
devices (for example, CPU, GPU, and Xeon Phi) in a heterogeneous environment.
The main steps of developing a common OpenCL program and corresponding new
basic operators (enclosed in quotation marks) added to SAA are the following.
1. Obtaining the list of available platforms and saving it to a variable plms: “Get all
available platforms (plms)”.
2. Obtaining the list of devices dvs of the platform plm: “Get all devices (dvs)
available on a platform (plm)”.
3. Creating the computing environment for the devices: “Create a context (cnt) for
devices (dvs)”.
4. Creating the queue of commands for the device: “Create a command queue
(cmdqueue) for a context (cnt) and a device (dv)”.
5. Designing a kernel function — a task to execute on the devices. The example of an
      </p>
      <p>
        SAA scheme for a kernel is given in Sect. 3.
6. Compilation of a file containing an OpenCL kernel: “Create a kernel (krnl) from a
source (file)”, where krnl is the name of a kernel function, file is the path to the file
with its source code.
7. Creating the buffer for data specified in a variable var: “Create a memory buffer
(buff) for data (var) on devices in the context (cnt)”.
8. Copying a buffer for a variable var to device memory: “Add commands to queue
(cmdqueue) writing a buffer of data (var) from host to device”.
9. Setting the value arg_value of the parameter with the index arg_index for a kernel:
“Set the argument value (arg_value) for a parameter (arg_index) of a kernel
(krnl)”.
10. Adding a kernel to a command queue and its asynchronous execution: “Add a
command to queue (cmdqueue) executing a kernel (krnl) (globalworksize)
(localworksize) on a device”, where globalworksize is the global number of
dimensions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for executing the kernel; localworksize is the number of dimensions of a
local subset of kernel instances (work-items) in a group.
11. Reading the data from a result buffer: “Add commands to queue (cmdqueue)
reading from a buffer of data (var) from device to host”.
      </p>
      <sec id="sec-2-1">
        <title>The use of the above constructs is illustrated in Sect. 3.</title>
        <p>
          In most computing problems, a large part of hardware resources are utilized by
computations inside loops, therefore the use of automatic parallelization of the cyclic
operators is most efficient for them. Further, we describe the developed method and
software framework named LoopRipper intended for semi-automatic parallelization
of programs containing loop operators for heterogeneous platforms. The framework is
based on the combined use of the rewriting rules system TermWare [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and IDS
toolkit. LoopRipper parallelizes the input compound loop of the following form:
“SEQUENTIAL LOOP”
==== FOR (i0 FROM 0 TO #I0 )
        </p>
        <p>FOR (i1 FROM 0 TO #I1)
...</p>
        <p>FOR (in FROM 0 TO #I N )</p>
        <p>“F (i , pin (i ), pout (i ))”</p>
      </sec>
      <sec id="sec-2-2">
        <title>END OF LOOP,</title>
        <p>(1)
where I0 , I1, ..., I N are the sets values of indices i0 , i1, ..., iN ; N +1 is the loop
nesting depth; i  {i j j  0,..., N} ; p(i )  P  {p j(i ) j  0... # P}, {in, out} are
the sets of input and output data variables; F is the function processing the data; the
iterations of the loop are independent.</p>
        <p>Let T be the number of threads which will be used in the execution of a kernel
function. The loop obtained as a result of the parallelization is the following:
“PARALLEL LOOP”
==== FOR (e FROM 0 TO L)
“ fill(inBuf )”;
“ push(inBuf )”;
“kernel(e, inBuf , outBuf )”;
“ pull(outBuf )”;
“unpack(outBuf )”</p>
      </sec>
      <sec id="sec-2-3">
        <title>END OF LOOP,</title>
        <p>where L is the number of kernel calls which is chosen the least possible so that
L T  G , G is the total number of initial loop iterations; fill is the function filling
the buffer of initial data inBuf ; push is moving the initial data buffer to the device
memory; kernel is the call of a kernel function; the kernel function contains the call
of the initial loop body “F(i ,inbufid (i ), outbufid (i ))”, where id is the thread
number; pull is the function moving the processed data from the device memory to
the buffer of processed data outBuf ; unpack is copying data from the buffer of
processed data to corresponding variables.</p>
        <p>
          The user of the LoopRipper framework (see Fig. 1) provides a source code or an
SAA scheme of the sequential program, specifies the loop to be parallelized and also
provides the list of input and output variables used in the loop. With the help of the
TermWare system, the framework generates kernel, fill and unpack functions and
additional data structures (buffers). IDS toolkit replaces the sequential loop in the
sequential program with corresponding kernel call and synthesizes the whole parallel
program from the above-mentioned functions and data structures.
In this section, we illustrate the use of the considered algebra-algorithmic tools with
an example of developing an OpenCL interpolation program used in numerical
weather forecasting [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The program performs the interpolation of meteorological
values defined on a macroscale grid to a mesoscale grid.
        </p>
        <p>The initial sequential interpolation scheme consists of four nested loops of the
form (1) with iterators h [0 ... Pk ] , k [0 ... Lmz] , j [0 ... Mmz] , i [0 ... Nmz] ,
where Pk , Lmz , Mmz , Nmz are input integer parameters defining the number of
points in finite-difference grid along altitude, longitude and latitude. The iterations of
the loops are independent and can be executed in parallel. The SAA scheme of the
OpenCL kernel obtained as a result of the transformation of the sequential
interpolation algorithm with the help of the LoopRipper framework is given below. The
computation is parallelized by indices h , k , j ; the main loop of interpolation iterates
through the index i . In the scheme, US , VS , TS , HS , QS are input arrays with
meteorological values (wind velocity, temperature, humidity, etc.); Qc is the output
array. The index for processing array elements is computed and stored in the
variable ind .
“interpolation_kernel(US, VS, HS, QS, TS, F_x, Zmz, Qc, Pk, Lmz, Mmz, Nmz)”
==== (h := “Get the global work-item identifier for dimension (0)”);
(k := “Get the global work-item identifier for dimension (1)”);
(j := “Get the global work-item identifier for dimension (2)”);
IF NOT((h &gt;= Pk) OR (k &gt;= Lmz) OR (j &gt;= Mmz))
THEN</p>
        <p>FOR (i FROM 0 TO Nmz)
LOOP
“(ind) := (h + Pk * k + Pk * Lmz * j + Pk * Lmz * Mmz * i)”;
“(a) := ((WZZ + US[ind] / 0.321) * Rs * VS[ind])”;
“(Tp) := (TS[ind] * pow(1000.0 / HS[ind], 2.0 / 7.0))”;
“(Tv) := (Tp * (1.0 + 0.6078 * QS[ind]))”;
“(Qc[ind]) := (a – (0.5*Tv + (1.0 – Zmz[k]) * g * F_x[j + I*Mmz]/0.321))”
END OF LOOP</p>
        <p>END IF</p>
        <p>The SAA scheme of the host code of the interpolation program uses the operators
described in Sect. 2 and is the following:
“interpolation_host”
==== “Read input parameters (Pk, Lmz, Mmz, Nmz) from the command line”;
“Get all available platforms (plms)”;
“Get all devices (dvs) available on a platform (plms[0])”;
“Create a context (cnt) for devices (dvs)”;
“Create a command queue (cmdqueue) for a context (cnt)</p>
        <p>and a device (dvs[0])”;
“Create a kernel (“interpolation”) from a source (“kernels/interpolation.cl”)”;
“Declare and fill continuous arrays for data (US,VS,HS,QS,TS,Qc,F_X,Zmz)”;
“Create memory buffers for data (US, VS, HS, QS, TS, Qc, F_X, Zmz) on
devices in the context (cnt)”;
“Add commands to queue (cmdqueue) writing buffers of data</p>
        <p>(US, VS, HS, QS, TS, F_X, Zmz) from host to device”;
“Set the argument values for parameters of a kernel (krnl)”;
“Add a command to queue (cmdqueue) executing a kernel</p>
        <p>(krnl) (3) ({Pk, Lmz, Mmz}) on a device”;
WAIT ‘All commands in (cmdqueue) were issued to device and completed’;
“Add commands to queue (cmdqueue) reading from a buffer of data (Qc)
from device to host”;</p>
        <p>
          IDS toolkit automatically translated the above SAA schemes into OpenCL code.
The obtained program was executed in two computing environments. The first
environment consisted of Intel Core i5-4210U CPU (2 cores) and Nvidia GeForce 840M
GPU (384 cores), the second one contained Intel Core i3-3110M CPU (2 cores) with
integrated Intel HD Graphics 4000 GPU (128 cores). Fig. 2(a) shows the
multiprocessor speedup Sp  Ts / Tp obtained in the first environment, where Ts is the execution
time of the sequential program on CPU, Tp is the execution time of the parallel
program (OpenCL and its previous version implemented in CUDA [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]) on GPU. The
previous version used multi-dimensional arrays for storing meteorological data
instead of one-dimensional. As can be seen, the OpenCL program significantly
outperforms the CUDA version, which is explained by the use of array linearization.
Fig. 2(b) shows the dependency of CPU and GPU execution time on the percentage of
data passed to CPU in the second environment. As can be seen from the diagram, the
execution of all computations on CPU is approximately 7 times longer than execution
on GPU only. However, if computations are performed simultaneously on CPU and
GPU, the optimization by about 7–10% is achieved.
        </p>
        <p>
          (a)
(b)
The proposed approach is related to works on the synthesis of programs from
specifications [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and automated generation of OpenCL programs [
          <xref ref-type="bibr" rid="ref10 ref11 ref12 ref8 ref9">8–12</xref>
          ]. In particular,
paper [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] presents an approach and a tool named Gaspard2 based on model-driven
engineering to specify, design, and generate OpenCL applications. In [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], a programming
tool called STEPOCL is proposed along with a new domain-specific language
designed to simplify the development of an application for multiple accelerators.
Paper [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] presents an automatic generator based on a C++ domain-specific embedded
language for the generation of OpenCL kernels defined by high-level specifications
provided by the user. In [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], an approach to the generation of OpenCL code based on
high-level functional expressions and rewriting rules is proposed. Par4All [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] is an
automatic parallelizing and optimizing compiler for C and Fortran which generates
OpenMP, CUDA and OpenCL source codes.
        </p>
        <p>
          The main difference of our approach from the mentioned related works is that it
uses algebraic specifications, based on Glushkov algebra of algorithms. Specifications
are represented in a natural linguistic form simplifying understanding of algorithms
and facilitating achievement of demanded software quality. Another advantage of our
tools is the method of automated design of syntactically correct algorithm
specifications [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ], which eliminates syntax errors during construction of algorithm schemes.
Unlike the above-mentioned Par4All, our parallelization framework LoopRipper
allows processing the amounts of data which exceed the GPU memory size and using
heterogeneous computing clusters.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>This paper proposes an approach and tools for automated designing and generation of
OpenCL programs based on algorithmic algebra. The method for semi-automatic
parallelization of cyclic operators is developed. The particular feature of the approach
consists in using high-level algebra-algorithmic program specifications, which are
represented in a natural linguistic form. The developed tools provide the construction
of algorithm schemes by superposition of predefined language constructs of
Glushkov’s algebra, which are considered as reusable components. The developed
tools automatically translate the specifications into source code in a programming
language. The approach is illustrated on developing the OpenCL interpolation
program used in a numerical weather forecasting.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>OpenCL</given-names>
            <surname>Overview</surname>
          </string-name>
          .
          <article-title>The open standard for parallel programming of heterogeneous systems</article-title>
          , https://www.khronos.org/opencl, last accessed
          <year>2019</year>
          /02/15.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Nvidia</surname>
            <given-names>CUDA</given-names>
          </string-name>
          technology, http://www.nvidia.com/cuda, last accessed
          <year>2019</year>
          /02/15.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Andon</surname>
            ,
            <given-names>P.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.Yu.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhereb</surname>
            ,
            <given-names>K.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yatsenko</surname>
            ,
            <given-names>O.A.</given-names>
          </string-name>
          :
          <article-title>Algebra-algorithmic models and methods of parallel programming</article-title>
          .
          <source>Akademperiodyka</source>
          ,
          <string-name>
            <surname>Kyiv</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhereb</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yatsenko</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Developing and optimizing parallel programs with algebra-algorithmic and term rewriting tools</article-title>
          . In: Ermolayev,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Mayr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.C.</given-names>
            ,
            <surname>Nikitchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Spivakovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Zholtkevych</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. (eds.) ICTERI</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>CCIS</article-title>
          , vol.
          <volume>412</volume>
          , pp.
          <fpage>70</fpage>
          -
          <lpage>92</lpage>
          . Springer, Cham (
          <year>2013</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -03998-
          <issue>5</issue>
          _
          <fpage>5</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Prusov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Computational techniques for modeling atmospheric processes</article-title>
          .
          <source>IGI Global</source>
          ,
          <string-name>
            <surname>Hershey</surname>
          </string-name>
          (
          <year>2018</year>
          ). https://doi.org/10.4018/978-1-
          <fpage>5225</fpage>
          -2636-0
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.Yu.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yatsenko</surname>
            ,
            <given-names>O.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beketov</surname>
            ,
            <given-names>O.G.</given-names>
          </string-name>
          :
          <article-title>Algorithm for automatic loop parallelization for graphics processing units</article-title>
          . Problems in programming,
          <source>(4)</source>
          ,
          <fpage>28</fpage>
          -
          <lpage>36</lpage>
          (
          <year>2017</year>
          )
          <article-title>(in Ukrainian)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gulwani</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Dimensions in program synthesis</article-title>
          .
          <source>In: 12th international ACM SIGPLAN symposium on Principles and practice of declarative programming</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>24</lpage>
          . ACM, New York (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guyomarc'h</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dekeyser</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>An MDE approach for automatic code generation from UML/MARTE to OpenCL</article-title>
          . Computing in Science and Engineering,
          <volume>15</volume>
          (
          <issue>1</issue>
          ),
          <fpage>46</fpage>
          -
          <lpage>55</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunet</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trahay</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parrot</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Namyst</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Automatic OpenCL code generation for multi-device heterogeneous architectures</article-title>
          .
          <source>In: 44th International Conference on Parallel Processing (ICPP</source>
          <year>2015</year>
          ), pp.
          <fpage>959</fpage>
          -
          <lpage>968</lpage>
          . IEEE, Piscataway, NJ (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tillet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rupp</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Selberherr</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An automatic OpenCL compute kernel generator for basic linear algebra operations</article-title>
          .
          <source>In: 2012 Symposium on High Performance Computing (HPC'12)</source>
          , pp.
          <volume>4</volume>
          :
          <fpage>1</fpage>
          -
          <issue>4</issue>
          :
          <fpage>2</fpage>
          . Society for Computer Simulation International, San Diego, CA (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Steuwer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fensch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindley</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubach</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Generating performance portable code using rewrite rules: from high-level functional expressions to high-performance OpenCL code</article-title>
          .
          <source>In: 20th ACM SIGPLAN International Conference on Functional Programming (ICFP'15)</source>
          ,
          <source>ACM SIGPLAN Notices</source>
          , vol.
          <volume>50</volume>
          , pp.
          <fpage>205</fpage>
          -
          <lpage>217</lpage>
          . ACM, New York (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. PIPS:
          <article-title>Automatic Parallelizer</article-title>
          and Code Transformation Framework, http://pips4u.org,
          <source>last accessed</source>
          <year>2019</year>
          /02/15.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>