<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>” CIRP Journal of Manufacturing Science and
Technology</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">2079-9292</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/STC-CSIT.2019.8929750</article-id>
      <title-group>
        <article-title>MapReduce Performance on Distributed Heterogeneous Platform through Integration with Hadoop Cluster</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vincent Karovič</string-name>
          <email>vincent.karovicml@fm.uniba.sk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mateusz Kaźmierczak</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr Pankiv</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maciej Górkiewicz</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maryana Zakharchuk</string-name>
          <email>maryana.zk@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roksolyana Stolyarchuk</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Comenius University in Bratislava, Faculty of Management</institution>
          ,
          <addr-line>820 05 Bratislava</addr-line>
          ,
          <country country="SK">Slovakia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>S. Bandera str., 12, Lviv, 79013</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>OpenCL</institution>
          ,
          <addr-line>CUDA, performance evaluation, Hadoop cluster</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Technical University of Lodz 90-924 Łódź</institution>
          ,
          <country country="PL">Poland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <volume>2</volume>
      <issue>11</issue>
      <fpage>32</fpage>
      <lpage>36</lpage>
      <abstract>
        <p>An effective processing of Big Data in various application areas is an important task today. Modern development of information technology provides the ability to calculate a large number of different tasks using a certain number of computers in distributed mode. MapReduce technology allows to perform distributed calculations on a huge amount of data by dividing them into parts, performing parallel calculations of each of them and combining the results. In this paper, experimental studies were performed to compare CUDA and OpenGL frameworks performance measurements for MapReduce operations on heterogeneous cluster. It has been found that CUDA is a more suitable framework that provide a significant advantage in this regard. It is determined that the greater the amount of processing data, the greater the delay caused by OpenCL. Further research will be conducted to determine the energy consumption of both technologies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Dramatic demand for usage of big data processing technologies has been observed during last few
years. Simultaneously, general-purpose GPU accelerated computation frameworks develops. Both
techniques are being used in similar fields eg. data science [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], data processing [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], data mining [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
machine learning [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], solving various biological [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], medical [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], physical [10] and
geographical [11] problems. Consequently, many research studies have been made to compose
scalability of big data processing systems with maximization of resources usage and performance of
GPU powered computations. Those attempts aim improving performance [12], power-usage [13], deal
with low capacity of GPU memory [14] and improve programmability [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] of this approach
      </p>
      <p>Computation using graphic processors can speed up any type of calculations. The technique is called
GPGPU – general-purpose GPU [15]. It is possible to implement MapReduce’s reduce operation using
previously mentioned GPU computing frameworks in order to speed-up whole algorithm. GPGPU
enables to run a code on way more cores than standard CPU. Although GPU is not designed for general
usage and has significantly less memory that CPU [16], well designed algorithm can speed up
computations of many different kinds</p>
      <p>The main goal of this paper is to compare performance of basic algorithm which aggregates results
by sum. We test both CUDA and OpenCL™ implementations. CUDA and OpenCL are frameworks</p>
      <p>2021 Copyright for this paper by its authors.
which enable using GPU for non-graphic related computing [17]. They provide GPU acceleration for
mass parallel computation that can potentially speed-up MapReduce algorithms execution [14].</p>
      <p>To achieve our goal we use Apache™ HadoopⓇ [18] powered heterogeneous cluster. Algorithms
are implemented according to MapReduce paradigm in CUDA and OpenCL. Apache Hadoop is a
software framework for distributed computing dedicated for big data processing. It environment consist
of connected workstations forming a cluster.</p>
      <p>MapReduce is a paradigm for constructing algorithms dedicated to run on distributed environment
[19]. It is supported by Apache Hadoop framework. It allows to spread process of computing among
workstations connected into Hadoop cluster. It enables relatively fast processing of large amount of
data by making the process parallel. Hadoop implements scheduling algorithm for optimal nodes
resources usage and techniques for error recovery</p>
      <p>We measure performance by measuring time taken to execute widely used MapReduce algorithms
implemented with two leading GPU general purpose computation frameworks - CUDA and OpenGL.
CUDA [17] is a framework dedicated for NVIDIA graphics cards while OpenCL [20] is multi-platform
framework which computes using every fond CPU and GPU resource available on a host machine. This
means that what we measure is an overhead introduced by OpenCL in the specific environment of
Hadoop based heterogeneous cluster.</p>
      <p>
        Other approaches includes using other distributed computation frameworks such as MPI. However,
these are lesser fault-tolerant than Hadoop [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [21]. On the other hand, consequence of using Hadoop
and MapReduce paradigm forces redesigning data processing algorithm so as they fit into specific
schema. Some techniques of visual programming may be applied to overcome this issue [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], but it does
not ensure correctness of designed solution
      </p>
      <p>There are many different systems that strive in the field of some kind of GPU powered MapReduce
paradigm implementation. For example:
1. HAPI [22]</p>
      <p>Hadoop combined with Aparapi – Java-to-OpenCL conversion tool developed and released by
AMD. Proposes easy and ready to use API for designing and implementation of GPU MapReduce
algorithms. By hiding complexity of GPU programming by system of annotations, allows programmers
to focus on developing good algorithms.</p>
      <p>
        2. HadoopCL [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
      </p>
      <p>Extension to Apache Hadoop, HadoopCL combines Hadoop and OpenCL by usage of Aparapi.
Provides easy and flexible programming interface, guarantees reliability and low power consumption.
States that it achieves nearly 3x overall speedup and 55x of computational sections of example
MapReduce application algorithms.</p>
      <p>3. GPMR [23]</p>
      <p>GPMR is a stand-alone library for MapReduce that is supposed to use GPU clusters for large scale
computing. By modifying MapReduce to combine large amounts of map and reduce items into chunks
and partial reductions and accumulation, they better utilize power of GPU.</p>
      <p>4. MITHRA [24]</p>
      <p>An architecture that combines power of NVIDIA CUDA and Apache Hadoop to create scalable
performance gains by utilizing MapReduce programming model. MITHRA was designed especially for
executing computing tasks of massive and independent data.</p>
      <p>5. MARS [25]</p>
      <p>Mars is a MapReduce framework that is supposed to improve and ease programming complexity of
GPU programming by a familiar MapReduce interface.</p>
      <p>Authors of those systems compared their solutions with other works in terms of performance –
however, reliable comparison of most popular GPGPU computation framework, that would be free of
overhead ensuing from their framework usage, is still lacking</p>
    </sec>
    <sec id="sec-2">
      <title>2. Proposed architecture</title>
      <p>For the purpose of our research we have built a cluster containing 10 workstations, each of them
hosting HortonWorksⓇ Apache Hadoop framework implementation. Workstations uses two different
physical configuration – differing with presence of GPU. Configuration consist of: IntelⓇ XeonⓇ CPU
E5-1630 v3 3.70 GHz, 16 GB RAM and optional NVIDIA Quadro K4200.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Aggregation algorithm development</title>
      <p>Basic idea behind our project is speeding up reduce operation (see Figure 2) using algorithms
designed specifically for GPUs. As an example, we would use modified algorithm for calculation of
exponent.</p>
      <p>Exponent can be calculated using following formula [27]:
∞
(1)
  = ∑</p>
      <p>It is easy to notice that sum factors are independent, and as a result it is easy to calculate overall
value by just summing up middle results (in any order).</p>
      <p>When creating algorithm of fast sum calculation we inspired with NVIDIA publication [28] on
parallel computation.</p>
      <p>Basic idea behind it is following: having an array of numbers on length N, we can sum up pairs of
elements, so complexity reduces from O(N) to around O(log_2 (n)) (Figure 3).</p>
      <p>As single operation is around simultaneous, speed-up is quite impressive. In order to achieve that
with MapReduce in Hadoop we created modified implementations of map and reduce.</p>
      <p>Map output was (key, value) pair where key was always 1, and value was n-th factor of the sum. As
a result, after reduce operation we always got single key, and value which represents a result</p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>In this section, we describe methodology of our time performance analysis and comparison of
MapReduce solution on both CUDA and OpenCl.</p>
      <p>To measure a performance of the system we run several tests of exponent calculation, each of them
differing number of jobs with and without GPU for both frameworks separately. The first test was run
on only one GPU powered node. Then we have tested every number of nodes each kind starting from
one to 10 total amount of nodes. All results were presented on charts below alongside with its
discussion.</p>
      <p>We measure performance, by measuring runtime of tests on all used nodes from sending data to each
node, to receiving data from each node on master machine. Results often are shown as an improvement
factor. This should be considered as execution time in certain case divided by execution time for
nonGPU solution.</p>
      <p>All the tests were run 5 times and presented results are average value. It was necessary due to the
cache memory misses. Nevertheless, some small differences in similar measurements can be still
observed.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>On the Figure 4 linear increase of computation time depending on the number of calculated elements
can be observed. For our comparison the most significant information is difference between CUDA and
OpenCL improvement. The results shows that usage of OpenCL generates latency in availability of
computation results that linearly depends on the number of elements.</p>
      <p>That means difference in improvement factor of execution time (Figure 5) between those two
frameworks grows. These data were measured for very small – in terms of big data processing – number
of elements. Big data algorithms usually work with millions of items. That implies that simple
summation and memory move operations powered by OpenCL in big data analysis, when it comes to
practical usage, can lead to significant superiority of CUDA over OpenCL</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>We have presented comparison of CUDA and OpenGL frameworks performance measurements for
MapReduce operations on heterogeneous cluster. The results shows noticeable superiority of CUDA in
this issue. Specifically, the biggest amount of is processed, the bigger OpenCL caused latency occurred
to be. Nevertheless, it should be mentioned that CUDA is supported only by GPUs of the only one of
two leading graphical cards manufacturers – namely NVIDIA. Therefore, this research does not solve
the problem of the right framework choice for particular task on particular hardware resources available.
Moreover, the comparison can be also done in the field of energy-saving [13].</p>
    </sec>
    <sec id="sec-7">
      <title>7. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Grossman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Breternitz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Sarkar</surname>
          </string-name>
          , “
          <article-title>HadoopCL: MapReduce on Distributed Heterogeneous Platforms through Seamless Integration of Hadoop and OpenCL</article-title>
          ,” in
          <source>2013 IEEE International Symposium on Parallel Distributed Processing</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>1918</fpage>
          -
          <lpage>1927</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Aneta</given-names>
            <surname>Poniszewska-Maranda</surname>
          </string-name>
          , Daniel Kaczmarek, Natalia Kryvinska, Fatos Xhafa, “
          <article-title>Studying usability of AI in the IoT systems/paradigm through embedding NN techniques into mobile smart service system”</article-title>
          , Computing,
          <year>November 2019</year>
          , Volume
          <volume>101</volume>
          , Issue 11, pp.
          <fpage>1661</fpage>
          -
          <lpage>1685</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>W.</given-names>
            <surname>Ma</surname>
          </string-name>
          , G. Agrawal,
          <article-title>A Translation System for Enabling Data Mining Applications on GPUs</article-title>
          ,
          <source>in Proceedings of the International Conference on Supercomputing, NY, USA</source>
          ,
          <year>2009</year>
          , p.
          <fpage>400</fpage>
          -
          <lpage>409</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R.</given-names>
            <surname>Tkachenko</surname>
          </string-name>
          , I. Izonin,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kryvinska</surname>
          </string-name>
          , I. Dronyuk, K. Zub, “
          <article-title>An Approach towards Increasing Prediction Accuracy for the Recovery of Missing IoT Data based on the GRNN-SGTM Ensemble”</article-title>
          ,
          <source>MDPI Sensors</source>
          <year>2020</year>
          ,
          <volume>20</volume>
          (
          <issue>9</issue>
          ), 2625, https://doi.org/10.3390/s20092625.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Grossman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Kaeli</surname>
          </string-name>
          , “
          <article-title>Mahout on Heterogeneous Clusters Using HadoopCL,”</article-title>
          <source>in Proceedings of the 2Nd Workshop on Parallel Programming for Analytics Applications</source>
          , New York, NY, USA,
          <year>2015</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yin</surname>
          </string-name>
          , et. al, “
          <article-title>Computing Platforms for Big Biological Data Analytics: Perspectives and Challenges,”</article-title>
          <string-name>
            <surname>Comput. Struct. Biotechnol. J.</surname>
          </string-name>
          , vol.
          <volume>15</volume>
          , no. Supplement C, pp.
          <fpage>403</fpage>
          -
          <lpage>411</lpage>
          , Jan.
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Khare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khare</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Khan</surname>
          </string-name>
          , “
          <article-title>HCudaBLAST: an implementation of BLAST on Hadoop</article-title>
          and Cuda,”
          <source>J. Big Data</source>
          , vol.
          <volume>4</volume>
          , no.
          <issue>1</issue>
          , p.
          <fpage>41</fpage>
          ,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Serrano</surname>
          </string-name>
          , et. al.,
          <article-title>“Architecture for the Execution of Tasks in Apache Spark in Heterogeneous Environments</article-title>
          ,” in Euro-Par
          <source>2016: Parallel Processing Workshops</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>504</fpage>
          -
          <lpage>515</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>O.</given-names>
            <surname>Tymchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Havrysh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Khamula</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Kovalskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vasiuta</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Lyakh</surname>
          </string-name>
          , '
          <article-title>Methods of Converting Weight Sequences in Digital Subtraction Filtration'</article-title>
          , in 2019 IEEE 14th International
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>