<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Experimental Considerations Towards E ective Memory Bandwidth Evaluation on Large-Scale ccNUMA Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pavel Drobintsev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vsevolod Kotlyarov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aleksei Levchenko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evgeniy Petukhov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Peter the Great St. Petersburg Polytechnic University</institution>
          ,
          <addr-line>Saint Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>40</fpage>
      <lpage>49</lpage>
      <abstract>
        <p>In order to predict the performance of a wide range of scienti c applications on current high-end ccNUMA architectures, this paper introduces benchmark-related modeling considerations for memory bandwidth and hybrid MPI/OpenMP performance. We use HPCG, state-ofthe-art benchmark, in order to create a workload representative for a multitude of computational and communication tasks. We ran our model validation experiments on real ccNUMA machine with 12Tb RAM in single operating system image mode to de ne the boundaries of problem size and demonstrate improved indicators for the target architecture as compared with the fundamental model. Our model will permit to evaluate reliably the performance of contemporary and future ccNUMA systems with more than 20Tb of RAM and to compare their experimental results with other problem-oriented architectures worldwide.</p>
      </abstract>
      <kwd-group>
        <kwd>benchmarking</kwd>
        <kwd>ccNUMA</kwd>
        <kwd>NUMA e ects</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        HPCG
memory bandwidth
The current Cache-Coherent Non-Uniform Memory Access (ccNUMA) systems
are able to provide a larger amount of random access memory per node with a
single operating system image than it is accessible on an usual cluster.
Asymmetric ccNUMA nature raises a number of potentially overwhelming strong NUMA
e ects such as memory hot-spotting, the substantial penalty of incorrect NUMA
assignment, varying complex multilevel structure of latency and mismatch of
data access models and actual distribution of data in memory [
        <xref ref-type="bibr" rid="ref22 ref4 ref9">9, 4, 22</xref>
        ]. These
factors have a multidirectional impact on memory bandwidth, which continues
to be a major system challenge for memory-bound scienti c applications.
Deducing memory bandwidth from the theoretical peak one for a speci c computing
procedure is a sophisticated problem [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Thereby, hypothetical prediction of
ccNUMA systems memory bandwidth is unconvincing.
      </p>
      <p>
        Our ultimate goal is to measure reliably the performance of current and
future ccNUMA systems. In this work, we present only preliminary considerations
for the experimental benchmarking, modeling and predicting of ccNUMA
memory bandwidth. The High Performance Conjugate Gradients (HPCG)
Benchmark was used for creating a workload with the low ratio of computations to
data access that is representative for the major communication and
computational patterns [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. As we extend the existing HPCG performance model, we
predict the e ective memory bandwidth of real system with a globally
addressable memory, so-called jumbonode equipped with 12Tb of RAM and loaded as
a single operating system image. We shall compare the obtained results with
other problem-oriented architecture worldwide and predict the e ective memory
bandwidth of future ccNUMA machines. We also demonstrate valuable technical
ccNUMA-related aspects of launching a hybrid HPCG.
      </p>
      <p>The remaining sections of this paper are organized as follows. In Section
2 we shall mention the most important previous works including the reference
model. In Section 3 we shall describe the factors considered by us that are able
to extend the existing general-purpose model for the ccNUMA architecture. The
model validation and experimental results are discussed in Section 4. Finally, we
summarize our conclusions in Section 5, where we also consider the aspects of
future development of the model.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background and Related Work</title>
      <p>We review the previous work in NUMA- and HPCG-related aspects, which will
help us to take into account more challenges proposed by the ccNUMA
architecture, namely (1) hybrid MPI/OpenMP performance modeling, (2) NUMA
e ects, which have impact on performance and (3) HPCG-related publications
including reference model of HPCG performance.</p>
      <p>
        Wang et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] presents a model, which predicts both memory bandwidth
usage and optimal core allocations. Luo et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] provides valuable insights
into o -socket and inter-socket bandwidth modeling to analyze performance of
di erent thread and data placements. A hybrid approach for the development
of high-level performance models of large-scale computing systems, which
combines mathematical modeling and discrete-event simulation has presented in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
Work [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] shows advantages of hybrid OpenMP/MPI programming on
largescale NUMA clusters. Other work on performance modeling of communication
and computation in hybrid MPI/OpenMP applications is carried out by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        As for the HPCG, we already have a number of important works since 2013.
Dongarra et. al [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] describes allowed and disallowed HPCG optimizations. Several
studies, [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref2 ref24 ref3 ref5">24, 3, 11, 10, 12, 5, 2</xref>
        ] have been done to describe an early experience of
HPCG optimizations on large systems like Tianhe-2, Angara, Sunway TaihuLight
System, etc.
      </p>
      <p>
        A general-purpose performance model [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] of the HPCG Benchmark
includes the execution time for main kernels, namely for Symmetric Gauss-Seidel
smoother (SymGS), Sparse Matrix Vector Multiplication (SpMV), Vector
Update, Global Dot Product (DDOT), as well as Multigrid preconditioner (MG).
Together with the model of two communication procedures, the complete model
allows us to predict HPCG performance reliably. As implied by the foregoing,
HPCG can provide insight into comparsion of ccNUMAs with the results of
other problem oriented architectures (non-ccNUMA). The evolutionary aspects
and experimental application of the mentioned works are contributions of this
work.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>The Extended Model Features</title>
      <p>
        The contribution made by our work is the prediction of ccNUMA system memory
bandwidth by using an reference model from work [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The main performance
challenges on ccNUMA are (1) locality of data access, (2) the amount of data
sharing between threads and (3) e ective memory bandwidth [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The e ective
memory bandwidth from main memory participating in the model of all
computing procedures is of a greater signi cance. Our contribution is also in using
hybrid HPCG, not only pure MPI like in model [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. In spite of the facts that
HPCG is well balanced at the MPI level, the performance of pure MPI
realization is higher and OpenMP does not provide support for ccNUMA, our core
point is that the hybrid version is an additional great challenge for ccNUMA
architectures per se, providing emergence of a number of e ects detrimental to
performance, such as memory hot-spotting. Table 1 shows the estimated range
of model options that have been considered by us or have such prospect. The
features of our model include (1) the execution time in seconds for main kernels
(SYMGS, SpMV, etc.) previously presented in [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and extended in this work to
take into account the e ective memory bandwidth and interconnect latency, and
(2) the e ects of hybrid MPI+OpenMP parallelism in ccNUMA environment.
In this paper, we describe only the experimental aspects of e ective bandwidth
evaluation.
      </p>
      <p>Extended
SY M GSexec time(sec) Considered +BWeff
SpM Vexec time(sec) Considered +BWeff
W AXP Bexec time(sec) Considered +BWeff
DDOTexec time(sec) Considered +BWeff
Allreduce; Haloexec time(sec) Considered +IClatency
Hybrid MPI+OpenMP Not considered Considered
E ective bandwidth Not considered Considered
IC latency Not considered Considered</p>
      <p>Optimization techniques Not considered Future work</p>
      <p>We already know total execution time from the non-hybrid HPCG model
Itertime(sec) = M G + SpM V (depth = 0) + 3(DDOT + W AXP B)
(1)</p>
      <p>
        Hybrid HPCG is more memory-bound, than pure MPI and can deliver better
performance [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], especially in case of the ccNUMA. For OpenMP, execution time
proposed by Wu and Taylor for hybrid MPI/OpenMP scienti c applications [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
is rewritten as follows:
      </p>
      <p>P erf = (RefMP I + OM P )</p>
      <sec id="sec-3-1">
        <title>T otalexec time(sec)</title>
      </sec>
      <sec id="sec-3-2">
        <title>Compexec time(sec) + Commexec time(sec)</title>
        <p>where OMP represent the model for intranode OpenMP performance:
OM P = Tc1 + (BWn
1) Tc2</p>
        <p>BW2</p>
        <p>
          Here we use Eqn. 3 to model the OpenMP application execution time on
n cores based on the performance for single and dual cores (Tc) and memory
bandwidth ratio (BWn) [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
        <p>
          The e ective memory bandwidth can be deduced from reference model [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
for every HPCG kernel as follows.
        </p>
        <p>BWSY MGS (Bytes=sec) =
(nx
ny
nz)=23 d
(20 + 20</p>
        <p>27)(Bytes)</p>
      </sec>
      <sec id="sec-3-3">
        <title>SY M GSexec time(sec)</title>
        <p>BWSpMV (Bytes=sec) =
(nx
ny
nz)=23 d
(20 + 20
27)(Bytes)</p>
      </sec>
      <sec id="sec-3-4">
        <title>SpM Vexec time(sec)</title>
        <p>BWW AXP B (Bytes=sec) =
BWDDOT (Bytes=sec) =
(nx
ny
nz)=23 d
24(Bytes)</p>
      </sec>
      <sec id="sec-3-5">
        <title>W AXP Bexec time(sec)</title>
        <p>(nx
ny
nz)=23 d
16(Bytes)</p>
      </sec>
      <sec id="sec-3-6">
        <title>DDOTexec time(sec)</title>
        <p>
          where the most expensive routine is SYMGS [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>
          While computing procedures were modeled exhaustively, important factors
obtained empirically remain. Second of them, after e ective memory bandwidth
from main memory, is IC latency, whose in uence on the prediction is considered
as insigni cant by the authors of work [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. We evaluate empirically IC latency
by KNEM, a Linux kernel module enabling high-performance intra-node MPI
communication for large messages [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
(2)
(3)
2
(4)
(5)
(6)
(7)
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Results and Discussion</title>
      <p>Since ccNUMA having more than 3Tb memory size are an exotic systems and
it seems complex to obtain a set of various gigantic ccNUMA systems, we use
our target system in di erent con gurations presented in Table 2.</p>
      <p>For a more in-depth study of NUMA-related challenges, we performed our
early-stage experiments with hybrid HPCG running on macronodes from 188Gb
of RAM (48 cores) with aggregation of macronode memory to 3Tb of RAM
(768 cores) and with subsequent integration into a single macronode with up to
12Tb of RAM (3072 cores) at the nal stage.</p>
      <p>A standalone server is based on AMD Opteron Processor 6380, interconnect
has a 3D Torus topology. We use Linux 4.12 with patchset for support of Block
Transfer Engine driver for NumaChip node controllers, which provide large
number of outstanding memory transactions, memory controller for the cache and
memory tags, a cross-bar switch for the interconnect fabric and a number of
interconnect fabric link controllers.</p>
      <p>
        Hybrid HPCG run on ccNUMA system with 12Tb is in itself nontrivial
problem, which has not been previously described, to the best of our knowledge.
Operating system as well as HPCG have been compiled with optimized libgomp,
which supports stack and thread local storage (TLS) to keep local to more than
1024 threads. A private stack with size up to 2Gb is allocated to each HPCG
thread for increasing of problem size, which is very relevant. All MPI processes
mapped by NUMA nodes to reduce memory tra c and keep the data close to the
cores [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Generation of instructions to prefetch memory is used for increasing
performance of loops that access large arrays. Load is balanced for improving
e ciency of OpenMP application, distributing threads through all accessible
NUMA nodes, using more FPU and reducing load on the memory interface and
L3 cache. Generation of instructions to prefetch memory is used for increasing
performance of loops that access large arrays. The largest allowable size of the
problem was 256 256 256. All start-up options described above have a
signi cant impact on HPCG performance on ccNUMA. Figure 1 shows the results
of modeling with the help of the fundamental model that does not take these
characteristics into account.
      </p>
      <p>
        Figure 2 compares our predictions with the actual measured results of hybrid
HPCG on the jumbonode with 12Tb of RAM, and the predictions by the
reference model have been put to comparison too. In contrast to the results of work
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the hybrid HPCG scales non-linearly on ccNUMA system; non-uniformity
of the system results in surface separation whose causes will be studied. And
nally in Figure 3 we show the comparison of the modeled ccNUMA bandwidth
and the STREAM Benchmark results.
      </p>
      <p>
        As to predicting the performance of future ccNUMA systems with more than
20Tb memory size, the view taken is that HPCG will remain memory-bound in
the future as well. Having about 7Tb memory consumption upon HPCG start
with the maximum jumbonode task size, we expect a proportionally high
memory consumption since future ccNUMAs will have at least 4Gb per core. IC
latency, whose weight in the general HPCG model is insigni cant, will grow. Based
on our model, we expect the performance of at least 400GFlops for macronode
with 20Tb of RAM. In respect of current non-ccNUMA machines, HPCG o ers
a single metric for comparing various problem-oriented architectures and reduces
the gap between them created by LINPACK. E.g., the experimental ccNUMA
demonstrates a satisfactory HPCG performance as compared to the results of
technical report [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for \The Sunway TaihuLight supercomputer" [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], suggesting
that the ccNUMA memory is slightly slower as compared to the current TOP500
leaders.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Concluding Remarks and Future Work</title>
      <p>In this work, we presented an experimental approach to contemporary ccNUMA
systems memory bandwidth evaluation. HPCG Benchmark was used to create a
workload comparable to the contemporary scienti c applications. The existing
HPCG performance model was extended by considering hybrid MPI/OpenMP
and supplemented by the factors in uencing memory bandwidth. As a result,
the e ective memory bandwidth of an real ccNUMA system with 12Tb of RAM
was predicted. The approach used by us can be applied to comparing of current
and future ccNUMA machines.</p>
      <p>
        As implied by the foregoing, the divergence between the actually obtained
using STREAM Benchmark results and the deduced from reference model ones
is up to 12%. As was demonstrated in Section 4, the whole software environment
was optimized on a wide scale, namely the Linux kernel, gcc, libgomp, etc.
However, large-scale optimizations of the HPCG itself are still possible. In the near
future we plan to concentrate for realization of the existing HPCG
optimizations for ccNUMA case as \improving the performance of HPCG will improve
the performance of real applications" (J.Dongarra, et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
      </p>
      <p>
        First of all, we consider the re nement of the cache locality model with
the help of the novel HPCG optimization technique proposed in the paper [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
namely coloring along two areas XY at a time in SymGS. Among other
improvements a number of works argue to replace the default CSR matrix storage
format with simpli ed SELLPACK for SpMV and SYMGS kernels [
        <xref ref-type="bibr" rid="ref2 ref24">24, 2</xref>
        ]. Table
3 shows the expected speedup. Also recent work [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] demonstrates new data
redeployment model which allows to reduce the remote memory access overhead
for computation-intensive applications with large size of problem in ccNUMA
architecture. These optimizations presume an analysis that will allow to study
better the challenges proposed by the ccNUMA architecture. Finally, we plan to
propose an IC latency model for ccNUMA systems in the near future.
Acknowledgments. This work was nancially supported by the Ministry of
Education and Science of the Russian Federation within the framework of the
Federal Targeted Programme for Research and Development in Priority Areas of
Advancement of the Russian Scienti c and Technological Complex for 2014-2020
(№ 14.584.21.0022, ID RFMEFI58417X0022). The results were obtained using
ccNUMA system in Supercomputer Center of Peter the Great St. Petersburg
Polytechnic University.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adhianto</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Performance modeling of communication and computation in hybrid mpi and openmp applications</article-title>
          .
          <source>In: 12th International Conference on Parallel and Distributed Systems - (ICPADS'06)</source>
          . vol.
          <volume>2</volume>
          , pp.
          <volume>6</volume>
          pp.
          <source>{</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Agarkov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semenov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simonov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Optimized implementation of HPCG benchmark on supercomputer with \Angara" interconnect</article-title>
          . In: Voevodin,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Sobolev</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the 1st Russian Conference on Supercomputing | Supercomputing Days 2015. CEUR Workshop Proceedings</source>
          , vol. Vol-
          <volume>1482</volume>
          , pp.
          <volume>294</volume>
          {
          <fpage>302</fpage>
          . Research Computing Center, Moscow State University, CEUR-WS.org,
          <source>Moscow (Sep</source>
          <volume>28</volume>
          {
          <fpage>29</fpage>
          ,
          <year>2015</year>
          ), http://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>1482</volume>
          /294.pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : HPCG:
          <article-title>Preliminary evaluation and optimization on Tianhe-2 CPU-only nodes</article-title>
          .
          <source>In: Computer Architecture and High Performance Computing (SBAC-PAD)</source>
          ,
          <year>2014</year>
          IEEE 26th International Symposium on. pp.
          <volume>41</volume>
          {
          <issue>48</issue>
          (Oct
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Diener</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cruz</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navaux</surname>
            ,
            <given-names>P.O.</given-names>
          </string-name>
          :
          <article-title>Modeling memory access behavior for data mapping</article-title>
          .
          <source>International Journal of High Performance Computing Applications</source>
          (
          <year>2016</year>
          ), http://hpc.sagepub.com/content/early/2016/04/13/1094342016640056.abstract
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dongarra</surname>
          </string-name>
          ,
          <source>J.: Report on the Sunway Taihulight System. Tech. Rep. UT-EECS16-742</source>
          , Oak Ridge National Laboratory,Department of Electrical Engineering and Computer Science, University of Tennessee (Jun
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dongarra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heroux</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luszczek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>High-performance conjugate-gradient benchmark: A new metric for ranking high-performance computing systems</article-title>
          .
          <source>International Journal of High Performance Computing Applications</source>
          <volume>30</volume>
          (
          <issue>1</issue>
          ),
          <volume>3</volume>
          {
          <fpage>10</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xue</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qiao</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hou</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , Zhang, J.,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>The Sunway TaihuLight supercomputer: system and applications</article-title>
          .
          <source>Science China Information Sciences</source>
          <volume>59</volume>
          (
          <issue>7</issue>
          ),
          <volume>1</volume>
          {
          <fpage>16</fpage>
          (
          <year>2016</year>
          ), http://dx.doi.org/10.1007/s11432-016-5588-7
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Goglin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreaud</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Knem: A generic and scalable kernel-assisted intra-node mpi communication framework</article-title>
          .
          <source>J. Parallel Distrib. Comput</source>
          .
          <volume>73</volume>
          (
          <issue>2</issue>
          ),
          <volume>176</volume>
          {188 (Feb
          <year>2013</year>
          ), http://dx.doi.org/10.1016/j.jpdc.
          <year>2012</year>
          .
          <volume>09</volume>
          .016
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robertazzi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Characterization of input/output bandwidth performance models in NUMA architecture for data intensive applications</article-title>
          .
          <source>In: 2013 42nd International Conference on Parallel Processing</source>
          . pp.
          <volume>369</volume>
          {
          <issue>378</issue>
          (Oct
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Reducing communication overhead in the high performance conjugate gradient benchmark on Tianhe-2</article-title>
          . In: Distributed Computing and Applications to Business, Engineering and Science (DCABES),
          <year>2014</year>
          13th International Symposium on. pp.
          <volume>13</volume>
          {
          <issue>18</issue>
          (Nov
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Accelerating HPCG on Tianhe-2: A hybrid CPU-MIC algorithm</article-title>
          .
          <source>In: 2014 20th IEEE International Conference on Parallel and Distributed Systems (ICPADS)</source>
          . pp.
          <volume>542</volume>
          {
          <issue>551</issue>
          (Dec
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <article-title>: 623 T op/s HPCG run on Tianhe-2: Leveraging millions of hybrid cores</article-title>
          .
          <source>International Journal of High Performance Computing Applications</source>
          <volume>30</volume>
          (
          <issue>1</issue>
          ),
          <volume>39</volume>
          {
          <fpage>54</fpage>
          (
          <year>2016</year>
          ), http://hpc.sagepub.com/content/30/1/39.abstract
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Compositional model of coherence and NUMA e ects for optimizing thread and data placement</article-title>
          .
          <source>In: 2016 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)</source>
          . pp.
          <volume>151</volume>
          {
          <issue>152</issue>
          (April
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Marjanovic</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gracia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glass</surname>
            ,
            <given-names>C.W.</given-names>
          </string-name>
          :
          <article-title>Performance modeling of the HPCG benchmark</article-title>
          . In: Jarvis,
          <string-name>
            <given-names>A.S.</given-names>
            ,
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.S.</given-names>
            ,
            <surname>Hammond</surname>
          </string-name>
          , D.S. (eds.)
          <source>High Performance Computing Systems. Performance Modeling</source>
          , Benchmarking, and Simulation: 5th International Workshop, PMBS 2014,
          <article-title>New Orleans</article-title>
          , LA, USA, November
          <volume>16</volume>
          ,
          <year>2014</year>
          . Revised Selected Papers. pp.
          <volume>172</volume>
          {
          <fpage>192</fpage>
          . Springer International Publishing,
          <source>Cham (Dec</source>
          <volume>5</volume>
          {
          <fpage>9</fpage>
          ,
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Nakajima</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Flat MPI vs</article-title>
          .
          <article-title>Hybrid: Evaluation of parallel programming models for preconditioned iterative solvers on \T2K Open Supercomputer"</article-title>
          .
          <source>In: 2009 International Conference on Parallel Processing Workshops</source>
          . pp.
          <volume>73</volume>
          {
          <issue>80</issue>
          (Sept
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smelyanskiy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vaidyanathan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heinecke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalamkar</surname>
            ,
            <given-names>D.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patwary</surname>
            ,
            <given-names>M.M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubey</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>E cient shared-memory implementation of high-performance conjugate gradient benchmark and its application to unstructured matrices</article-title>
          . In: SC14:
          <article-title>International Conference for High Performance Computing, Networking, Storage and Analysis</article-title>
          . pp.
          <volume>945</volume>
          {
          <issue>955</issue>
          (Nov
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Pllana</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benkner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xhafa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barolli</surname>
            ,
            <given-names>L.:</given-names>
          </string-name>
          <article-title>Hybrid performance modeling and prediction of large-scale computing systems</article-title>
          .
          <source>In: Complex, Intelligent and Software Intensive Systems</source>
          ,
          <year>2008</year>
          . CISIS 2008. International Conference on. pp.
          <volume>132</volume>
          {
          <issue>138</issue>
          (March
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Tsuji</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sato</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Performance evaluation of OpenMP and MPI hybrid programs on a large scale multi-core multi-socket cluster, T2K open supercomputer</article-title>
          .
          <source>In: 2009 International Conference on Parallel Processing Workshops</source>
          . pp.
          <volume>206</volume>
          {
          <issue>213</issue>
          (Sept
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>So</surname>
            <given-names>a</given-names>
          </string-name>
          , M.L.:
          <article-title>Predicting the memory bandwidth and optimal core allocations for multi-threaded applications on large-scale NUMA machines</article-title>
          .
          <source>In: 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA)</source>
          . pp.
          <volume>419</volume>
          {
          <issue>431</issue>
          (March
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , V.:
          <article-title>Performance modeling of hybrid MPI/OpenMP scienti c applications on large-scale multicore cluster systems</article-title>
          .
          <source>In: Computational Science and Engineering (CSE)</source>
          ,
          <year>2011</year>
          IEEE 14th International Conference on. pp.
          <volume>181</volume>
          {
          <issue>190</issue>
          (Aug
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antony</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rendell</surname>
            ,
            <given-names>A.P.:</given-names>
          </string-name>
          <article-title>A simple performance model for multithreaded applications executing on non-uniform memory access computers</article-title>
          .
          <source>In: High Performance Computing and Communications</source>
          ,
          <year>2009</year>
          . HPCC '
          <volume>09</volume>
          . 11th IEEE International Conference on. pp.
          <volume>79</volume>
          {
          <issue>86</issue>
          (
          <year>June 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jin</surname>
          </string-name>
          , H.:
          <article-title>A Data-Centric Tool to Improve the Performance of Multithreaded Program on NUMA</article-title>
          , pp.
          <volume>74</volume>
          {
          <fpage>87</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ren</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Optimization of computation-intensive applications in cc-NUMA architecture</article-title>
          .
          <source>In: 2016 International Conference on Networking and Network Applications (NaNA)</source>
          . pp.
          <volume>244</volume>
          {
          <issue>249</issue>
          (
          <year>July 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Optimizing and scaling HPCG on Tianhe-2: Early experience</article-title>
          . In: Sun,
          <string-name>
            <given-names>X.h.</given-names>
            ,
            <surname>Qu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Stojmenovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Guo</surname>
          </string-name>
          , Huaand Min,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (eds.)
          <article-title>Algorithms and Architectures for Parallel Processing</article-title>
          : 14th International Conference,
          <year>ICA3PP 2014</year>
          , Dalian, China,
          <source>August 24-27</source>
          ,
          <year>2014</year>
          . Proceedings, Part I. pp.
          <volume>28</volume>
          {
          <fpage>41</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>