<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>PARALLEL ALGORITHMS FOR STUDYING THE SYSTEM OF LONG JOSEPHSON JUNCTIONS</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>M. Bashashin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Nechaevskiy</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Podgainy</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>I. Rahmonov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yu. Shukrinov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>O. Streltsova</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. Zemlyanaya</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Zuev</string-name>
          <email>zuevmax@jinr.ru</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bogoliubov Laboratory of Theoretical Physics</institution>
          ,
          <addr-line>JINR, 6 Joliot-Curie St., Dubna, 141980</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dubna State University</institution>
          ,
          <addr-line>19 Universitetskaya St., Dubna, 141980</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Laboratory of Information Technologies, JINR</institution>
          ,
          <addr-line>6 Joliot-Curie St., Dubna, 141980</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Maxim Bashashin</institution>
          ,
          <addr-line>Andrey Nechaevskiy, Dmitry Podgainy, Ilhom Rahmonov, Yury Shukrinov, Oksana Streltsova, Elena Zemlyanaya, Maxim Zuev</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>4</volume>
      <issue>2019</issue>
      <fpage>392</fpage>
      <lpage>396</lpage>
      <abstract>
        <p>The results on studying the efficiency of parallel implementations of the computing scheme for calculating the current-voltage characteristics of the system of long Josephson junctions are presented in the paper [1-4]. The following parallel implementations were developed: the OpenMP implementation for computing on systems with shared memory, the CUDA implementation for computing on Nvidia graphics processors. The development, debugging and profiling of parallel applications were performed on the education and testing polygon of the HybriLIT heterogeneous computing platform, while computations were carried out on the “Govorun” supercomputer [2].</p>
      </abstract>
      <kwd-group>
        <kwd>long Josephson junctions</kwd>
        <kwd>parallel computations</kwd>
        <kwd>HPC</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Formulation of the problem</title>
      <p>A generalized model that takes into account the inductive and capacitive coupling between
long Josephson junctions (LJJs) is considered [1]. The system of N coupled LJJs is supposed to consist
of superconducting (S) and intermediate dielectric (I) layers with a length L (Fig. 1).
(l  1, 2,N ) . In a dimensionless form, the system of equations has the form [1]:
</p>
      <p> C V ,
 t
V  1 2
 t x2  V  I t ;
1  V1 
   
  2 , V  V2 , 0  x  L, t  0,
   
 N  VN 
where  is the inductive coupling matrix, С is the capacitive coupling matrix:
 1 S 0 0 S   Dc sc 0
  
   0 S 1 S 0 , C   0 sc Dc sc
  
 S 0 0 S 1   sc 0
0
0
sc 


,

</p>
      <p>Dc 
layer; sc is the capacitive coupling parameter, I t  is the external current.</p>
      <p>The system of equations is supplemented with zero initial and boundary conditions:
l 0,t  l  L,t 
Vl 0,t   Vl  L,t  

 0, l  1, 2</p>
      <p>
        x x
The problem when boundary conditions in the direction x were defined by the external magnetic field
was also considered [
        <xref ref-type="bibr" rid="ref2">3</xref>
        ].
      </p>
      <p>When calculating the current-voltage characteristics (CVC), the dependence of the current on
time was selected in the form of steps (a schematic representation is given in Figure 2), i.e. the
problem is solved at the constant current  I  I j  , while the found functions l , Vl l  1, 2
N  are
taken as initial conditions to solve the problem for the current I j1 .</p>
    </sec>
    <sec id="sec-2">
      <title>2. Computing scheme</title>
      <p>values l , Vl l  1, 2
method.</p>
      <p>A uniform grid by the spatial variable (the number of grid nodes NX) was built for the
numerical solution of the initial-boundary problem. In the system of equations (1), the second-order
derivative by the coordinate x is approximated using three-point finite-difference formulas on the
discrete grid with the uniform step x . The obtained system of differential equations relative to the</p>
      <p>N  in nodes of the discrete grid by x is solved by the fourth-order Runge-Kutta</p>
      <p>To calculate CVC, averaging Vl  x,t  over the coordinate and time is performed. To do this, at
each time step, the integration of voltage over the coordinate using the Simpson method and the
averaging are carried out</p>
    </sec>
    <sec id="sec-3">
      <title>3. Parallel scheme</title>
      <p>When numerically solving the initial-boundary problem by the fourth-order Runge-Kutta
method over the time variable, at each time layer the Runge-Kutta coefficients  Ki  can be found
independently (in parallel) for all NX nodes of the spatial grid and for all N Josephson junctions.
Meanwhile, the coefficients Ki i  1, 2,3, 4 are defined one by one (sequentially). Thus, the
parallelization is efficiently performed on the NX  N points. When carrying out averaging in CVC
computing, the calculation of integrals can be performed in parallel as well.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Parallel implementations</title>
      <p>To speed up CVC computing, parallel implementations of the computing scheme described
above were developed. The results on studying the efficiency of parallel implementations performed at
the values of the following parameters: L  10 , Imin  0 , Imax 1.1,   0.2 , N  1, Tmax  200 ,
t  0.04 – are presented below; the number of nodes by the spatial variable is NX  20048 .</p>
      <sec id="sec-4-1">
        <title>4.1 OpenMP implementation</title>
        <p>The computations were performed:
 on computing nodes with processors Intel Xeon Phi 7290 (KNL: 16GB, 1.50 GHz, 72 cores, 4
threads per core supported – total 288 logical cores) and the Intel compiler (Intel Cluster Studio
18.0.1 20171018);
 on dual-processor computing nodes with processors Intel Xeon E5-2695 (Broadwell; 45 MB
Cache, 2.1 GHz, 18 cores, 2 threads per core supported – total 72 logical cores per node);
 on dual-processor computing nodes with processors Intel Xeon Gold 6154 (Skylake; 24.75 MB</p>
        <p>Cache, 3.00 GHz, 18 cores, 2 threads per core supported – total 72 logical cores per node);
 on dual-processor computing nodes with processors Intel Xeon Platinum 8268 (Cascade Lake;
35.75 MB Cache, 2.9 GHz, 24 cores, 2 threads per core supported – total 96 logical cores per
node).</p>
        <p>The graphs of the dependence of the calculation speedup obtained using the parallel algorithm:
T
S  1 ,</p>
        <p>Tn
(where T1 is the computation time using one core, Tn is the time of computations on n-logical cores)
on the number of threads, the number of which is equal to the number of logical cores, and the graph
of the dependence of efficiency of using computing cores by the parallel algorithm:</p>
        <p>T
E  1 100%,</p>
        <p>nTn
characterizing the scalability of the parallel algorithm, are presented below.</p>
        <p>Figure 3 shows the dependence of speedup on the number of threads when performing
computations on nodes with KNL without instructions AVX-512 and using instructions AVX-512,
while Figure 4 illustrates the dependence of efficiency of these computations. It is noteworthy that the
use of this instruction allowed us to reduce the computation time in 1.8 times.</p>
        <p>The computation time on CPU Intel Xeon E5-2695, Intel Xeon Gold 6154 and Intel Xeon
Platinum 8268 is presented in Figure 5.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2 CUDA implementation</title>
        <p>A CUDA implementation of the parallel algorithm was developed for computing on Nvidia
graphics accelerators. The parallel reduction algorithm using shared memory was applied for
calculating integrals. The computation time on the graphics accelerators Nvidia Tesla K40 and Nvidia
Tesla K80 is presented in Figure 6.</p>
        <p>The study was supported by the Russian Science Foundation (the project № 18-71-10095).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Comparative analysis of parallel implementations</title>
      <p>For the above parameter values to calculate CVC of LJJs, the best computation time on nodes
with the KNL processor is 2108.911 minutes on 150 OpenMP threads with instructions AVX-512;
using Nvidia K80 instead of Nvidia K40 reduced the computation time by 1.08 times or the above
parameter values for calculating CVC of LJJs.</p>
      <p>When comparing the processors Intel Xeon E5-2695, Intel Xeon Gold 6154 and Intel Xeon
Platinum 8268, the minimal computation time is 83.23 seconds on Intel Xeon Platinum 8268, and the
speedup of computing reached 2.15 times in comparison with Intel Xeon E5-2695.
[1] Atanasova P.H., Bashashin M.V., Rahmonov I.R., Shukrinov Yu.M., Zemlyanaya E.V. Influence
of the inductive and capacitive coupling on the current-voltage characteristic and electromagnetic
radiation of the system of long Josephson junctions // Journal of Experimental and Theoretical
Physics. 2017. Vol. 151. No 1. Pp. 151-159 (in Russian).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Adam</given-names>
            <surname>Gh</surname>
          </string-name>
          .,
          <string-name>
            <surname>Bashashin</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belyakov</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirakosyan</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matveev</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Podgainy</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sapozhnikova</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Streltsova</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torosyan</surname>
            <given-names>Sh.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vala</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valova</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vorontsov</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaikina</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zemlyanaya</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuev</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>IT-ecosystem of the HybriLIT heterogeneous platform for high-performance computing and training</article-title>
          of IT-specialists // CEUR Workshop Proceedings.
          <year>2018</year>
          . Vol.
          <volume>2267</volume>
          . Pp.
          <volume>638</volume>
          -
          <fpage>644</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Bashashin</surname>
            <given-names>M.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zemlyanaya</surname>
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahmonov</surname>
            <given-names>I.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shukrinov</surname>
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Atanasova</surname>
            <given-names>P.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volokhova</surname>
            <given-names>A.V.</given-names>
          </string-name>
          <article-title>Numerical approach and parallel implementation for computer simulation of stacked long Josephson Junctions /</article-title>
          / Computer Research and Modeling.
          <year>2016</year>
          . Vol.
          <volume>8</volume>
          . No.
          <issue>4</issue>
          . Pp.
          <volume>593</volume>
          -
          <fpage>604</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Zemlyanaya</surname>
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bashashin</surname>
            <given-names>M.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahmonov</surname>
            <given-names>I.R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Shukrinov</given-names>
            <surname>Yu</surname>
          </string-name>
          .M.,
          <string-name>
            <surname>Atanasova</surname>
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Kh</surname>
          </string-name>
          .,
          <string-name>
            <surname>Volokhova</surname>
            <given-names>A. V.</given-names>
          </string-name>
          <article-title>Model of stacked long Josephson junctions: Parallel algorithm and numerical results in case of weak coupling /</article-title>
          / AIP Conference Proceedings.
          <year>2016</year>
          . Vol
          <volume>1773</volume>
          .
          <fpage>110018</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>