<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>THE DEPENDENCE OF SITES SPEED FACTOR FROM THE DECLARED HS06</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>V. A. Matskovskaya</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Sciabà</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Geneva</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Switzerland</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Plekhanov Russian University of Economics</institution>
          ,
          <addr-line>36 Stremyanny per., Moscow, 117997</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>95</fpage>
      <lpage>98</lpage>
      <abstract>
        <p>The work started with the existing data analysis for the ATLAS experiment, designed to measure the processing speed of various ATLAS sites. The main task was to adapt the analysis so that it was fully automated and could be integrated into the ATLAS monitoring system. Another goal was to check whether the processing speeds declared by sites as “corepower” (i.e. the HEPSpec06 benchmark score divided by “corepower”) are well correlated with the rates established by this method and measure this correlation quantitatively. As a result, it was decided to continue the study, repeat the analysis and test individual sites for which there is a significant discrepancy from the reference values, measure and adjust the speed of various CPU models with their HS06 estimates.</p>
      </abstract>
      <kwd-group>
        <kwd>HepSPEC06</kwd>
        <kwd>Grid sites</kwd>
        <kwd>ATLAS</kwd>
        <kwd>Kibana</kwd>
        <kwd>speed factor</kwd>
        <kwd>corepower</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Proceedings of the VIII International Conference "Distributed Computing and Grid-technologies in Science and</p>
      <p>Education" (GRID 2018), Dubna, Moscow region, Russia, September 10 - 14, 2018</p>
    </sec>
    <sec id="sec-2">
      <title>1. Background</title>
      <p>ATLAS developed a monitoring system based on Elasticsearch for storing data on all
jobs it runs at different sites worldwide. Kibana, its graphical user interface, is available for
data exploration and rapid prototyping of analyses and it allows to easily create complex
queries.</p>
      <p>In our analysis, data is aggregated in different “buckets” using these variables</p>
      <p>JEDI task ID (a task is a collection of similar jobs, each one running on a




fraction of a dataset)</p>
      <sec id="sec-2-1">
        <title>Site</title>
      </sec>
      <sec id="sec-2-2">
        <title>CPU model</title>
        <p>Processing type (which describes the type of job)</p>
        <p>Inside each bucket, numerical metrics (sums, averages, standard deviations) are
calculated for the: CPU time, wallclock time, number of events, number of jobs, number of
cores etc. Given a set of jobs running on a set of sites (and each one on a certain CPU), we
can assume that if the jobs belong to the same JEDI task, their average CPU time per event is
inversely proportional to the speed of the CPU. We call speed factor a dimensionless number
proportional to the speed of the CPU. If our set of jobs includes several JEDI tasks, one can
think of using simultaneously these tasks to fit the same set of speed factors.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Speed factor calculation</title>
      <p>In the following, α indicates an index running over tasks and i an index running over sites or
only on sites or CPU types where the task runs.</p>
      <p>CPU times. The population average of the CPU time per event for task α at site or CPU i is   and the
average we measure is   . Of course, not all sites or CPU types appear on all tasks, so sums over i run
It is assumed that the   values we measure are Gaussian-distributed around  
with standard
deviation   . It is also assumed that the speed factors   do not depend on α, and in particular that
    =   , where   is a constant with respect to i. It is assumed that also  
=   
 ,∀i,α.</p>
      <p>To make the analysis easier, we choose to use   as errors on   , while it would be more
correct to use   /√  , where   is the number of jobs run at site or CPU type i for task α. We also
assume that   =S, where S is constant for all tasks.</p>
      <p>The  2 we want to minimise is the following:
 2 = ∑ (</p>
      <p>2
−   ) = ∑ (</p>
      <p>2
−   ) = ∑
 
1

= ∑
∝
 2 ∑ (</p>
      <p>− 1) =
   
2
1
 2 ∑

1
 2
∝</p>
      <p>1
 2 ∑ (</p>
      <p>
        2
− 1) =
∑(    −   )
2
types in task α.
meaningful.[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
  . We use as initial values   = 1 and   = (∑   )/  , where   is the number of sites or CPU
As S is a constant, we can minimise a function  =  2 2 and the free parameters are   and
The  2 does not change if we rescale all   and   by the same factor. Just for esthetic reasons
we normalise the speed factors so that their sum is equal to their number (so to be varying around 1);
therefore there is, no absolute scale and only ratios between speed factors are quantitavely
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. HepSPEC06</title>
      <p>
        PanDA is the ATLAS workload management system, and in its architecture it has
the concept of Panda queue (PQ) as an entry point to a computing resource. A PQ has an
attribute called corepower, which is supposed to be the average HS06 score per core for
the CPUs in the worker nodes accessed by that PQ. For each site, HEPSPEC06 was
obtained from AGIS, the ATLAS database containing all information on the computing
infrastructure. We know that these numbers have a low “trust level”, as they are not
accurately validated.[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
      </p>
      <p>
        For each CPU, the HS06 score is obtained from the HEPiX benchmarking working
group. It’s calculated in “ideal” conditions, e.g. HS06 run at boot time on the physical
node. CPUs seen by real jobs can be different: virtual machine overhead, overcommitting,
and most importantly, not knowing if the site enabled hyperthreading (which gives a 1.6
factor of difference in the “corepower”). It is fair to assume that in most cases HT is
enabled.[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ][
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. Stages of work</title>
      <p>Firstly we do data aggregation from ES. Then we do the above speed factor calculations, then
for each job we look at sites which have outstanding speed factor (very high or very low). Then we
made plots for each site with abnormal CorePowers which differ from the most common values by
more than 20％. We gave a closer look to sites that have a speed factor very far from the value that
would correspond to its corepower according to the linear fit. The linear function was chosen with one
parameter so that it passes through the point (0,0), because an site with corepower equal to zero can
have only zero speed factor. For each processing type (job) we identify a few sites which have more
than two errors to make linear fit and count correlation coefficient. So the correlation plots show how
speed factors change with different CPUs.</p>
      <p>For example you can see on Figure 1 that Intel Xeon E5-2650 v2 with Ivy Bridge architecture
gets a very low speed factor compared to Haswell and Broadwell architectures. However, on Figure 2,
you can see that the same CPU looks much faster than others. From this we can conclude that the
difference may be due to factors different from the architecture.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Conclusion</title>
      <p>The accuracy of the speed coefficients for sites compared to the main power is about
20%. this may occur due to many systematic uncertainties and inaccuracies of the main
indicators. This is due to the variability of the core power of sites, for which the speed factor
is checked, which affects the linear fit. However, with respect to speed factors for processors,
the picture is less consistent.</p>
      <p>Correlation graphs for different types of processing on the same sites look very similar
to each other, which indicates that the analysis does not produce “random” numbers. The
points of older CPUs tend to lie below points of newer CPUs on the correlation graphs. But
this is not always the case, indicating the differences between the performance observed by
the application and the “theoretical” performance of HS06 due to other factors.</p>
      <p>This analysis can also be used for other Grid sites, but with adaptation to the site
specifics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Sciabà</surname>
            <given-names>A.</given-names>
          </string-name>
          et al.
          <article-title>CPU benchmarking with production jobs /</article-title>
          /https://indico.cern.ch/event/ 578967/contributions/2455901/attachments/1408462/2153742/CPU_
          <article-title>benchmarking_with_production_ jobs</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Benchmarking</surname>
            <given-names>WG</given-names>
          </string-name>
          // https://w3.hepix.org/benchmarking.html
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Sciabà</surname>
            <given-names>A.</given-names>
          </string-name>
          et al.
          <article-title>Passive benchmarking of ATLAS Tier-</article-title>
          0 CPUs // https://indico.cern.ch/ event/614359/contributions/2510792/attachments/1425728/2187149/Passive_benchmarking_of_ATL AS_
          <article-title>Tier-0_CPUs</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] A comparison of HEP code with SPEC1 benchmarks on multi-core worker nodes</article-title>
          // http://iopscience.iop.org/article/10.1088/
          <fpage>1742</fpage>
          -6596/219/5/052009/meta
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>