<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Statistics of Software Package Usage in Supercomputer Complexes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pavel Shvets</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vadim Voevodin</string-name>
          <email>vadim@parallel.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergey Zhumatiy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Research Computing Center of Lomonosov Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>20</fpage>
      <lpage>29</lpage>
      <abstract>
        <p>Any modern high-performance computing system is a very sophisticated complex that contains a variety of hardware and software components. Performance of such systems can be high, but it reaches its maximum only if all components work consistently and parallel applications e ciently use provided computational resources. The steady trend of all recent years concerning large supercomputing centers { constant growth of the share of supercomputer applications that use o -the-shelf application packages. From this point of view, supercomputer software infrastructure starts to play a signi cant role in the e cient supercomputer management as a whole. In this paper we will show why the analysis of software package e ciency should be taken seriously, what is worth paying attention to, and how this type of analysis can be conducted within the supercomputer complex.</p>
      </abstract>
      <kwd-group>
        <kwd>parallel computing</kwd>
        <kwd>high-performance computing</kwd>
        <kwd>super- computer</kwd>
        <kwd>e</kwd>
        <kwd>ciency</kwd>
        <kwd>e</kwd>
        <kwd>ciency analysis</kwd>
        <kwd>software package</kwd>
        <kwd>job</kwd>
        <kwd>ow monitoring data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Low e ciency of supercomputer system functioning has been an actual problem
in high-performance computing eld for many years. Over the years, many
possible approaches and methods have been developed for analysis and optimization
of both individual supercomputer applications and entire supercomputer
functioning as a whole. Due to increasing large-scale use of this area, i.e. growing
number of users of supercomputer systems, as well as the progress in the
development of HPC packages, o -the-shelf software packages are more and more often
used for solving di erent scienti c problems. For example, around 500-1000
applications are run on Lomonosov and Lomonosov-2 supercomputers daily. Part
of these applications is written by the users themselves with traditional parallel
programming techniques: MPI, OpenMP, OpenCL, CUDA, TBB, etc. The other
part is based on ready-to-use packages from di erent subject areas: Gromacs,
Gaussian, FireFly, CP2K, Amber, VASP, OpenFOAM and many others, and the
share of such applications is steadily growing. This being said, the vast majority
of users also rely on the other types of program packages { system software, for
example, various versions of compilers. And the peculiarities of their usage can
also greatly a ect the performance of user applications.</p>
      <p>
        This leads to the fact that the task of studying the performance of applied
application packages becomes more and more important, since both the speed
of conducting particular scienti c experiments and the overall e ciency of
supercomputer system functioning depend on it. The experience of supporting
and maintaining such systems accumulated in Research Computing Center of
Lomonosov Moscow State University (RCC MSU) shows that such e ciency in
many cases is unfortunately really low, which leads to computational resources
being idle [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. It is necessary to develop and implement special software tools
aimed at analyzing this particular aspect of supercomputer functioning. It is
worth noting that similar researches were conducted earlier (for example, [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]),
but they were originally intended to study only speci c computing systems and
were not portable.
      </p>
      <p>Therefore, RCC MSU decided to study the statistics of software package
usage in MSU Supercomputing Center. The main goal of this study is to
understand how e cient and in-demand di erent software packages are on each
machine. Using the examples of Lomonosov and Lomonosov-2 supercomputers,
in this paper we show what information can be collected and how it can be
obtained. First of all, this is of interest to management and administrators of
supercomputer complexes, since it allows analyzing and improving the e ciency
of the whole supercomputer. In particular, this helps to make a decision whether
a particular package should be further used (e.g. if nobody uses it, there is no
need to renew the license) or whether there is a need of more ne tuning or more
detailed analysis of a particular package with very low usage e ciency. But such
kind of information can be useful for common users as well, since it allows to
make a choice in favor of more e cient package. For example, a user can compare
the performance of two linear algebra packages installed on a particular cluster
and choose a more suitable one, or check the performance of locally deployed
proprietary solution.</p>
      <p>The results shown in this paper were obtained in this Center, but the overall
approach can be quite easily implemented in other supercomputing centers as
well, which is a part of our plans for future.
2</p>
      <p>Collecting Statistics on the Usage of Software Packages
To perform the analysis of package usage, it is necessary to collect information
on all jobs running on the supercomputer; that is, we need to constantly monitor
the overall supercomputer job ow. And it is necessary to collect, integrate and
analyze di erent types of data:</p>
    </sec>
    <sec id="sec-2">
      <title>1. general information on the job ow; 2. data on the e ciency of each particular job; 3. detailed information on the job itself that allows to classify used package and associate it with obtained e ciency data.</title>
      <p>
        Information on the job ow describes when and on which nodes each job was
launched. Such information can be easily collected from the resource manager;
this process is already well known and implemented. In MSU Supercomputing
Center the Slurm manager [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is being used.
      </p>
      <p>
        The data on job execution e ciency is collected via di erent monitoring
systems. These systems provide information on the dynamics of job execution as a
set of characteristics which describe CPU load, intensity of communication
network usage, number of active processes on the node, presence of GPU activity,
etc. This set of required data can di er and depends on the desired criteria of
package usage e ciency. Depending on this, di erent monitoring systems
(Collectd [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Zabbix, Nagios, Zenoss, Cacti or DiMMon system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] being developed
in RCC MSU) may be more appropriate. A set of proprietary tools is currently
used in MSU Supercomputing Center for collecting needed monitoring data.
Since these are the rst steps in this study, for now we are interested only in the
most general indicators of computational resource usage { CPU and GPU user
load as well as average system load (loadavg), but in the future we also plan to
analyze such characteristics as the intensity of memory subsystem and network
usage.
      </p>
      <p>
        Note that in order to collect statistics on the package usage, it is su cient to
collect and store only integral characteristics for each launch, and this task can
be successfully solved by existing monitoring systems. Moreover, this information
was being stored for a while in MSU Supercomputing Center for other purposes
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], so this data was already available for our research.
      </p>
      <p>
        Thus, in this study the main di culty in collecting the required data was to
solve the last issue { to collect detailed information on each running job about
the packages used. We were interested in information about used compilers,
application packages, linked libraries, etc. It was decided to use XALT tool [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]
developed by Texas Advanced Computing Center (TACC) for solving this issue.
This tool was chosen due to several reasons:
      </p>
    </sec>
    <sec id="sec-3">
      <title>1. it is an o -the-shelf system software product that is ready for use;</title>
      <p>2. it has all the functionality needed for our purposes;
3. it is open source.</p>
      <p>This tool collects the data by intercepting calls to the linker and job launcher.
For this purpose, wrappers are created that override standard front-end scripts.
Each time a user links or launches his program, XALT stores all necessary
information on performed actions in user home directory in JSON format. In
particular, used compiler and its options, list of linked static libraries, full path
to the compiled le and compilation time are stored during linking operation.
When the job launcher is used, XALT stores the name on executed le, the way
it was launched, used MPI library version, linked dynamic libraries, values of
environmental variables and many other useful information. Other ways of
storing the data (besides writing JSON les in home user directories) are possible {
using syslog or direct connection to the database.</p>
      <p>The collected data periodically (or by administrator request) is stored in a
single MySQL database which accumulates all statistics on the executed jobs.
After that, all that needed is to compose correct SQL query in order to obtain
corresponding data slice required for analysis. Though it is necessary to learn
the (quite simple) database format in order to manually create needed SQL
queries, many ready-to-use examples of such queries enough to perform rst
quick analysis are provided in the XALT manual. This MySQL database is built
to aggregate data from di erent clusters, so it is quite easy to analyze the data
about the whole supercomputer complex as well as about one particular machine.
3</p>
      <p>Analysis of Results Obtained in MSU Supercomputing
Center
The speci ed XALT tool was installed in MSU Center in June 2017, so some
information is available only from that moment. However, during XALT
installation and testing it was discovered that XALT scripts for de ning software
packages can be applied to all previously collected data for executed jobs, so the
results presented in Chapters 3.1 and 3.2 are based on data since the beginning
of year 2017. Below are the main interesting results that were obtained in this
study.
3.1</p>
      <sec id="sec-3-1">
        <title>Statistics on the Popularity of Application Packages</title>
        <p>To ensure the e ciency of supercomputing centers we need to move from \our
feelings what is needed" to real facts. From this point of view, even a simple
general statistics on the frequency of using application packages is already very
useful. Figure 1a and 1b show the distribution of the frequency of running certain
packages on Lomonosov and Lomonosov-2 supercomputers. The frequency of
package usage is shown as a share (in percents) of the total number of runs on
a supercomputer. It can be seen that in both cases top 3 places in total provide
more than 50% of all package launches, but there is only one package that is
in both top 3 lists { Fire y package for quantum chemical calculations. In the
whole list some overlaps can be seen: for example, packages like VASP, Gromacs,
NAMD and DL POLY are frequently used on both systems.</p>
        <p>In general statistics on both supercomputers is quite di erent, and this
illustrates the fact that there is no big reason to compare such results on di erent
machines. This is because the architecture of supercomputers, as well as the
overall research interest of users (and therefore the package usage) can vary greatly
in each case. The main goal of analyzing such kind of information is to get the
basic understanding of package usage on each particular cluster independently.</p>
        <p>Now let's look at the use of packages from the other point of view and evaluate
the CPU hours consumed by various packages (Fig. 2a and 2b for Lomonosov
and Lomonosov-2 supercomputers correspondingly). The picture changes
signi cantly; for example, the leading position has changed { in both cases the
rst place is now occupied by the Gromacs package, which is used to model
physicochemical processes in molecular dynamics. It can also be noted that the
DL POLY package is not represented on these charts, although it was in top
lists in Fig. 1. This indicates that this package is launched very often, but on
average it requires not so many computational resources.</p>
        <p>Taking into account that dozens of packages are installed and used on
supercomputers, such analysis helps to prioritize their installation, updating and
ne-tuning. Together with similar data, but speci ed for certain users, this
information shows the actual demand for particular packages.
Another huge part of information is provided by the binding of dynamic
characteristics describing the e ciency of parallel job execution to speci c software
packages. Figures 3a and 3b describe average value of loadavg for jobs that use
di erent packages on Lomonosov and Lomonosov-2 supercomputers. Loadavg
parameter estimates the number of processes on a node that are ready for
execution, which describes degree of parallelism in an application from a system
point of view. If an application makes the best use of hardware platform
capabilities, this value is usually equal to the number of cores on a node. Nodes
of Lomonosov supercomputer have two 4-core or 6-core processors, nodes of
Lomonosov-2 { one 14-core processor. The corresponding optimal loadavg
values based on the number of physical cores are shown with horizontal lines in
Fig. 3. Note that processors in both systems have Hyper-Threading
technology enabled, which doubles the number of logical cores, so potentially optimal
loadavg value can be twice as high.</p>
        <p>It can be seen in Fig. 3 that many packages show high loadavg values
meaning that they use available resources quite actively, although there are notable
exceptions like Molpro package used on Lomonosov or VASP used on
Lomonosov2. It is also interesting to note that Gromacs package on Lomonosov seems to
use Hyper-Threading quite often (its loadavg value is higher than the number
of physical cores on a node), since apparently this allows further performance
increase.</p>
        <p>Such data is very important: it re ects the accumulated experience and helps
to understand the way the packets are used by the computing community. This
is necessary as well for evaluating the performance of the package on a
particular supercomputer, and this also helps to make adjustments to the use of the
particular package by individual users.</p>
        <p>The degree of parallelism is de ned not only by the number of used cores
on a node, but also by an overall number of used nodes. Our practice shows
that, in an attempt to reduce the running time of applications, users increase
the number of nodes, not always paying attention to the limited scalability of
applications. As a result { part of supercomputer resources is wasted. A good
method for evaluating e ciency is to analyze how CPU cores are loaded by each
package. An example showing average core load for Lomonosov and Lomonosov-2
supercomputers is shown in Fig. 4.</p>
        <p>Comparison of Fig. 3 and 4 shows that loadavg and core load characteristics
correlate well enough { relative change of loadavg for di erent packages is similar
to corresponding change in the values of the core load. This indicates that on
average each packet equally loads each active process. However, there are
exceptions { the loadavg value for WRF package on Lomonosov-2 supercomputer is
much higher than the values for other packages, but its core load is at average
level. Perhaps this is due to very active memory usage or data transfer using
communication network, but it is also likely that this package uses resources
rather ine ciently. We should notice that this is true only for this particular
system; the behavior of WRF package on other systems can di er.</p>
        <p>In general, low core load can be explained by di erent reasons: bad
package scalability, errors made by user or his desire to allocate more memory per
process on a node. There are many possible reasons, and each case must be
considered separately. In any case, such information should be available to system
administrators to control the situation.</p>
        <p>Similarly, package usage e ciency can be analyzed using other monitoring
data, for example, GPU usage load (also shown in Fig. 4). The package can be
optimized for GPU usage, but only real-life statistics can show, how e ectively
someone actually uses it and whether he forgot to specify the \use the GPU"
option. In particular, Fig. 4b explains low loadavg value for VASP package:
the graph shows that this package uses GPU very actively, and apparently the
computations are almost not performed on the CPU side, so it does not need to
start a lot of CPU of processes.</p>
        <p>Many packages are sensitive to hardware parameters, and the analysis of
dynamic characteristic values helps to optimize both overall functioning of
supercomputer complex and behavior of particular user applications. In this
regard, it is very important to detect anomalous values of dynamic characteristics
in time, which again may happen due to di erent reasons like incorrect
package installation by the system administrator or incorrect package usage by the
common user. There are many examples of anomalies: excessively high or low
loadavg value, high core idle value, large number of cache misses, low GPU load,
high intensity of I/O operations, high uctuation of dynamic characteristics, etc,
but all of them indicate potential problems with the e ciency of supercomputer
complex. In this research it is planned to develop a methodology for detecting
such anomalies in future.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Distribution of Used Compilers</title>
        <p>Results shown earlier relate directly to the job launch process. But linking
process is also of interest. In particular, one of the main questions { which compilers
are most in demand? Figure 5 shows the distribution of compilers based on the
number of uses on Lomonosov-2 supercomputer. It can be seen that GNU
compilers are used much more often than Intel compilers. For instance, GNU C
compiler is used in almost half the cases. Interestingly, Fortran language,
according to this statistics, is used almost as often as C language (44% vs 50%),
and C++ language is far less common (only 6%).
This paper describes the research being conducted in RCC MSU aimed at
studying the e ciency of software package usage in supercomputer systems. This
research on the given examples shows how such analysis can be conducted within
the majority of modern supercomputers, and what information can be obtained
on its basis. We engage the management and system administrators of di erent
supercomputer centers to perform such study, since this helps to improve the
e ciency of the supercomputer in general, and also help users to select more
suitable software packages for their tasks.</p>
        <p>This research is based on the analysis of di erent types of information: data
on the overall job ow that helps to determine when and on which nodes each
job was launched; data on the e ciency of each particular job; detailed data on
the package usage. The task of collecting rst two types of data is easy enough
and was already implemented in MSU Supercomputing Center, so only the last
question was to be solved. XALT package developed in TACC was used for
this purpose. This package intercepts calls to the linker and job launcher, which
allows to save required information each time they are used.</p>
        <p>As a result, various statistics on package usage on Lomonosov and
Lomonosov2 supercomputers were collected. In particular, the distribution of packages by
the number of launches, consumed CPU hours, average CPU load and loadavg,
as well as the frequency of using di erent compilers was analyzed. Such analysis
allows to perform holistic estimation of the popularity and usage e ciency of
di erent software packages.</p>
        <p>The analysis of the results described in this paper is just the rst step in
this direction. A variety of services intended to help common users and system
administrators to optimize their work can be developed based on this data, which
will be the focus of further work in this study.</p>
        <p>The data on average dynamic characteristics for a package will help the user
to correctly interpret the values obtained by his applications. One possible way
of implementation { automatically generated noti cation based on the results
of application execution, which contains a comparison of the current launch
characteristics with the average characteristics for this package. This can be a
mass service for all users with built-in analytics for each application package.
This service can also be con gured to compare only with the previous launches of
the same user, so a deviation from \standard" values speci c for this particular
user can be clearly seen.</p>
        <p>System administrators also require similar services that constantly analyze
the integral characteristics of di erent packages. Such services can help to nd
out when a package reinstallation, software update or issue of a \How to e
ciently use a package" guide is needed. Slightly more complex options are
searching for more e cient alternative packages (based on price and e ciency criteria)
or planning a software update budget for the next year.</p>
        <p>Acknowledgments. The results described in all sections except Section 2 were
obtained in the Lomonosov Moscow State University with the nancial support
of the Russian Science Foundation (agreement № 17-71-20114). The research
presented in Section 2 was supported by the Russian Foundation for Basic Research
(№ 17-07-00664).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Voevodin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vadim</given-names>
            <surname>Voevodin</surname>
          </string-name>
          .
          <article-title>E ciency of Exascale Supercomputer Centers and Supercomputing Education</article-title>
          .
          <source>In High Performance Computer Applications: Proceedings of the 6th International Supercomputing Conference in Mexico (ISUM</source>
          <year>2015</year>
          ), pages
          <fpage>14</fpage>
          {
          <fpage>23</fpage>
          . Springer, Cham,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Nikitenko</surname>
          </string-name>
          , Vladimir Voevodin, Alexey Teplov, Sergey Zhumatiy, Vadim Voevodin, Konstantin Stefanov, and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Shvets</surname>
          </string-name>
          .
          <article-title>Supercomputer Application Integral Characteristics Analysis for the Whole Queued Job Collection of Large-scale HPC Systems. In Parallel Computational Technologies (PCT'</article-title>
          <year>2016</year>
          ), pages
          <fpage>20</fpage>
          {
          <fpage>30</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Matthew</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>Joseph P.</given-names>
          </string-name>
          <string-name>
            <surname>White</surname>
            , Martins Innus, Robert L. DeLeon, Nikolay Simakov, Je rey
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Palmer</surname>
          </string-name>
          ,
          <string-name>
            <surname>Steven M. Gallo</surname>
          </string-name>
          ,
          <string-name>
            <surname>Thomas R. Furlani</surname>
            , Michael Showerman, Robert Brunner, Andry Kot, Gregory Bauer, Brett Bode, Jeremy Enos, and
            <given-names>William</given-names>
          </string-name>
          <string-name>
            <surname>Kramer</surname>
          </string-name>
          .
          <article-title>Workload Analysis of Blue Waters</article-title>
          . mar
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Slurm</given-names>
            <surname>Workload</surname>
          </string-name>
          <article-title>Manager</article-title>
          . URL: http://slurm.schedmd.
          <source>com. Cited 08 Aug 2017</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Collectd {
          <article-title>The system statistics collection daemon</article-title>
          . URL: https://collectd.org.
          <source>Cited 08 Aug 2017</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Konstantin</given-names>
            <surname>Stefanov</surname>
          </string-name>
          , Vladimir Voevodin, Sergey Zhumatiy, and
          <string-name>
            <given-names>Vadim</given-names>
            <surname>Voevodin</surname>
          </string-name>
          .
          <article-title>Dynamically Recon gurable Distributed Modular Monitoring System for Supercomputers (DiMMon)</article-title>
          .
          <source>Procedia Computer Science</source>
          ,
          <volume>66</volume>
          :
          <fpage>625</fpage>
          {
          <fpage>634</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Kapil</given-names>
            <surname>Agrawal</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mark R. Fahey</surname>
          </string-name>
          ,
          <string-name>
            <surname>Robert McLay</surname>
            ,
            <given-names>and Doug</given-names>
          </string-name>
          <string-name>
            <surname>James</surname>
          </string-name>
          .
          <article-title>User Environment Tracking and Problem Detection with XALT</article-title>
          .
          <source>In 2014 First International Workshop on HPC User Support Tools</source>
          , pages
          <volume>32</volume>
          {
          <fpage>40</fpage>
          . IEEE, nov
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>8. XALT tool homepage</article-title>
          . URL: https://www.tacc.utexas.edu/researchdevelopment/tacc-projects/xalt. Cited 08 Aug 2017
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>