<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PARTICIPATION OF RUSSIAN INSTITUTES IN THE PROCESSING AND STORAGE OF ALICE DATA</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A.K.Kiryanov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A.K. Zarochentsev</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey Kiryanov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey Zarochentsev</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Petersburg Nuclear Physics Institute of NRC “Kurchatov Institute”</institution>
          ,
          <addr-line>1 Orlova Rocha, Gatchina, 188300</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Saint Petersburg State University</institution>
          ,
          <addr-line>7-9 Universitetskaya emb., Saint Petersburg, 199034</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>5</fpage>
      <lpage>9</lpage>
      <abstract>
        <p>This report presents the results of the work of Russian institutes in the processing of ALICE experiment data during the last 3 years of the operation of the Large Hadron Collider (LHC) including the end of the LHC Run2 and the 1st year of the COVID-19 pandemic. The main problems and tasks facing both ALICE Grid Computing and its Russian segment before the LHC Run3 including the problems of support and modernization of existing resources are considered. Also, plans for the preparation for the operation of the LHC in the HL (High Luminosity) mode are presented.</p>
      </abstract>
      <kwd-group>
        <kwd>GRID</kwd>
        <kwd>ALICE</kwd>
        <kwd>CERN</kwd>
        <kwd>LHC</kwd>
        <kwd>WLCG</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Challenges</title>
      <p>In 2018 the second stage of the LHC or Run2 was completed and the "Long Shutdown" began
- a long break and preparation for the 3rd stage - Run3 [1]. In addition to updating the experimental
installations themselves, significant improvements in the data storage and data processing system were
also required. Many central services and systems have undergone changes, and in this paper, we will
discuss how these changes affected Russian Tier 1 and Tier 2 sites for ALICE.</p>
      <p>By the end of 2018, RDIG (Russian Data Intensive Grid) [10] provided the following sites for</p>
      <sec id="sec-1-1">
        <title>ALICE:</title>
      </sec>
      <sec id="sec-1-2">
        <title>Tier 1: RRC-KI-T1 Tier 2: RRC-KI-T2, JINR, IHEP, ITEP, PNPI, SPbSU, Troitsk, SARFRTI (SARFTI site was commissioned for ALICE production earlier in 2018, but was not yet officially registered in WLCG) [fig. 1].</title>
        <p>At the end of 2018 the coordinators of Russian sites were facing the following major
challenges, which needed to be addressed in order to continue supporting the ALICE experiment in
Russia:
●
●
●
●</p>
      </sec>
      <sec id="sec-1-3">
        <title>Complete the commissioning of the SARFTI site for WLCG</title>
        <p>Ensure the connection of Russian sites to the LHC ONE network [4] via IPv6 protocol
Return the MEPhI site back to production
Move away from unsupported CentOS 6 and CREAM CE [2] [3], replacing them by CentOS
7 and one of the supported CE flavours (ARC CE [6] or HTCondor CE [5])</p>
        <p>The last bullet in this list required a complete reinstallation and reconfiguration of the majority
of the site’s computing resources.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Work done</title>
      <p>To solve the last task - the change of computing resource manager - the analysis of the
available options for the new managers was conducted, specifically the comparison between ARC 6
and HTCondor. After evaluating both options, ARC 6 was selected as a recommended option for
RDIG sites because of its flexibility and ability to work together with various cluster level resources
managers, including the Torque [7] which is very popular among RIDG. In 2020 the first ARC 6 test
installation was conducted on the SPbSU site. It was registered and tested for the OPS virtual
organization, and later, after successful functional testing, for the ALICE virtual organization as well.
The documentation covering in detail the installation and configuration procedures for the ALICE
Grid sites was created [8], indicating a number of common problems and proposed solutions,
including the code changes that are necessary to support new versions of Torque resource manager [9].
This patch was integrated into the ARC 6 codebase in October 2021. Another problem identified only
on production workloads was an incorrect interpretation of a virtual memory requirement from the
XRSL script that was translated into physical memory requirement for Torque script, causing
submission error due to the lack of resources. This pach was integrated into the ARC 6 codebase
earlier in 2021. Some other patches related to publishing of resource and queue information to the
LDAP information system are still being tested on RDIG sites.</p>
      <p>The transition of RDIG computing resources to the new operating system did not cause
substantial technical problems, but rather organizational - some sites could not start this process for a
long time. On a number of sites the OS upgrade on storage systems was combined with equipment
upgrades, in some cases due to obsolescence of the old equipment, in other cases because of a
necessary maintenance like hard disk replacement. In this situation it was necessary to plan and act
carefully to avoid any damage or loss of the precious experiment data on the storage systems. Initially
one of the disk servers containing the data files was drained with all of the files moved to the other
servers. After the drain the server was replaced with a new one with the old disks replaced and the OS
upgraded. Then the same procedure was repeated with other servers. In the case of the JINR site with
8 servers filled by 65%, the process of draining a single server took up to several days and the overall
migration process took about a month. In the case of the SRFTI site with 4 servers the process of
draining a single server took several times longer because of the problems with hardware. Given that
the drained servers already had problems with disks and a damaged RAID array, it was decided that
the time required for migration does not justify the saving of existing files. At the same time there
were still sites (e.g. PNPI) with an outdated pure XRootD storage system, the migration of which to a
new storage system will require some additional work in the future.</p>
      <p>In 2021 the SPbSU site, the last of the Russian sites working for ALICE, was finally
connected to the LHC ONE network with IPv6 protocol support. It required both an installation of a
new network equipment and a reconfiguration of an existing one because up to that point SPbSU
network did not support the IPv6 protocol. As a result all ALICE RDIG sites are currently connected
to the LHC ONE network with IPv6 protocol.
As it was already mentioned, plans for the 2018 included the return of the MEPhI site back to
production, but the COVID-19 pandemic that began in 2020 has made strong changes to the work
plans of many. In a number of institutions, for a long time access to the servers was physically limited,
just like the number of employees who had the opportunity to conduct the necessary work. As a direct
consequence of these circumstances, the MEPhI site was not returned to production, and, moreover,
one of the sites that was working in 2018, RRC-KI-T2 did not manage to promptly upgrade its
working nodes to the CentOS 7 despite ALICE requirements and was excluded from the ALICE
experimental data processing. As a result, the pledged resources from Russia for 2022 decreased in
comparison to the previous years, despite the fact that some sites including JINR have managed to
increase resources for ALICE [tab.1]. Also, as a consequence of pandemic and inability to service
resources in a timely manner the availability and reliability of RDIG sites decreased as well [fig. 2].</p>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusions</title>
      <p>According to the results of the last three years, of which the last year and a half went under the
COVID-19 pandemic, it can be said that most of the RDIG sites fulfilled their obligations for the
ALICE experiment in terms of computing and storage resources and software updates, but there is
still quite a lot of room for improvement.</p>
      <p>2020 2021 2022
DISK(TB) CPU(HEPSPEC) DISK(TB) CPU(HEPSPEC) DISK(TB) CPU(HEPSPEC)
1200 12000 2000 13500 2300 15525
316 4488 316 4488 0 0
297 2631 314 3017 314 3017
180 2700 180 2700 180 2700
168 2640 168 2640 168 2640
113 641 113 641 113 641
158 3696 158 3696 158 3696
210 7466 210 7466 210 7466
2642 36262 3459 38148 3443 35685</p>
    </sec>
    <sec id="sec-4">
      <title>4. Acknowledgements</title>
      <sec id="sec-4-1">
        <title>This work is supported by the SPbSU grant ID:75252518</title>
        <p>[1] https://lhc-commissioning.web.cern.ch/schedule/LHC-long-term.htm
[2] https://wlcg-ops.web.cern.ch/cream-ce
[3] https://twiki.cern.ch/twiki/bin/view/LCG/CreamMigrationTaskForce
[4] https://lhcone.web.cern.ch/
[5] https://research.cs.wisc.edu/htcondor/
[6] http://www.nordugrid.org/arc/arc6/</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>