<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CURRENT STATUS OF DATA CENTER FOR COSMIC RAYS BASED ON KCDC</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>V.A. Tokareva</string-name>
          <email>victoria.tokareva@kit.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D.G. Kostunin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Haungs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Nuclear Physics, Karlsruhe Institute of Technology</institution>
          ,
          <addr-line>Hermann-von-Helmholtz-Platz 1, Eggenstein-Leopoldshafen, 76344</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>405</fpage>
      <lpage>409</lpage>
      <abstract>
        <p>We present the current status of a data center based on KCDC (KASCADE Cosmic-ray Data Center), which was designed for providing open access to the events measured and analyzed by KASCADEGrande, a cosmic-ray experiment located at KIT, Karlsruhe. In the frame of the German-Russian Astroparticle Data Life Cycle Initiative we extend KCDC in order to provide an access to different cosmic-ray experiments and make possible aggregation and joint querying of heterogeneous airshower data. In the present talk we discuss the description of data and metadata structures, implementation of data querying and merging, and first results on including data of experiments located in Tunka, Russia, in this common data center.</p>
      </abstract>
      <kwd-group>
        <kwd>astroparticle physics</kwd>
        <kwd>data life cycle management</kwd>
        <kwd>cosmic rays</kwd>
        <kwd>metadata</kwd>
        <kwd>open data</kwd>
        <kwd>distributed computing</kwd>
        <kwd>cloud computing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Nowadays, important results in the field of astroparticle physics are achieved because of a
rapid evolution of detector electronics and algorithms and modern computational methods (distributed
computing, machine learning), which has come in use in recent years. According to Astroparticle
Physics European Coordination committee (APPEC), the common data rate for astrophysical
experiments all together is a few PBytes/year, which is comparable to the current LHC output [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Hence, from one side we face a big amount of data, and from the other side there is a growing
complexity of the computations we would like to perform. To achieve the stable reliable work of data
processing at every stage, a proper data life cycle management is required. Thereby, we discuss
organizing such a cycle in a framework of German-Russian astroparticle data life cycle initiative and
discuss issues of the data integration and data workflow management.
      </p>
      <sec id="sec-1-1">
        <title>1.1. The project’s aim and objectives</title>
        <p>The aim of the initiative is implementing the data life cycle, i.e. the data management pipeline,
which would suit well for the specific requirements we face in the field of astroparticle physics today,
to name the major:
 Relatively big (from 100 to 1000 Tb/year) amounts of data modern astroparticle experiments
have to collect in order to observe the messengers of our interest in the most precise way;
 Relatively small amount of events, which are in high interest of us;
 Interest in joint data analysis from various observations for statistics increasing;
 Potentially different structure of data, accumulated with different types of detectors, for example,
for the high energy cosmic ray data collected with Cherenkov light detectors and data about the
same messenger taken with scintillator arrays;
 General high demand for open access data sharing for research, outreach and education, which is
rapidly gaining ground in particle physics nowadays.</p>
        <p>For gaining the aim proposed we develop the common data life cycle for two astroparticle
physics projects KASCADE-Grande [2] and TAIGA [3], organized for at studying high-energy
cosmic rays by observing extensive air showers (EAS).</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2. KASCADE-Grande experiment</title>
        <p>The KASCADE [2] experiment has definitely become a significant milestone in astroparticle
physics. It was started in the beginning of 90-s and aimed at studying high energy cosmic rays by
observing extended air showers. The setup consisted of several detectors: scintillation counters of
various techniques, in particular, numerous KASCADE and Grande scintillation detectors; an hadronic
calorimeter; and a digital radio array LOPES. The experiment has been running for more than 20 years
and has collected important data. It’s analysis has already led to several notable physical results (see
[4−6] and refs. therein) and is still ongoing.</p>
        <p>KASCADE-Grande is the first and still the only big cosmic ray experiment that has
completely published it’s data. Specialized web-portal KCDC [7] (https://kcdc.ikp.kit.edu) has been
developed for this purpose and we suppose that it’s infrastructure could become a good basis for joint
data access and analysis for both the experiments and the possible extensions.</p>
      </sec>
      <sec id="sec-1-3">
        <title>1.3. TAIGA experiment</title>
        <p>TAIGA [3] detectors are located in Russia in Tunka Valley near Baikal lake. It consists of
various detector types, including: Cherenkov photomultipliers of various sensitivity; scintillation
counters; a radio detector array; and modern Imaging Air cherenkov Telescopes (IACT). The setup is
currently operating and still being extended.</p>
        <p>The combined data analysis from all these detectors is already a challenging task, but we
decided to start with a bit different approach and perform a combined analysis of data from
KASCADE-Grande and Tunka-133. Both the detectors have observed cosmic rays, but in different
time, in different locations and different atmospheric conditions. One of them consists of scintillating
counters, the other one from Cherenkov photo multipliers.</p>
        <p>The data amount for this task is fairly limited but looking forward into the perspective of
extending our solution for the whole experiments data we are considering the ways to make our
solution scalable enough.</p>
      </sec>
      <sec id="sec-1-4">
        <title>1.4. KASCADE Cosmic Ray Data Center</title>
        <p>The web portal KCDC [7] provides the access to data of the KASCADE-Grande setup for the
interested public, i.e. for professional physicists from astroparticle physics community, lecturers,
students, and all broad audience. The IT infrastructure of the portal includes highly-demanded
technologies, such as interactive data selections, high-availability message exchanging and dynamic
task distribution.</p>
        <p>We work on extending KCDC by integrating the TAIGA data into the data center workflow,
see Figure 2. For a fast data search the metadata-based approach described further is proposed.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Data integration</title>
      <p>The problem of data integration is of interest because there is no generally accepted solution in
the field of astroparticle physics so far. Since astroparticle physics is an observational science, and one
cannot influence the phenomena that generate data, the experiments tend get as much as possible from
the data they are able to obtain, which results in a data-driven approach for the data analysis. Both
conventional algorithmic analysis and deep learning using neural networks can be used, but the
problem of scaling of heterogeneous data arises in both cases.</p>
      <sec id="sec-2-1">
        <title>2.1. Data storage and retrieval</title>
        <p>The data corresponding to different experiments possess many essential distinctions, that
makes the problem challenging. The data format itself is different, as each detector measures its own
set of observables. Then, the software for the initial steps of data analysis is strongly depending on the
detector and could require special software environment of libraries installed. These differences result
in the first problem that is creating a unified software interface for working with data. Currently, the
whole analysis procedure is designed separately and incompatibly for both experiments, KASCADE
and TAIGA. Our goal is to develop the unified interface, where the common analysis steps would be
designed in a common way, and provide well-enough encapsulation of low-level details to hide them
from end user.</p>
        <p>The basic functionality of the data management pipeline is to provide data to scientists on
demand. A request is a set of conditions and logical operations on them that determine which data the
user wants to receive.</p>
        <p>Data storage is going to be organized distributedly by the means of a virtual file system
deployed on top of existing servers, the most likely candidate being considered is CernVM-FS [8].</p>
        <p>Since the data size is huge and its structure is diverse, a direct search within the data would be
extremely slow and resource-consuming, and thus is not going to be implemented. Fortunately, the
data have a common metadata format, which includes time, place, atmospheric conditions, etc. A
centralized database containing the metadata of all events from both experiments will be used to
process data-retrieval requests. The proposed database structure is presented in Figure 1. In case any
kind of requests using the properties not included in this database prove to be necessary, the
appropriate information must be extracted from the data and inserted into the metadata registry.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Data analysis</title>
        <p>Besides the delivery of requested data, the task of data analysis requires access to computer
facilities and software for the analysis. Data analysis generally includes two types of sub-tasks:
simulations and analyzing experimental data.</p>
        <p>Simulation for the physics of air showers consists of simulating two consecutive processes:
first, the shower propagation in the atmosphere, and second, the detector response. Both ones are
performed independently for each event, require a lot of computing power and produce data volumes
comparable with that of experimental data. Thus, parallel computing is a natural solution in this case.
Simulations of a shower uses standard CORSIKA software [9] for any experiment and does not
depend on detector features. Instead, detector simulations require the software that is usually
developed separately for each particular experimental setup, and a special computing environment for
it, like libraries or OS versions.</p>
        <p>The same requirement of a dedicated environment applies to the software for analyzing
experimental data. Distributed execution of such type of tasks could be organized using virtualization
techniques, when an image or container with all necessary software is deployed at each processing
node.</p>
        <p>The demands to keep the uniform CPU load for the computing nodes and to minimize the data
transfer overheads result in the necessity to employ a workload management system (WMS) for the
distributed analysis. We are going to undertake a comparative performance analysis for various WMS
in order to find the one most suitable for the extended KCDC portal integrating KASCADE and
TAIGA data.</p>
        <p>The workflow scheme for the joint data analysis is presented in Figure 2 and includes the
following steps. First, data acquisition is being done for each experiment, resulting in a set of
registered events. For each event dedicated simulations are performed. These data are used to
reconstruct the shower properties and the characteristics of an initial particle.</p>
        <p>In order to perform a further joint analysis of events from different experiments, data mapping
is required, that consists of comparing preliminary distributions of various observables, finding a
common scale and normalizing the data according to it, followed by a new iteration of data
reconstruction when necessary.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusion</title>
      <p>The constantly growing amount of accumulated astroparticle data and the request for the
multi-messenger astronomy as well as sophisticated analysis methods, like machine learning, requires
to develop a unified system for astroparticle data storage and processing.</p>
      <p>In a framework of German-Russian astroparticle data life cycle initiative we proposed the
concept of astroparticle data center based on KCDC. Our plans include organization of distributed data
storing and processing, creating a platform for joint data analysis for both experiments and providing
the data and analysis results to the public open access, as well as using them for the educational and
outreach activity. With taking into account the specific features of the research field and keeping in
mind the data-oriented approach, we proposed the structure of the metadata database and the possible
data workflow scheme.</p>
      <p>The built-up infrastructure is supposed to be used to analyze combined data sets with large
statistics, allowing to study galactic sources of high-energy γ-rays and cosmic rays, which could be a
notable step forward in multi-messenger astroparticle physics.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] Berghöfer T.,
          <string-name>
            <surname>Agrafioti</surname>
            <given-names>I.</given-names>
          </string-name>
          et al.
          <article-title>Towards a Model for Computing in</article-title>
          European Astroparticle Physics // Astroparticle Physics European Consortium − 2016
          <source>− arXiv:1512.00988 [astro-ph.IM]</source>
          [2]
          <string-name>
            <surname>W.-D.</surname>
          </string-name>
          Apel et al. // Nucl. Instr.
          <source>Meth. Phys. Res. A −</source>
          2010
          <string-name>
            <surname>− V.</surname>
          </string-name>
          620 − P.
          <volume>202</volume>
          [3]
          <string-name>
            <given-names>N.M.</given-names>
            <surname>Budnev</surname>
          </string-name>
          et al. // Phys. Part.
          <string-name>
            <surname>Nucl</surname>
          </string-name>
          . − 2018
          <string-name>
            <surname>− V.</surname>
          </string-name>
          49 − P.
          <volume>589</volume>
          [4]
          <string-name>
            <surname>W.-D.</surname>
          </string-name>
          Apel et al. // Astropart. Phys. − 2012
          <string-name>
            <surname>− V.</surname>
          </string-name>
          36 − P.
          <volume>183</volume>
          [5]
          <string-name>
            <surname>W.-D.</surname>
          </string-name>
          Apel et al. // Phys.
          <string-name>
            <surname>Rev. D −</surname>
          </string-name>
          2013
          <string-name>
            <surname>− V.</surname>
          </string-name>
          87 − P.
          <volume>081101</volume>
          [6]
          <string-name>
            <surname>W.-D.</surname>
          </string-name>
          Apel et al. // Astroph. J. − 2017
          <string-name>
            <surname>− V.</surname>
          </string-name>
          848 − P.
          <volume>1</volume>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Haungs</surname>
          </string-name>
          et al. (
          <string-name>
            <surname>KASCADE-Grande</surname>
            <given-names>Collaboration</given-names>
          </string-name>
          ) // Eur. Phys.
          <string-name>
            <surname>J. C −</surname>
          </string-name>
          2018 −
          <volume>78</volume>
          :
          <issue>741</issue>
          [8]
          <string-name>
            <given-names>J</given-names>
            <surname>Blomer</surname>
          </string-name>
          et al.
          <article-title>Delivering LHC Software to HPC Compute Elements with CernVM-FS // ISC High Performance 2017</article-title>
          . Lecture Notes in Computer Science / eds. Kunkel et al. − Springer −
          <year>2017</year>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Heck</surname>
          </string-name>
          et al. CORSIKA:
          <string-name>
            <given-names>A Monte</given-names>
            <surname>Carlo Code to Simulate Extensive Air Showers − Karlsruhe: Forschungszentrum Karlsruhe GmbH − 1998</surname>
          </string-name>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>