<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A DISTRIBUTED DATA WAREHOUSE SYSTEM FOR ASTROPARTICLE PHYSICS</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minh-Duc Nguyen</string-name>
          <email>nguyendmitri@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Kryukov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julia Dubenskaya</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Korosteleva</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stanislav Polyakov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evgeny Postnikov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Bychkov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey Mikhailov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexey Shigarov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleg Fedorov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yulia Kazarina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dmitry Shipilov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dmitry Zhurov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Applied Physics Institute, Irkutsk State University</institution>
          ,
          <addr-line>Irkutsk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lomonosov Moscow State University, Skobeltsyn Institute of Nuclear Physics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Matrosov Institute for System Dynamics and Control Theory, Siberian Branch of Russian Academy of Sciences</institution>
          ,
          <addr-line>Irkutsk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>2018 Minh-Duc Nguyen, Alexander Kryukov, Julia Dubenskaya, Elena Korosteleva</institution>
          ,
          <addr-line>Stanislav Polyakov, Evgeny Postnikov, Igor Bychkov, Andrey Mikhailov, Alexey Shigarov, Oleg Fedorov, Yulia Kazarina, Dmitry Shipilov, Dmitry Zhurov</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>419</fpage>
      <lpage>423</lpage>
      <kwd-group>
        <kwd>data warehouse system</kwd>
        <kwd>remote data access</kwd>
        <kwd>online data analysis</kwd>
        <kwd>astroparticle physics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In this paper, we describe the current status of the data warehouse system at
Astroparticle.online. Astroparticle.online 0 is a project funded by the German-Russian Astroparticle
Data Life Cycle Initiative to support physical experiments in astroparticle physics such as TAIGA 0
and KASCADE-Grande 0. The primary goal of the project is to create a smart mechanism to distribute
data at each site where the experiments take place to scientists so that they can use together data from
different experiments in their research for analysis as if the files are locally available in their
computers.</p>
      <p>There are some requirements for such a data warehouse system due to the computing facility
at the experiments' sites. First of all, the implementation of the system must not lead to any changes in
the existing hardware and software infrastructure at the sites. All new components must be added to
the sites as independent modular add-ons. Second, all work that might cause high CPU loading to the
computing facility of the cites should be done somewhere else. Third, only the exact amount of
necessary data should be transferred from the sites due to limited bandwidth and a large amount of
accumulated data. And finally, the data are read-only for all users. Changes to data must be done only
on-site not online.</p>
      <p>All described above requirements narrow the searching scope and lead to three major existing
open-source solutions: CernVM-FS 0, HDFS 0 and OpenAFS 0. In this paper, we will describe our
attempt to build the targeted data warehouse system using CernVM-FS and the problems we are facing
during the process. The structure of this paper is as follow. In the second section, we take a deep dive
into CernVM-FS to see how it works. In the third section, we explain how we build our data
warehouse system using CernVM-FS. In the last section, we point out what we managed to do and the
future plan.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Cern VM-FS</title>
      <p>CernVM-FS is a widely used file system at CERN, the European Organization for Nuclear
Research, to distribute the software to be used in physical experiments to scientists. The key idea is
that CernVM-FS indexes only the metadata and the directory structure and distribute them to users.
The metadata and directory structure are very small, so it is easier and faster to be transferred. Users
can browse through the remote catalogue right after mounting it to a mount point in the OS using
CernVM-FS client. The content of a file is fetched and delivered to users only on actual reads so a lot
of unnecessary data transfer can be avoided.</p>
      <p>C ERN V M -FS D A TA</p>
      <p>LI FEC YC LE</p>
      <sec id="sec-2-1">
        <title>Release M anager</title>
      </sec>
      <sec id="sec-2-2">
        <title>Dat a St orage</title>
      </sec>
      <sec id="sec-2-3">
        <title>HTTP Server</title>
        <p>The data life cycle in Cern VM-FS is one directional. Data are available as read-only to users.
Changes can be made only from the server side. After a careful review, a data release manager creates
a new release of the data and upload it to the central repository controlled by a CernVM -FS server.
This process is called a transaction. During a transaction, the new catalogue of files and folders is
indexed and added as flat objects into a "big bag" similar to a Git-repository 0. After the confirmation
from the release manager, the transaction is published.</p>
        <p>C ERN V M -FS D A TA</p>
      </sec>
      <sec id="sec-2-4">
        <title>D I STRI BU TI O N</title>
        <p>CVM -FS
CVM -FS
CVM -FS</p>
        <p>PRO XY
SERVER</p>
        <p>PRO XY</p>
        <p>SERVER</p>
        <p>The CernVM-FS repository is distributed through an https server. In terms of CernVM-FS an
https server at the root level where the data catalogue itself is located is called stratrum-0. Since the
content is distributed through an https server, it can be cached and redistributed through a proxy
server. The redistribution proxy server is called stratum-1. Different combinations of stratum-0 and
stratum-1 servers are possible as shown in figure 2. A user can then use the CernVM-FS client to
mount the remote catalogue to a mount point in her computer. The content of files is fetched from the
stratum-0 or stratum-1 server to the user’s machine on-read. As long as the files are in the local cache
of the CernVM-FS client they can be accessed offline.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The astroparticle.online data warehouse system</title>
      <p>At each site of the physical experiments (TAIGA and KASCADE-Grande), the result after
each run is saved as files in specific binary formats. When a physicist investigates an event, she
collects all files related to the event and analyses them. The process when the physicist searches for
necessary files is called preliminary search. The goal of the data warehouse system is to make this
process as fast as possible. Once necessary files have been found, we would want the files to be able to
be reused as long as possible without retransferring them from the root server. We also want to
implement a light data access policy to prevent users from seeing the entire catalogue of data from all
sites. Users can see only files related to their queries each time. Taking into account all requirements
above, we started building a prototype of the targeted data warehouse system based on CernVM-FS.
The overall architecture of the system is shown below.</p>
      <sec id="sec-3-1">
        <title>HTTP Server</title>
      </sec>
      <sec id="sec-3-2">
        <title>HTTP Server</title>
        <p>Each data storage server at each experiment site contains an instance of the CernVM-FS
server. The existing data storage is plugged into Cern VM-Fs server as an external data source. In this
case, CernVM-FS is still able to index the data catalogue, but the data catalogue itself is still managed
by the old accepted way at the site. We don't want to make a full copy of the whole data catalogue in
another partition for indexing it by CernVM-FS. All stratum-0 instance data catalogues are mounted to
a stratum-1 proxy server. We call it the aggregator. At this level, we fetch the metadata from the raw
files and create the indices based on a set of parameters that later are used by users in their search
queries such as event start time, event end time, energy range, site location etc. The data access policy
is also implemented at the aggregator level. Another critical function of the aggregator is caching the
most frequently accessed files and search result. The aggregator also serves as an API server. From the
user side, one can use a customised client that talks with the aggregator using the API to find a
collation of files, then the client uses CernVM-FS client to mount those files to the local computer.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>At the moment we were able to merge all data catalogues from different sites into a big one
and redistribute it to users. Data caching is done by CernVM-FS out-of-the-box. Metadata indexing is
still an ongoing work. The most significant issue with CernVM-FS is that the whole data catalogue is
visible and accessible for everyone by design. We are looking for a solution to implement the data
access policy that requires a minimum change in the CernVM-FS core.</p>
      <p>In the future, we plan to build other prototypes using HDFS and OpenAFS with the same
requirements. After that, we will compare the prototypes by metrics including the complexity, spent
time and effort to adapt each solution to build the target data warehouse system and to maintain it. We
also plan to conduct a performance benchmark.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgment</title>
      <p>This work is supported by the RSF grant #18-41-06003.</p>
      <p>Initiative.</p>
      <p>Available
Guide.</p>
      <p>Available at:
Guide.</p>
      <p>Available</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>German-Russian Astroparticle</surname>
          </string-name>
          Data Life https://astroparticle.online/about (accessed
          <volume>08</volume>
          .10.
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Budnev</surname>
            <given-names>N.</given-names>
          </string-name>
          et al.
          <article-title>The TAIGA experiment: From cosmic-ray to gamma-ray astronomy in the Tunka valley // Nuclear Instruments and</article-title>
          Methods in Physics Research Section A: Accelerators, Spectrometers,
          <source>Detectors and Associated Equipment February</source>
          <year>2017</year>
          : Vol.
          <volume>845</volume>
          , pp
          <fpage>330</fpage>
          -
          <lpage>333</lpage>
          . - DOI: 10.1016/j.nima.
          <year>2016</year>
          .
          <volume>06</volume>
          .041
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Apel W.D</surname>
          </string-name>
          . et al.
          <source>The KASCADE-Grande Experiment // Nuclear Instruments and Methods in Physics Research Section A 620 April</source>
          <year>2010</year>
          : pp
          <fpage>202</fpage>
          -
          <lpage>216</lpage>
          - DOI: 10.1016/j.nima.
          <year>2010</year>
          .
          <volume>03</volume>
          .147
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Buncic</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aguado Sanchez</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomer</surname>
          </string-name>
          .,
          <string-name>
            <surname>Franco</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harutyunian</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mato</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            <given-names>Y</given-names>
          </string-name>
          . CernVM
          <article-title>- a virtual software appliance for LHC applications //</article-title>
          <source>Journal of Physics: Conference Series</source>
          <volume>219</volume>
          (
          <year>2010</year>
          )
          <fpage>042003</fpage>
          - DOI: 10.1088/
          <fpage>1742</fpage>
          -6596/219/4/042003
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Borthakur</surname>
            <given-names>D.</given-names>
          </string-name>
          <article-title>Hadoop Distributed File System Architecture https</article-title>
          ://hadoop.apache.
          <source>org/docs/r1.2</source>
          .1/hdfs_design.html.
          <source>(accessed 08.10</source>
          .
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>IBM</given-names>
            <surname>Pittsburg</surname>
          </string-name>
          <article-title>Labs</article-title>
          . OpenAFS User http://docs.openafs.org/UserGuide/index.html.
          <source>(accessed 08.10</source>
          .
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Chacon</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Straub</surname>
            <given-names>B. Pro</given-names>
          </string-name>
          <string-name>
            <surname>Git</surname>
          </string-name>
          . Available at: https://git-scm.com/book/en/v2
          <source>(accessed 08.10</source>
          .
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>