<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Aggregation in the Astroparticle Physics Distributed Data Storage?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Minh-Duc Nguyen</string-name>
          <email>nguyendmitri@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Kryukov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julia Dubenskaya</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Korosteleva</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Bychkov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey Mikhailov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexey Shigarov</string-name>
          <email>shigarov@icc.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University, Skobeltsyn Institute of Nuclear Physics</institution>
          ,
          <addr-line>1/2 Leninskie Gory, 119991, Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Matrosov Institute for System Dynamics and Control Theory, Siberian Branch of Russian Academy of Sciences</institution>
          ,
          <addr-line>Lermontov st. 134, Irkutsk</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>German-Russian Astroparticle Data Life Cycle Initiative is an international project whose aim is to develop a distributed data storage system that aggregates data from the storage systems of di erent astroparticle experiments. The prototype of such a system, which is called the Astroparticle Physics Distributed Data Storage (APPDS), has been being developed. In this paper, the Data Aggregation Service, one of the core services of APDDS, is presented. The Data Aggregation Service connects all distributed services of APPDS together to nd the necessary data and deliver them to users on demand.</p>
      </abstract>
      <kwd-group>
        <kwd>Distributed storage Data aggregation Data lake</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The amount of data being generated by various astroparticle experiments such
as TAIGA [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], KASCADE-Grande [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], MAGIC [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], CTA [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], VERITAS [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and
HESS [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is tremendous. Processing data and, more importantly, delivering data
from di erent experiments to end-users are one of the real issues in open science
and particularly in open access to data. To address this issue, the
GermanRussian Astroparticle Data Life Cycle Initiative had been started in 2018. The
primary goal of the project is to develop a prototype of a distributed storage
system where data of two physical experiments, TAIGA and KASCADE-Grande,
are aggregated in one place and to provide a uni ed access mechanism to
endusers.
      </p>
      <p>Such a prototype has been being developed and is called Astroparticle Physics
Distributed Date Storage (APPDS). The design of the system is based on two
key principles. The rst one is creating no interference with existing data
storage systems of the physical experiments. This principle is critical since most
existing collaborations in astroparticle physics, not only the collaborations of
TAIGA and KASCADE-Grande experiments, have been historically using their
own stack of technologies and approach to data storage. Any change to the
existing data storage systems might cause potential problems which are hard to
be discovered. The second principle is using no computing resources at the sites
of the storage systems to handle users' queries. This principle leads to creating
a global metadata database where the data description from all storage systems
is aggregated in one place. All searching and ltering operations are performed
within the global metadata database. The operation results are delivered to users
via a web interface. Actual data transfer from the existing data storage systems
takes place only when users want to access the inside content of a le. Thus,
APPDS causes no load to the computing resources at each site; the only load on
the data storage systems is data delivery.</p>
      <p>This article is organized as follows. In the second section, the architecture
overview of APPDS is presented. Section 3 is dedicated to the design of the Data
Aggregation Service, which is the core of APPDS, and the stack of technologies
that have been used. In conclusion, the current state of the service is presented.
2</p>
    </sec>
    <sec id="sec-2">
      <title>APPDS architecture</title>
      <p>The architecture overview of APDDS is presented in Fig.1. S1, S2, S3 are data
storage instances of the physical experiments. In1 is the original data input. In2
indicates the case when the original data from In1, which are already stored in a
storage instance, are being reprocessed. At the level of data storage instances, to
preserve the original data processing pipeline, a program called Extractor (E1) is
injected into the pipeline. In most cases, input data are les. After standard
processing, the les are passed to the Extractor. The Extractor retrieves metadata
from the les using the metadata description (MDD) provided by the
development groups of the physical experiments, sends the metadata to the Metadata
Database using its API, and passes the les back to the pipeline. If the data
need to be reprocessed, the same pipeline is applied but with a di erent type of
the Extractor (E2).</p>
      <p>
        The les of each storage instance are delivered to the Data Aggregation
Service by the Adapter, which is a wrapper of the CernVM-FS server [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>To retrieve necessary les, the user forms a query using the web interface
provided by the Data Aggregation Service. When the Data Aggregation Service
receives the user query, it asks the Metadata Database for the answer. When
the Metadata Database answers, the Data Aggregation Service generate a
corresponding response and delivers it to the user.</p>
      <p>
        All components of APPDS are talking to each other via RESTful API [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
3rd party application services can also talk to APPDS via RESTful API. The
key business logic of APPDS is implemented in the Data Aggregation Service
whose design and implementation are considered in the next section.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Data Aggregation Service</title>
      <p>There are two types of queries that a user can make using the Web Interface of
Data Aggregation Service. Identi cation queries are used to retrieve the basic
information about the resources provided by APPDS such as facilities,
clusters, detectors, data channels, data authorities, permissions, and available data.</p>
      <sec id="sec-3-1">
        <title>HTTP Server</title>
      </sec>
      <sec id="sec-3-2">
        <title>Cache Manager</title>
      </sec>
      <sec id="sec-3-3">
        <title>Local File Buffer</title>
        <p>(CernVM-FS</p>
      </sec>
      <sec id="sec-3-4">
        <title>Repositories)</title>
        <p>ech iss
.2aC it/H</p>
        <p>M</p>
        <p>3.LookGupra/ApnhsQwLer
Search queries are used when the user wants to look for the available data
using a list of lters like data availability interval, energy range, facility location,
detector types, data channel speci cation, weather condition, etc. Typical data
lookup scenarios that users might create are:
{ data obtained by one facility or all available facilities for a certain period;
{ season data which start from September to the end of May;
{ data obtained in a testing period or a speci c run;
{ regular monthly or weekly data.</p>
        <p>
          The obvious choice to implement the dialogue between the Web Interface and
the Aggregation Service as separate microservices is to create a RESTful API.
While identi cation queries are the best candidates to be implemented using
a set of endpoints as in the standard RESTful API, the search queries with
their complex lters are not. Using the standard RESTful API to implement
the search queries leads to complex query sets containing redundant data that
users do not need. The best approach to implement search queries is to use the
Facebook GraphQL [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] instead of the standard RESTful API. By using GraphQL
all typical search queries can be formatted exibly as a JSON object that later
is sent to the Metadata Database, and the responses to the queries contain the
exact information that the user needs with no redundancy. It is also quite easy
to implement additional lters without breaking compatibility using GraphQL.
3.2
        </p>
        <p>
          Query Processing Pipeline
Whenever the Core Controller receives a query from the GraphQL Backend,
it calculates the query checksum using the MD5 algorithm. The checksum is
used as the query ID. After that, the Core Controller check against the Cache
Manager to nd if such a query is already registered. If not the Core Controller
registers the query with the Cache Manager. The Cache Manager is implemented
based on Redis [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a popular caching mechanism. The query body is converted
into a Redis hash which is a map composed of elds corresponding to the lters,
each of which contains the lter value.
        </p>
        <p>If the response to the query is already cached and the active period does
not expire, it will be delivered to the user. If the response to the query has not
been cached, the Core Controller forwards the query to the Metadata Database.
The Metadata Database generates a SQL-query according to the original query
and executes it. Depending on the data sources, the response from the Metadata
Database can vary. The response, including data from the TAIGA experiment,
is a set of les which match the lters speci ed in the original query. The
Metadata Database also calculates the checksum of the response, adds it to the nal
response and sends it to the Core Controller. The Core Controller, in turn,
appends the response to the query in the Cache Manager. After that, the Core
Controllers sends the response to the Web Interface so that the user could look
at it quickly.</p>
        <p>At the same time, the Core Controller starts preparing the full response
for delivery. Files from all Data Storage instances are exported to the Data
Aggregation Service as separate CernVM-FS repositories. All repositories are
located inside a Global Mount Point where each subdirectory is the mount point
of each repository. To prepare the data, the Core Controller nds the les in the
Global Mount Point and copies them to a directory in the Local File Bu er. The
directory is then con gured as a CernVM-FS repository for export. The name
of the repository is the same as the query ID. If the user query requires not
the original les but a composition of them, the Core Controller scans through
the original les, composes a new set of les, and puts them into the directory.
Later, if the user wants to work with the les locally, she can mount the prepared
CernVM-FS repository to her computer. The Core Controller also generates an
archive containing all les. The archive is put into the repository and can be
downloaded separately. If the full response to a query is not used more than
a speci c time, it will be deleted from the Local File Bu er and the Cache
Manager.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>Currently, a beta version of APPDS, including the Data Aggregation service, the
Metadata Database, the Adapter, and the Web Interface, has been implemented
and being tested against the storage system of the TAIGA experiment. In the
next release, support for the storage system of the KASCADE-Grande
experiment will be included. The rst working prototype of the system is planned to
be released this fall. The proof-of-concept tests and performance benchmarks
are also planned after the release.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bundev N</surname>
          </string-name>
          . et al.:
          <article-title>The TAIGA experiment: From cosmic-ray to gamma-ray astronomy in the Tunka valley</article-title>
          .
          <source>In: Nuclear Instruments and Methods in Physics Research Section A: Accelerators</source>
          , Spectrometers,
          <source>Detectors and Associated Equipment February</source>
          <year>2017</year>
          , vol.
          <volume>845</volume>
          , pp
          <fpage>330</fpage>
          -
          <lpage>333</lpage>
          . https://doi.org/10.1016/j.nima.
          <year>2016</year>
          .
          <volume>06</volume>
          .041
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Apel W.D</surname>
          </string-name>
          . et al.:
          <article-title>The KASCADE-Grande Experiment</article-title>
          .
          <source>In: Nuclear Instruments and Methods in Physics Research Section A 620 April</source>
          <year>2010</year>
          : pp
          <fpage>202</fpage>
          -
          <lpage>216</lpage>
          . https://doi.org/10.1016/j.nima.
          <year>2010</year>
          .
          <volume>03</volume>
          .147
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Anderhub H</surname>
          </string-name>
          . et al.:
          <article-title>MAGIC Collaboration</article-title>
          .
          <source>In: 31st International Cosmic Ray Conference (ICRC</source>
          <year>2009</year>
          ). https://doi.org/10.15161/oar.it/1446204371.89
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Cherenkov</given-names>
            <surname>Telescope</surname>
          </string-name>
          <article-title>Array: Exploring the Universe at the Highest Energies</article-title>
          . https://www.cta-observatory.org/
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Very</given-names>
            <surname>Energetic Radiation Imaging Telescope Array System</surname>
          </string-name>
          https://veritas.sao.arizona.edu
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>6. High Energy Stereoscopic System</article-title>
          . https://www.mpi-hd.mpg.de/hfm/HESS/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Buncic</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aguado Sanchez</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blomer</surname>
          </string-name>
          .,
          <string-name>
            <surname>Franco</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harutyunian</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mato</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            <given-names>Y</given-names>
          </string-name>
          .:
          <article-title>CernVM { a virtual software appliance for LHC applications</article-title>
          .
          <source>In: Journal of Physics: Conference Series</source>
          <volume>219</volume>
          (
          <year>2010</year>
          )
          <article-title>042003</article-title>
          . https://doi.org/10.1088/
          <fpage>1742</fpage>
          - 6596/219/4/042003
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Fielding</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Th</surname>
          </string-name>
          .:
          <article-title>Architectural Styles and the Design of Networkbased Software Architectures</article-title>
          . https://www.ics.uci.edu/ elding/pubs/dissertation/ elding dissertation.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Facebook</given-names>
            <surname>Inc</surname>
          </string-name>
          .: GraphQL. https://graphql.org
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. Redis Labs:
          <article-title>Redis documentation</article-title>
          . https://redis.io/documentation
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>