<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>THE VISUALIZATION METHOD PIPELINE FOR THE APPLICATION TO DYNAMIC DATA ANALYSIS</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>T. Galkin</string-name>
          <email>tpgalkin@mephi.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Popov</string-name>
          <email>dmitry.popov@skoltech.ru</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V. Pilyugin</string-name>
          <email>vvpilyugin@mephi.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Grigorieva</string-name>
          <email>Maria.Grigorieva@cern.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Research Nuclear University MEPhI</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Skolkovo Institute of Science and Technology</institution>
          ,
          <addr-line>Moscow</addr-line>
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Timofei Galkin</institution>
          ,
          <addr-line>Dmitry Popov, Victor Pilyugin, Maria Grigorieva</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>4</volume>
      <issue>2019</issue>
      <fpage>295</fpage>
      <lpage>299</lpage>
      <abstract>
        <p>The new era of scientific research brings an enormous amount of data for scientists. These complex and multidimensional data structures are used for the verification of scientific hypothesis. Exploring such data by researchers requires the development of new technologies for its efficient processing, investigation and interpretation. Intellectual data analysis and statistical methods are rapidly developing, and this is where visualization methods are getting their place. This work describes mathematical basis of the developed visualization tool for the analysis of multidimensional dynamic data. This tool provides the pipeline of methods, which combined, allow to cope with a set of practical tasks (anomalies detection, cluster, trends and variation analysis) using visualization method. Authors provided mathematical models of geometrical operations under the data domain, algorithms for solving the mentioned classes of tasks and several use-cases with technological and economic data based on visualization method.</p>
      </abstract>
      <kwd-group>
        <kwd>visual analysis</kwd>
        <kwd>dynamic data</kwd>
        <kwd>time-variant data</kwd>
        <kwd>multidimensional data</kwd>
        <kwd>multivariate data</kwd>
        <kwd>visualization</kwd>
        <kwd>data analysis</kwd>
        <kwd>multidimensional analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Proceedings of the 27th International Symposium Nuclear Electronics and Computing (NEC’2019)</p>
      <p>Budva, Becici, Montenegro, September 30 – October 4, 2019</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Intelligent computer algorithms are state-of-the-art of data analysis today.</p>
      <sec id="sec-2-1">
        <title>Artificial</title>
        <p>intelligence, machine learning and neural networks create the trend of discourse in the data science.
However, at the same time some research point out the problem of understanding, interpretation and
verification of the research results [1]. Various methods are available for these purposes. One of them
is data visualization. Industry and science bring us tasks which include complex multidimensional data
analysis, and data visualization can provide deep understanding of data based on its graphical
representation. This paper describes the experience of applying the visualization method for
multidimensional dynamic data analysis.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Background</title>
      <p>The overview of visualization techniques for time-dependent multidimensional data can be
found in [2]. The authors divide these techniques into static and dynamic. Moreover, they considered
interactions</p>
      <p>with the visual representations. The recent overview [3] considers the different
visualization techniques and data transformations. Data model for the visualization can be represented
as the multidimensional Euclidian space and affine transformations within this space [4].</p>
      <p>Various data structures from different research fields are successfully investigated by imaging,
which provides analyst with the advanced and interactive means for data exploration. This paper
describes the visualization pipeline, developed by the authors, based on 3D scatter plot diagram with
colored distances between data objects in multidimensional space.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Visualization Method for the Dynamic Data</title>
      <p>Dynamic multidimensional data is represented as a set of parameters of objects, changing in
time. This data is stored as a set of numeric data values given for some periods of time.</p>
      <sec id="sec-4-1">
        <title>3.1 Task formulation</title>
        <p>Thus, the following task formulation is to solve:</p>
        <p>Let m objects given, each of them is characterized by n parameters. The data is organized as a
set of tables, such as follow:</p>
        <sec id="sec-4-1-1">
          <title>Parameter 1</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Parameter 2</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Given:</title>
        <p>Time = j
Object 1</p>
        <sec id="sec-4-2-1">
          <title>Object 2 … Object m</title>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Task:</title>
        <p>This tabular data representation contains object parameter values at a specific point in time.
of parameter l of object i in table j at the moment of time tj (l = (1,n), i = (1,m) , j = (1,k)).
Thus, table j is filled with parameter values for time tj. It is stated that t1 &lt; t2 &lt; … &lt; tk, and    – value
To make the formulations easier,    for l = (1,n) is called a n-tuple within the fixed j.</p>
        <p>Find the subsets of similar objects, explore these subsets at each given point in time and make
judgements about their behavior in time.</p>
        <p>11
 21
…

 1</p>
      </sec>
      <sec id="sec-4-4">
        <title>3.2 Data samples</title>
        <p>data.</p>
        <p>In this research two data samples with dynamic data were used: technological and financial
Technological data sample were taken from Kaggle's dataset "CareerCon 2019 - Help
Navigate Robots"1. It represents sensors data gathered while driving a small mobile robot over
different floor surfaces: orientation, velocity, acceleration, etc. These data may help robots to
recognize the floor surface. The dataset has ~4K objects and 128 time measurements. The total length
of dataset is about 400K records. The visualization pipeline was applied to these data to visualize
sensors data in order to explore the surface features.</p>
        <p>Financial data sample represents data of banking system. This data was obtained from the
open sources2 and describes 81 banks for a period of 13 month by a set features like sales profit,
deposits, overdue debt and others. The main idea of the exploration of these data is to uncover
suspicious, anomalous banks visually, and detect a point in time or a period when the anomalous
behavior takes place.</p>
      </sec>
      <sec id="sec-4-5">
        <title>3.3 Task solving method</title>
        <p>For solving the data analysis task, the scientific visualization method was used [5].</p>
        <p>Both data samples were represented as a set of dynamic objects. Each object with all the
corresponding features at each point in time is stored in
a table. A set of tables for all objects at different points
in time form a preprocessed dataset for loading into the
visualization application.</p>
        <p>The visual analysis of data has two stages. The
first stage is the visualization itself: data tables are
transformed into geometrical objects on screen. This
transformation implies four steps: sourcing (obtaining
the data from the source), filtering (getting the data
ready for the application), mapping (corresponding
geometric objects placement on the scene) and rendering
(making the resulting picture of the scene). After the
visualization is ready, the second stage of the analysis
the interpretation of images is performed by the analyst. Figure 1. The visualization pipeline
Parameters of each step of the visualization pipeline
may be changed by the analyst interactively in order to
generate another visualization sample. This makes the
process of data analysis iterative and interactive.</p>
      </sec>
      <sec id="sec-4-6">
        <title>3.3.1 Visualization pipeline</title>
        <p>Figure 1 shows the transformation between the
data table and the visualization. Each line in the table
represents the data object. Features of objects are
transformed into multidimensional coordinates. The
objects are then projected into spheres.</p>
        <p>The application backend calculates distances
between all pairs of objects in multidimensional space,
and display it as segments between corresponding pairs</p>
        <sec id="sec-4-6-1">
          <title>1 https://www.kaggle.com/c/career-con-2019/data 2 https://www.banki.ru/</title>
          <p>of spheres. Segments are colored from blue to red,
depending on how close the objects are in the original
multidimensional space. This allows to observe
similarity of objects in multidimensional space looking
at the 3D visualization.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Implementation and Application</title>
      <sec id="sec-5-1">
        <title>The algorithms of the visualization method pipeline were implemented in C# programming language, using Unity graphics engine.</title>
        <sec id="sec-5-1-1">
          <title>4.1 Technological data sample visualization</title>
          <p>At the figure 2, X Y and Z values are the
orientation parameter, in degrees. The spheres in the
picture form a circle. That is the key point of visual
analysis. Human experts are good at interpretation of
graphical images, which is hard to be programmed
automatically. Moving the time slider allows to observe
changes of parameters values in time and can be useful
in the detection of specific points of time when
anomalous behavior takes place.</p>
          <p>One more thing that can be visually discovered
is presented in the figure 3. An analysist found a cluster
of spheres, and this cluster does not change in time.
They are marked with red color at the picture. Further
investigation showed that these spheres correspond to
the specific type of floor surfaces. Therefore, such
visualization is useful for the problem of clusterization.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>4.2 Economic data sample visualization</title>
          <p>The figure 4 shows that the spheres lay on a
plane, which noticeably rotates over the time. The
analyst may observe visually the direction and the
velocity of movement of some financial parameters,
making conclusions about common financial situation
for banks. Also, such visualization allows to catch
anomalous banks, which features are changing in time
along other trajectories.</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>4.3 Other applications</title>
          <p>This application was also tested on metadata
from ATLAS Grid Information System3 as shown on
figure 5. The visualization shows the appearing of
computing queues, and the duration of these queues in
time, and can be used to observe some specific
tendencies, which may be unobvious without graphic
representation.</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>3 http://atlas-agis.cern.ch/agis/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>