<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Dubna, Russia, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>ON THE WAY FROM VIRTUAL COMPUTING TO VIRTUAL DATA PROCESSING</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A. Bogdanov</string-name>
          <email>a.v.bogdanov@spbu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Degtyarev</string-name>
          <email>ba.degtyarev@spbu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>N. Shchegoleva</string-name>
          <email>cn.shchegoleva@spbu.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>V. Khvatov</string-name>
          <email>dvalery.khvatov@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>DGT Technologies AG.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>http://dgt.world/</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Alexander Bogdanov</institution>
          ,
          <addr-line>Alexander Degtyarev, Nadezhda Shchegoleva, Valery Khvatov</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>RUE, SL CTABD</institution>
          ,
          <addr-line>36, Stremyanny lane, Moscow, Russia, 117997</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>St. Petersburg University</institution>
          ,
          <addr-line>7-9, Universitetskaya emb., St. Petersburg, Russia, 199034</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>16</volume>
      <issue>2020</issue>
      <fpage>25</fpage>
      <lpage>30</lpage>
      <abstract>
        <p>Concept of a virtual personal supercomputer is proposed to solve the problems of distributed data processing. It boils down to data virtualization and a two-level processing system that allows data to be processed at their location. On this basis, a new computing platform DGT has been created, which can serve as the basis for the Big Data ecosystem.</p>
      </abstract>
      <kwd-group>
        <kwd>Big Data</kwd>
        <kwd>virtual private supercomputer</kwd>
        <kwd>data virtualization</kwd>
        <kwd>kappa paradigm for data processing</kwd>
        <kwd>DGT platform</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Efficient processing of large amounts of data has come to the fore in recent years. The main
problem of distributed data processing is described by the CAP theorem (Brewer's theorem) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It says
that simultaneously when working with distributed data, only two of the three requirements can be
satisfied: Consistency (data do not contradict each other in all computational nodes at the same time),
Availability (any request to a distributed system finalizes with a correct response), Partition tolerance (the
system continues to operate despite an arbitrary number of messages being dropped by the network
between nodes). It is also possible to reformulate and extend the CAP theorem. It is called the PACELC
theorem, according to which, in the case of network partition (P) in a distributed computer system, it is
necessary to choose between availability (A) and consistency (C), but in any case, even if the system
works normally in no separation, you have to choose between latency (L) and consistency (C).
      </p>
      <p>The development of information technologies in the direction of finding a solution to this
problem led to the emergence of the Data Lakes concept. This concept is that storage can be of different
types, including portals, archives, storefronts, databases of different kinds, data clouds, and networks.
These stores can have synchronous or asynchronous computer connections. Since the data type is often
not known in advance, there is a need for a highly flexible storage system that allows you to easily switch
between different sources and systems. However, moving away from some problems, the user ultimately
encounters other, equally complex problems. First of all, with a low speed of work with data. However,
no other way of working in a distributed environment has yet been proposed.</p>
      <p>The intensive development of information technology determines the annual increase in data
processing capabilities. At the same time, users of large computing and data centers are faced with the
fact that the architecture with which they have to deal remains old. Permanent replacement of equipment
requires both high financial resources and, equally important, very high personnel qualifications. A
disaster occurs when new data types arise and the existing architecture is completely unattended to these
challenges. However, you have to somehow use this user-pressured architecture to solve your problems.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Virtual private supercomputer</title>
      <p>
        How to solve this problem in the current conditions? A similar question was raised by the authors
in the field of computing more than ten years ago [
        <xref ref-type="bibr" rid="ref2 ref3">2,3</xref>
        ]. At that time, it seemed that there was an
unsolvable problem that more and more computer power was required to solve complex problems, but
they could be achieved in those conditions only by building a large cluster, distributed or hybrid system.
All these solutions led to high network losses, which made it impossible to use a large number of nodes to
solve almost any real problem in which there is at least a small interaction between processors. The way
out was found in the concept of a personal virtual supercomputer [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], when all possible computer
elements are virtualized: processors, memory, network, address space, etc. This approach did not allow,
of course, to create a large universal supercomputer similar to the NEX SX, but made it possible to
organize a high-performance virtual SMP system. Such system can solve any single complex computing
problem basing on the standard computing equipment available to the researcher. Such approach takes
possibility also to develop a methodology for restructuring the system to a specific task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>The idea of virtual private supercomputer allows you to circumvent the problem of dynamic load
balancing, since you can work statically in a virtual environment. In addition, with this approach, it is
possible to migrate the task to the data, and not vice versa, as is only possible with the MPP computing
paradigm. Therefore, one possible solution for today's distributed data problem is to add a data
virtualization element. Thus, combining the concept of a virtual private supercomputer with a
classification of Big Data, taking into account various storage schemes, would solve this problem.</p>
    </sec>
    <sec id="sec-3">
      <title>3. New paradigm of data processing</title>
      <p>Let us consider how this approach can be applied to working with data. To do this, first of all, you
need to understand that now researchers are forced to abandon the previously proposed paradigm of
working with data. The old -paradigm described monolithic systems based on the old architecture, when
different data types could only be processed separately using their own instrumentation for each type, and
then summarize all the results into a single source. The recent emergence of a large number of new data
types has led to the transition to a new -architecture, when data processing is sent to an external level,
and work with tools, and not with data. In practice, this allows you to send business analytics to the
periphery of the system and receive only the result of data processing.</p>
      <p>The proposed new paradigm is that we do not work with domains where data is located, but
directly with data. What should actually be done for this?
1. You need to locate the data, you want to process
2. Process data where it is
3. Insert them into the required infrastructure</p>
      <p>At the same time, it is not necessary to collect everything in a single center, as it is implemented
in the classic example – a distributed database. It is not possible to perform a transaction without
changing the system as a whole.</p>
      <p>The idea is not to work with the system "as a whole," but to be able to change something
somewhere locally and get only the result. That is, abandoning the Big Data concept, we allocate only
small parts of the data that are required to solve the task for processing. They can be isolated and
combined only on the basis of virtualization technology.</p>
      <p>
        This paradigm moves from a "data lakes" to a mesh data network. In fact, the requirements for
this network are dictated by the need to solve the following problems: when you virtualize a system, each
operation collects only those parts of the data that are needed to perform the operation and eliminate the
overwhelming amount of data that is not involved in this particular operation. This is similar to a virtual
private supercomputer in computing [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        This naturally leads us to the following data virtualization paradigm [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (see Figure 1):
- the layer where the data are produced;
- the layer where the data are processed;
- the layer where the data are utilized.
      </p>
      <p>
        In the framework of this approach, the user on the one hand can work only with a part of the data
belonging to one huge array, and on the other hand he can "bypass" the requirements of the CAP (C
consistency, A - availability, P - partition tolerance) theorem, providing two (CA, AP or CP) of the three
properties of the theorem for different data sets. This allows you to work with the selected part of the data
as local data, and to apply different combinations of tools for its processing, which best correspond to
their type (CA, AP or CP) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Since data processing is done in according with the Data Lakes concept on
the same server where they are located, it is not necessary to develop a data isolation technique, since
they are not stored in the DBMS.
      </p>
      <p>
        This approach is in line with the modern concept of Data marketplaces developed by leading
corporations such as Amazon, Intel, etc. The structure of the Data marketplace [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is shown in Figure.2.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Decentralized storage platforms</title>
      <p>
        A recent example of such a decentralized storage platform is the DGT Network [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which forms
a virtual data network, connecting different data sources across the boundaries of corporate information
into a single analytic system accessed by authorized users in a way that ensures differentiated
confidentiality. This platform is developed in accordance with Byzantine tolerance protocol [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and this is
the first in the world, and so far the only protocol supported by proven mathematical theorems.
      </p>
      <p>DGT Network provides horizontal integration by creating separate clusters of enterprise-managed
nodes that communicate through the secure F-BFT protocol and write them to a single ledger (Direct
Acyclic Graph). Although this ledger serves as a "single source of truth" for its participants, differential
anonymity protects the corporate confidentiality of the source data, while allowing analysts to provide
relevant information to participants in real time.</p>
      <p>In some ways, the competitor to the DGT Network is Data Marketplace, launched by the IOTA
Foundation in 2017. IOTA has launched an open source distributed ledger that connects IoT devices to
microtransactions processing in exchange for cryptocurrency. Unlike the blockchain, the distributed
IOTA - IOTA Tangle ledger does not group transactions into blocks, as a typical blockchain. Instead of
this it considers them as a stream of separate transactions connected together using a relatively simple
network algorithm. For participation, a node must perform a small amount of computational work to
verify the two previous transactions.</p>
      <p>Since the Internet of Things does not require any consensus, the DGT Network has a significant
advantage, allowing you to implement projects of different levels of complexity from cryptocurrencies to
Internet banks.</p>
      <p>Why is this so important right now? When working with Big Data, the central issue is data
quality. If the user works in the field where the CAP theorem is true, then he must sacrifice something.
When working with distributed data, data quality is a key issue.</p>
      <p>Two critical trends guide the direction in which Big Data quality structures are growing: data
decentralization and virtualization. The first trend illustrates the need to adapt distributed ledger
technologies for data quality control, and the second shows the need to abandon verification according to
a given data structure (since it can vary).</p>
      <p>A two-tiered approach to data processing is therefore proposed:
• Pre-processing of incoming data with identification of main information objects and verification
of their attributes;
• Handling quality attributes for all available data based on differences in transactional information
versions.</p>
      <p>This is a solution to the problem, since the data analysis task is much simpler. It means that it
does not require large loads and more local. Division into two stages allows you to remove a large amount
of unnecessary data from the process, which makes it possible to significantly reduce time costs and
ensure a high level of parallelization. The principle in this approach is that the quality of the data does not
suffer.</p>
    </sec>
    <sec id="sec-5">
      <title>5. DGT Quality Framework</title>
      <p>
        The basic strategy of DGT Quality Framework is based on allocating Master Data processing to a
separate data processing type for a distributed environment [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In fact, master data refer to all static
information that is used to identify critical elements of an organization and its business processes.
Assigning inbound operational information to objects requires identification, as does generating
consistent data sets for analysis. Therefore, master data, transactional data, and analytical data are
interdependent and part of the same context. Errors and discrepancies in master data can cause the same
or even greater damage than differences in transaction data.
      </p>
      <p>Master data support the consistency of a common information array between different
information systems, departments, and organizations. The most important characteristic of master data is
the slow rate of change in information exchange between several participants. When working with master
data, you can select the following management styles:
 on the basis of transactions;
 centralized master data;
 common master data.</p>
      <p>The following information exchange characteristics shall be taken into account
 limitations of centralized solutions
 data access in real time
 smart data processing
 storage of logs</p>
      <p>As part of the DGT Quality Framework approach, these challenges are solved through innovative
technologies that provide fast decision-making and reduce data mismatch loss:</p>
      <p> The integration layer of the system is built on a high-performance DGT core, which ensures the
formation of a single Master Data leger and its distribution among participants in information exchange.</p>
      <p> “Smart” modules that track data in real time and participate in the creation of consistent datasets
while measuring quality measures.</p>
      <p> A developed API that can connect not only to various enterprise systems and analytical tools, but
also to various data management and profiling tools.</p>
      <p>
        Base architecture of the framework [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is shown on Figure 3.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>The use of distributed ledger technologies to support basic data between organizations will
provide a single information space for groups of companies that are integrated horizontally or vertically.
This technology enables real-time quality indicators to be calculated and information exchanged
effectively in operational data, to improve the quality of analytical data and, ultimately, to make the
decision-making process qualitative.</p>
      <p>Data virtualization is a method of organizing of access to data without requiring knowledge of its
structure or location in a particular information system. This makes it possible to achieve the main goal
to simplify access to and use of data by turning it into a service, thus significantly shifting the paradigm
from storage to use. This is provided by the proposed virtualization concept, which supports the
scalability and operational efficiency required for Big Data environments through the implementation of:
- portioning, that is sharing resources and transitioning to streaming data;
- insulation, which is an object-oriented approach to data taking into account the application of the
subject area;
- encapsulation that allows you to save logical storage as a separate object.</p>
      <p>All this ensures differentiated data security and confidentiality. Therefore, the proposed data
virtualization is more than just a modern approach, it is a completely new view of the data and how to
work with it.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Eric</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Brewer</surname>
          </string-name>
          .
          <article-title>Towards robust distributed systems</article-title>
          .
          <source>In Proceedings of the nineteenth annual ACM symposium on Principles of distributed computing (PODC '00)</source>
          . ACM, New York, NY, USA.
          <year>2000</year>
          . DOI:
          <volume>10</volume>
          .1145/343477.343502
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Gankevich</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , et al,
          <source>Constructing Virtual Private Supercomputer Using Virtualization and Cloud Technologies //Lecture Notes in Computer Science</source>
          ,
          <volume>8584</volume>
          ,
          <fpage>341</fpage>
          -
          <lpage>354</lpage>
          (
          <year>2014</year>
          ) DOI:
          <fpage>10</fpage>
          .1007/978-3-
          <fpage>319</fpage>
          - 09153-2_
          <fpage>26</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] Virtual Supercomputer as basis of Scientific Computing / A</article-title>
          . Bogdanov,
          <string-name>
            <given-names>A.</given-names>
            <surname>Degtyarev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Korkhov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gaiduchok</surname>
          </string-name>
          , I. Gankevich // Horizons in Computer Science Research - New York: Nova Science Publishers, Inc.,
          <year>2015</year>
          . - Vol.
          <volume>11</volume>
          , -
          <fpage>203р</fpage>
          .,
          <fpage>159</fpage>
          -
          <lpage>198</lpage>
          p.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Bogdanov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Degtyarev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Korkhov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Desktop supercomputer: what can it</article-title>
          do? // Phys. Part.
          <source>Nuclei Lett</source>
          .
          <volume>14</volume>
          (
          <issue>7</issue>
          ),
          <fpage>985</fpage>
          -
          <lpage>992</lpage>
          (
          <year>2017</year>
          ) DOI:
          <fpage>10</fpage>
          .1134/S1547477117070032
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Bogdanov</surname>
            <given-names>A.</given-names>
          </string-name>
          , et al,
          <source>Evolving Principles of Big Data Virtualization. //Lecture Notes in Computer Science</source>
          ,
          <volume>12254</volume>
          ,
          <fpage>67</fpage>
          -
          <lpage>81</lpage>
          (
          <year>2020</year>
          ) DOI:
          <fpage>10</fpage>
          .1007/978-3-
          <fpage>030</fpage>
          -58817-
          <issue>5</issue>
          _
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Bogdanov</surname>
            ,
            <given-names>A. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shchegoleva</surname>
            ,
            <given-names>N. L.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Ulitina</surname>
            ,
            <given-names>I. V.</given-names>
          </string-name>
          ,
          <year>2019</year>
          ,
          <article-title>Database ecosystem is the way to data lakes</article-title>
          <source>Proceedings of the 27th Symposium on Nuclear Electronics and Computing (NEC</source>
          <year>2019</year>
          ). Korenkov,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Strizh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Nechaevskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            &amp;
            <surname>Zaikina</surname>
          </string-name>
          ,
          <string-name>
            <surname>T</surname>
          </string-name>
          . (ed.). RWTH Aahen University, pp.
          <fpage>147</fpage>
          -
          <lpage>152</lpage>
          (CEUR Workshop Proceedings ; vol.
          <volume>2507</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>[7] DGT, the Decentralized Enterprise Platform</article-title>
          . http://dgt.world/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bogdanov</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <source>et al A DLT Based Innovative Investment Platform. //Lecture Notes in Computer Science</source>
          ,
          <volume>12251</volume>
          ,
          <fpage>72</fpage>
          -
          <lpage>86</lpage>
          (
          <year>2020</year>
          ) DOI:
          <fpage>10</fpage>
          .1007/978-3-
          <fpage>030</fpage>
          -58808-
          <issue>3</issue>
          _
          <fpage>7</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Bogdanov</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <source>et al Data Quality in a Decentralized Environment. //Lecture Notes in Computer Science</source>
          ,
          <volume>12251</volume>
          ,
          <fpage>58</fpage>
          -
          <lpage>71</lpage>
          (
          <year>2020</year>
          ) DOI:
          <fpage>978</fpage>
          -3-
          <fpage>030</fpage>
          -58808-
          <issue>3</issue>
          _
          <fpage>6</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>