<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ALGORITHM FOR SOLVING PROBLEM SYNTHESIS THE OPTIMAL LOGICAL STRUCTURE DISTRIBUTED DATA IN ARCHITECTURE OF GRID SERVICE</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nurmatova E.V.</string-name>
          <email>nurmatova@mirea.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gusev V.V.</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Elena Nurmatova, Victor Gusev</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>MIREA - Russian Technological University</institution>
          ,
          <addr-line>20, Stromynka, Moscow, 107996</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Research Center “Kurchatov Institute” - Institute for High Energy Physics</institution>
          ,
          <addr-line>1, Science sq, Protvino, 142281</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>5</fpage>
      <lpage>9</lpage>
      <abstract>
        <p>The questions of constructing optimal logical structure of a distributed database (DDB) are considered. Solving these issues will make it possible to increase the speed of processing requests in DDB in comparison with a traditional database. In particular, such tasks arise for the organization of systems for processing huge amounts of information from the Large Hadron Collider. In these systems various DDB are used to store information about: the system of triggers of data collection from physical experimental installations, the geometry and the operating conditions of the detector while collecti ng experimental data. Two interrelated stages in the synthesis algorithm are proposed. At the first stage, the problem of distribution of database clusters between the server and clients, followed by the problem of optimal distribution of data groups of each node by types of logical records are addressed. At the second stage the problem of database localization on the nodes of the computer network is solved, in addition to the results of the first stage, the characteristics of the DDB are taken into account. Optimal logical structure of DDB will ensure the efficiency of the information system on computational resources. As a result of its solution, the local network of the DDB is decomposed into a number of clusters that have minimal information connectivity with each other. Solving the problem of synthesis of the optimal logical structure is also of great practical importance for the automated design of logical structures, for the automated formation of query specifications and adjustments of the DDB.</p>
      </abstract>
      <kwd-group>
        <kwd>data warehouse</kwd>
        <kwd>optimal logical data structure</kwd>
        <kwd>applications</kwd>
        <kwd>large data volume</kwd>
        <kwd>synthesis algorithm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Previous work</title>
      <p>For a grid architecture with a large number of requests, users, and large amounts of data, it is
advisable to formulate the problem of synthesizing the optimal logical structure of a DDB based on
the criterion of the minimum total time for implementing a set of user requests. Indeed, along with the
necessary information that is really required by the user, redundant information, which arises as a
result of the localization of information elements that are not required by the user in one record, is
transmitted from the database server. Excessive information "clogs up" the communication channels,
which in the future will require an increase in network bandwidth due to the power of hardware and
software.</p>
      <p>The variety of options for alternative solutions for the choice of not fully defined evaluation
criteria, is a rather weak side of the problem of determining the optimality of the developed logical
data structure in a distributed architecture.</p>
      <p>
        When working with quantitative criteria, which include request response time, update cost,
memory cost, time to create, reorganization cost, the contradiction of the criteria to each other can
cause the difficulty [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>There are optimality criteria, which are immeasurable properties, poorly represented in
quantitative terms or in the form of an objective function. The qualitative criteria for evaluating a DB
include flexibility, adaptability, availability for new users, compatibility with other systems, the ability
to convert for use on another computing platform, the possibility of recovery, the possibility of
fragmentation and expansion of the structure.</p>
      <p>It is expedient to consider the synthesis of the logical structure of the DDB as a sequential
solution of three particular problems, the results of which are the determination of:
1) optimal localization of data groups, providing a minimum of total traffic in the system and
satisfying the specified restrictions;
2) the structure of the optimal distribution of groups by types of logical records, providing a
minimum of the total time of local processing of information on the servers of network nodes
under the given restrictions;
3) the structure of the database localization by the system nodes, providing the minimum value
of the total time of access to the localization and DB processing nodes.</p>
      <p>The solution of the problem of data processing in real time requires a special organization of
the logical and physical structure of the DDB to ensure the response time of the system on the order of
1-2 seconds or less (the minimum implementation time for operational requests).
2. Stages of an approximate algorithm for synthesizing the optimal logical
structure of the DDB</p>
      <p>Consider an approximate algorithm for solving the problem of synthesizing the optimal logical
structure of the DDB and the structure of localization of the database according to the criterion of the
minimum total time for the implementation of a set of user requests , consisting of a sequence of
stages (figure 1).</p>
      <p>Stage 1. At this stage, the localization of data groups in the computing system is determined
by the criterion of the minimum total traffic. To solve this problem, an approximate algorithm for
distributing DDB clusters between the server and clients of the local network is used. At the first step
of the stage, the graph of the canonical structure of the DDB is reduced to a disconnected graph with
the calculation of the "weight" of each data group. The weight of each group consists of the weight of
the data group itself and the weight of the arcs, taking into account the requirements of users:
the canonical structure of the DDB.</p>
      <p>where,   гр  the total weight of the data group;   св′ the weight of the arcs of the graph of
 0  0
 =1  =1

 гр = ∑
∑ 

з з </p>
      <p>0  0
 =1  =1
  св′ = ∑
∑ 
з з 

∑   ′ 
Г
  ′
 0  0
  = ∑
∑ </p>
      <p>=1  =1
з з 
 (1 + ∑   ′</p>
      <p>Г ′)

 ′</p>
      <p>′


  =   + ∑    ′
 0
 ′
Then the weight of the i-th group is
where, з  the frequency of usage of queries by users; з  the elements of the matrix for
using queries by DDB users; 
the semantic contiguity matrix of data groups.</p>
      <p> the matrix for using data groups when executing queries;  Г ′ </p>
      <p>At the second step of the stage, the computer network graph is transformed to a disconnected
graph with the calculation of the "weight" of each node:
time of decomposition of the query into subqueries, route selection and connection establishment, etc.;
where</p>
      <p> the total average duration of data processing in the r-th node, consisting of the
matrix of logical distances between the servers of the nodes of the computer network.</p>
      <p>′  the average duration of data transmission between nodes, determined based on the
  =   ×   for  = ̅1̅,̅̅;  = 1̅̅,̅̅̅0̅.</p>
      <p>At the third step of the first stage, the matrix  = ‖  ‖ is formed, whose elements are equal:
At the fourth step of the first stage, the problem
is solved under the constraints:</p>
      <p>0
min∑
{  }  =1  =1</p>
      <p>∑    
∑     ,  = 1̅̅, ̅̅0̅</p>
      <p>=1
 0
 =1
∑     ,  = ̅1̅,̅̅</p>
      <p>=1
∑     взу



by the number of data groups, the localization of which is possible on one node
on the admissible duplication of groups by network nodes  xir  M i ,
r0
r1
on the amount of available external memory of the network servers for storing data
where, 
host;  
otherwise.</p>
      <p> the vector of group lengths in bytes;   the vector of number of
instances in groups; взу  the amount of available memory on the server of the  -th
= 1, if the  -th data group is included in the r-th network node;   = 0 –</p>
      <p>This is a linear integer programming problem. Its solution makes it possible to determine the
optimal localization of data groups by network nodes.</p>
      <p>Stage 2. At this stage, the problems of optimal distribution of data groups of each node
according to the types of logical records are solved by the criterion of the minimum total time of local
data processing in each network node. The number of synthesis tasks for this stage is determined by
the number of network nodes.</p>
      <p>
        The initial data are subgraphs of the graph of the canonical structure of the DDB, as well as
the temporal and volume characteristics of the subgraphs of the canonical structure of the DDB, the set
of requests from users and network nodes [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        The synthesis problem for this stage is solved using exact or approximate algorithms [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The
following restrictions are used: restrictions on the number of groups in a logical record, on the
onetime inclusion of groups in records, on the cost of storing information, on the required level of
information security of the system, on the time of performing operational transactions on DB servers,
on the total time of servicing operational requests on servers. As a result, we determine the logical
structures of the DB for each node in the network.
      </p>
      <p>Stage 3. Localization of the DB by network nodes. Initial data of the stage: results of the
previous stages and characteristics of the DDB. Restrictions: on the total number of synthesized
logical records located on the server of the r-th node of the computer network; on the amount of
available external memory of the network servers for storing the database; on the number of copies of
logical records placed on the network.</p>
      <p>As a result of the proposed algorithm (figure 2), localization matrices of the set of data groups
by the types of logical records are formed (result of stage 1) and then groups of records by network
nodes (result of stage 2, table 1) are formed. The timing of the algorithms is also evaluated.</p>
    </sec>
    <sec id="sec-2">
      <title>3. Similar solutions</title>
      <p>
        As alternative solutions for comparing the results of synthesizing the data structure according
to various criteria, we analyzed the analogue that solves the NP-hard nonlinear integer discrete
optimization problem from the DDB domain [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and implements 3 neural network algorithms for
synthesizing the optimal logical structure DDB according to the criterion of the minimum total time of
sequential processing of a set of user requests:
      </p>
      <p>Further work within the framework of this topic is the development of software that
implements the search algorithm for a variant of the logical structure of the DDB, which ensures the
optimal value of the specified criterion for the efficiency of the functioning of the grid system and
satisfies the main system, network, and structural constraints.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Kosterin</surname>
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Minin</given-names>
            <surname>Yu</surname>
          </string-name>
          .V.,
          <string-name>
            <surname>Ivanova</surname>
            <given-names>O.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Matari</surname>
            <given-names>N.A.</given-names>
          </string-name>
          <string-name>
            <surname>Kh</surname>
          </string-name>
          .
          <article-title>Statement of the problem of synthesizing the optimal logical structure of a network database in fuzzy conditions</article-title>
          .
          <article-title>- Information and security</article-title>
          , Voronezh,
          <year>2014</year>
          . - 574 - 579 p.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Nurmatova</surname>
            <given-names>E.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gusev</surname>
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kotliar</surname>
            <given-names>V.V.</given-names>
          </string-name>
          <article-title>Analysis of the features of the optimal logical structure of distributed databases// Collection of works the 8th International Conference “Distributed Computing and Grid-technologies in Science and Education”</article-title>
          .
          <source>- Dubna</source>
          ,
          <year>2018</year>
          .- 167 p.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Amaru</surname>
            <given-names>L.G.</given-names>
          </string-name>
          <article-title>New Data Structures and Algorithms for Logic Synthesis</article-title>
          and Verification.- Springer,
          <year>2016</year>
          . - 262 p. -
          <source>ISBN: 9783319431734</source>
          , EISBN:
          <fpage>9783319431741</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>