<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Concept and Evaluating of Big Data Quality in the Semantic Environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oleksandr Novytskyi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>, Institute of Software Systems of the National Academy of Sciences of Ukraine</institution>
          ,
          <addr-line>Academician Glushkov Avenue</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>40</institution>
          ,
          <addr-line>Kyiv, 03187</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Big data refers to large volumes, complex data sets with various autonomous sources, characterized by continuous growth. Data storage and data collection capabilities are now rapidly expanding in all fields of science and technology due to the rapid development of networks. Evaluating the quality of data is a difficult task in the context of big data, because the speed of semantic data reasoning directly depends on its quality. The appropriate strategies are necessary to evaluate and assess data quality according to the huge amount of data and its rapid generation. Managing a large volume of heterogeneous and distributed data requires defining and continuously updating metadata describing various aspects of data semantics and its quality, such as conformance to metadata schema, provenance, reliability, accuracy and other properties. The article examines the problem of evaluating the quality of big data in the semantic environment. The definition of big data and its semantics is given below and there is a short excursion on quality assessment. The model and its components which allow to form and specify metrics for quality have been developed. This model includes such components as: quality characteristics; quality metric; quality system; quality policy. A quality model for big data that defines the main components and requirements for data evaluation has already been proposed. In particular, such evaluation components as: accessibility, relevance, popularity, compliance with the standard, consistency, etc. are highlighted. The problem of inference complexity is demonstrated in the article. Approaches to improving fast semantic inference through materialization and division of the knowledge base into two components, which are expressed by different dialects of descriptive logic, are also considered below. The materialization of big data makes it possible to significantly speed up the processing of requests for information extraction. It is demonstrated how the quality of metadata affects materialization. The proposed model of the knowledge base allows increasing the qualitative indicators of the reasoning speed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Big data quality, semantic big data, reasoning optimization in the semantic big data</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>The concept of Big Data in the broad sense of this word is used to define data processing, spread,
and analytics [1]. The main special feature of this data is increased exponentially. Many efforts are
aimed at solving the problem of big data, this is due to the need to develop new methods and
algorithms for BD processing.</p>
      <p>Defining big data is primarily related to the difficulty of defining a quantitative definition of a set
of information objects. The most accepted definition is indicated in the report [2], where the problem</p>
      <p>
        2022 Copyright for this paper by its authors.
added to the definition of big data. This term was clarified and supplemented with criteria that
affected the complexity and unstructuredness of the data [
        <xref ref-type="bibr" rid="ref1">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">5</xref>
        ]. A number of big data definitions
came from real business problems. However, we assume that the semantics and structure are given
through external ontologies and fixed through metadata for semantic big data. We do not consider the
problem of normalization and data extraction but evaluate the quality of such data. But this does not
solve the problems of operating with such data and creates additional problems related to the
reasoning of information from such a BD set. Our semantic data model must satisfy such
requirements as Findable, Accessible, Interoperable and Reusable data or metadata [
        <xref ref-type="bibr" rid="ref3">6</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. Big Data Semantics</title>
      <p>
        The issue of semantics was studied in works [
        <xref ref-type="bibr" rid="ref4">7</xref>
        ], where big data was considered on the basis that
data semantics refers to the meaningful and effective use of a data object to represent a concept or
object in the real world. Such a general concept unites a wide variety of applications [
        <xref ref-type="bibr" rid="ref5">8</xref>
        ]. Big Data
semantic knowledge refers to numerous aspects of rules, expert knowledge and domain information
[
        <xref ref-type="bibr" rid="ref6">9</xref>
        ]. One of the specific properties of big data in the semantic environment is the increasing
complexity of reasoning even though this data not to big for the first view. Online web-application is
very sensitive for delay for response and union approach reasoning and web technology provide high
requirement to velocity big data. Our article surveys the problem of big data quality for web
application and means for increasing velocity.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Model quality of Dig Data</title>
      <p>The practical suitability of BD is determined primarily by its quality. The urgency of solving the
BD quality problem is determined by the scale of its creation and distribution.</p>
      <p>Let us consider the main concepts related to the quality of BD [10] some concepts was taken from
the digital library domain and adapting to big data. Quality is a set of properties of objects that give
them the ability to satisfy the stipulated or anticipated needs of the consumer following the purpose.</p>
      <p>The quality characteristic is a property or a set of object properties, with the help of which quality
can be described and evaluated. Each object has its nomenclature characteristic. A characteristic can
be a composition of other characteristics, forming a hierarchical structure.</p>
      <p>Metric is a formula or rule for determining the degree to which an object possesses a characteristic.</p>
      <p>A quality indicator is a quantitative or qualitative value, obtained as a result of the procedure for
evaluating the quality of a characteristic according to the evaluation methodology. Quantitative
indicators have a numerical expression within a certain scale. Qualitative indicators have a verbal
expression within a certain verbal ordered scale.</p>
      <p>Quality level is the degree of acceptability of the obtained quality indicator from the view of the
expected (planned) quality.</p>
      <p>The quality system is a set of organizational structures, methods, processes, procedures and
resources necessary for the general direction and management of quality by established methods. It
includes quality policy, quality model; quality achievement system; quality system documentation.</p>
      <p>The quality policy is a document developed by the responsible management. It expresses the goals
in the quality field, the acceptable level of quality, the duties of various persons and structures for
quality assurance, a set of measures to achieve quality. The quality policy is defined based on tasks
set in the quality field.</p>
      <p>Quality model is a set of objects for which it is described, evaluated and supported. Also, it
includes quality characteristics, methods and means of quality assessment, metrics and algorithms for
determining quality indicators. A specific quality model is selected based on the developed quality
policy and other factors.</p>
      <p>Achieving quality is a set of organizational structure, responsibilities, procedures, processes and
resources that implement general quality management [11].</p>
      <p>The quality management system is an organizational structure that includes personnel who
implement quality management functions using established methods.</p>
      <p>Quality management is the general management of quality provided by resources, particularly
human resources. It organizes quality assurance work, interacts with the external environment, defines
policies, goals and plans in the quality field, and makes strategic and important operational decisions
regarding quality.</p>
      <p>Also an quality assurance is creating confidence that quality requirements will be met. It includes
administrative and procedural measures carried out within the framework of the quality system to
ensure the fulfillment of requirements and goals. This is a systematic measurement, comparison with
a standard, process monitoring, making technological or any other process adjustments to achieve the
required quality.</p>
      <p>Quality control is a set of measures, procedures, methods and means that allow performing a
systematic and independent analysis. It is possible to determine the compliance of activities and
results in the quality field with the planned measures and the effectiveness of their implementation
and compliance with the set goals. The quality assurance system is the subject of the system analysis.</p>
      <p>Quality assurance</p>
      <p>Manage</p>
      <p>Quality management
Monitors execution</p>
      <p>Evaluates efficiency</p>
      <p>Manage
Quality control</p>
      <p>Quality assessment</p>
      <p>Tools of support
Methods</p>
      <p>Approach</p>
      <p>Tools
• the quality of the product itself, without taking into account its behavior with the external
environment (internal quality);
• product quality regarding its behavior in the external environment (external quality);
• the quality of technological processes of product development (process quality);
• the quality of the product to its use in different contexts (and the quality experienced by
the user in specific scenarios of product use (quality during use)).</p>
      <p>B. The quality model should include all stages of the BD development and use life cycle starting
from requirements development and ending with the industrial operation.</p>
      <p>С. The quality model is relevant to all structural elements of BD. It contains all types of support
for the software system — functional, informational, mathematical, technical, etc.</p>
      <p>D. An important component of the quality model is the structure of quality characteristics and
metrics that assess elementary characteristics.</p>
      <p>BD consist of two components are data and data base application, information is retrieved from a
computerized BD by using a computer program.</p>
      <p>The semantic information model for BD defines as a set of information objects in which each
predicate define through top-level ontology.</p>
      <p>пм1,
c (M i, p) = нппп0,
о
object Mi has property p;
another case.</p>
      <p>.</p>
      <p>Then the estimate of the degree to which the set of objects M has the property p is equal to:
If the objects M i ( i = 1,..., N ) are unequal and their weighting factor K i : 0 Ј K i Ј 1 ( i = 1,..., N ),
is given for each of them, which determines the relative importance of the objects, then the above
formula takes the following form:</p>
      <p>N</p>
      <p>Similarly, a metric can be defined for a situation where one object can have multiple properties
and it is necessary to determine to what extent they are inherent to the object.</p>
      <p>Establishing acceptable values for certain characteristics and adding a qualitative measure to the
appropriate range is important for metrics. This range can be determined experimentally or
algorithmically. An expert establishes it in many cases. For example, let's imagine as j an expert with
K j competence specifying a range of values клйX ij ,Y ij ъыщfor the i characteristics, where Y ij - the optimal
value of the characteristic is X ij - its worst value.</p>
      <p>M experts evaluated the characteristics. The final score for the range of values is calculated as
follows:</p>
      <p>N
е c (M j , p)
M (p) = j = 1</p>
      <p>N</p>
      <p>.</p>
      <p>N
е K j Чc (M j , p)
M (p) = J = 1</p>
      <p>.</p>
      <p>M
е K j ЧX ij
X = j = 1
i</p>
      <p>M
е K j
j = 1</p>
      <p>M
е K j ЧY ij
Y = j = 1
i</p>
      <p>.</p>
      <p>M
е K j
j= 1</p>
      <p>Each information IO object in the BD environment is defined as a certain directed acyclic graph
where the information object consists of a list of statements in the triplet «subject - predicate
object». The set of such triplets forms a directed graph, in which vertices are subjects and objects, and
edges are predicates. Certain metadata describes each node of such a graph. That is, the model of the
information object in the BD environment is defined as IO = (s(m ), p(m ),o(m )).</p>
      <p>Evaluating the quality of elementary characteristics involves determining their metrics represented by
formulas or rules for determining the degree to which an object has an elementary characteristic [13]. The
metric of an elementary characteristic reflects the degree to which an object or a set of objects possesses a
certain property. Let a set of equivalent objects M = {M i } where ( i = 1,..., N ), be given, which may or may
not have a certain property. We define the following characteristic function:</p>
      <p>It should be noted that intervals клйX ij ,Y ij ъыщare set by experts or determined algorithmically only for
elementary characteristics. At other levels, i.e. for integral characteristics, the minimum and
maximum values are calculated according to the defined formulas based on the given or calculated
values of the previous levels [13].</p>
    </sec>
    <sec id="sec-5">
      <title>4. Quality properties of information objects in Big Data</title>
      <p>Next, the issues of evaluating the quality of semantic information objects are considered. IO
quality characteristics.</p>
      <p>Accessibility is a complex function that depends on many factors, including:
(3)
(4)
(1)
(2)
• the IO is actually available in the BD (the information object may be in the BD, but for
some reasons, it may be removed from public access or due to the amount of data, it may
not be identified among a set of objects);
• there is a service that can find the IO (one of the ways to remove an information object
from public access is to deactivate its searching characteristics);
• it is the network and data transmission system in the network operational;
• there are no restrictions on access to the IO or if there are such restrictions they do not
apply to specific persons or groups of persons.</p>
      <p>It should be noted that in the given context, they talk about the availability of the IO to perform a
single operation as reading. Our review does not include other possible operations with IO (changes,
deletion, administration).</p>
      <p>For BD this is availability provided by a specific service that interacts with BD. As a rule, a
distinction is made between availability for all and certain services. In this case, the restriction of
access rights A cc (Si , IOj ) where the S i service for IO j , means a function that acquires the following
values: 1 — the service does not have access restrictions or it belongs to the group to which access is
open; 0 — otherwise.</p>
      <p>Now, if we mark other availability indicators as Pi except for access rights restrictions which take
the following values: 1 — the indicator is satisfied, 0 — the indicator is not satisfied, then the general
availability formula is calculated as follows:</p>
      <p>MIN (P1, ..., Pn , A cc (Si , IOj )).
(5)</p>
      <p>Relevance is the measure to which the information content of the information object meets the
information needs of the user. Both cannot be strictly formalized. This assessment largely depends on
the depth of the user's knowledge about their information needs at the current time and the tasks
facing them. The user's information needs at the current moment are expressed through his
information search query as a result of knowledge reasoning. The query implicitly defines the context
in which relevance is evaluated. The user carries out an evaluation of this compliance as a result of
receiving a response to the request (the user can be a group of people).</p>
      <p>The relevance evaluation function is as follows R elevance (IOi , S j ,Queryk ):
R elevance (IOi , S j ,Queryk ) = нппм1 - Servise Sj appove that IOi, is relevant for Queryk ьпп
оппп0 - anot her case эпппю
(6)</p>
      <p>Accuracy of storage. In the process of existence, the object can go into different states caused
by the transition to other software and technology platforms. Big data is characterized by constant
changes, and errors in these data also tend to accumulate and scale, including changing the storage
format, using newer versions of BD, etc. All this can lead to a loss of storage accuracy of the new
version of the information object compared to the old one. This characteristic requires assessing the
loss degree of storage accuracy based on comparing states in the dynamic environment that in general
is a complex task and required additional research [13].</p>
      <p>Credibility means that the IO has the ability to confirm that it is what it should be. The ability
to verify and measure the extent to which an IO is what it is claimed to be is fundamentally important
in its correct perception and use. Reliability determines the extent to which the IO can be relied upon.
This is largely determined by the developer's credibility and origin source. The credibility of the IO
can be measured by:
• the attitude of users towards the IO itself;
• the attitude of users towards the source of the IO;
• the availability of information on the chronology of IO changes;
• the attitude of users to the BD in which the IO is located.</p>
      <p>Integrity determines to what extent the IO is complete and correct from the point of view of the
software object it represents. Integrity contributes to increasing trust in the IO [14]. Accuracy of
reproduction determines the degree of accuracy of the reproduction of the IO of its original. For
example, a text document reproducing an ancient book can accurately reproduce the text and
completely ignore its artistic design.</p>
      <p>Timeliness indicates that the IO is introduced and updated on time, as this issue is specific to
BD. This characteristic evaluates how quickly the set s(m ), p(m ),o(m ) in IO is updated compared to
the real state of affairs.</p>
      <p>The characteristic is measured by the ratio of the actual delay time compared to the permissible
one:</p>
      <p>T imeliness (IO(s, p,o)) =</p>
      <p>real time delay
exp ected time delay .</p>
      <p>(7)</p>
      <p>Origin is a characteristic of the quality of an IO. It indicates how well (correctly, completely,
qualitatively) the entire prehistory of the origin and change of an IO is presented, and how accurately
and during what period it is possible to trace the prehistory of the existence of an IO. This is an
important characteristic since inference over semantic data depends on the data itself. Understanding
the historical information about the data helps to determine the reasons for changing the system's
behavior, which is not a trivial task in the BD environment.</p>
      <p>Susceptibility indicates how easily a person can understand and accept IO. It can be used to
analyze which set of IO is most easily perceived by a group of persons due to the solved tasks.</p>
      <p>Practical aspects of assessment of the quality of BD. One of the most challenging tasks in
achieving data quality metrics is the early detection of data-related problems. Typical problems
include completeness, the integrity of data and lack of contradictions. The problem lies in that in the
conditions of the BD, the time to detect such issues may exceed the time requirements for receiving a
response to the information from the BD. That is why it is necessary to develop methods that will
allow the detection of such problems at an early stage. There are various approaches to deal with the
task, like the way to control all data entered into the system through the ontology. In practice, it is
often not known what the data model should be since the requirements for the BD system can change
as the data increases. These requirements can be constantly updated. This means that data previously
entered into the BD management environment in the previously specified structure may not
correspond to the quality model after some time. Identifying these problems due to the scale is a
difficult problem.</p>
      <p>One of the criteria of the quality model is the ability of BD to give a quick response to user
requests. The most effective method of increasing such speed is materialization [15]. Materialization
can be used to improve performance at query time by making the required information explicit in
advance. Thus, recalculation of the necessary information for each separate request is avoided.
However, this method can be ineffective if there is excessive materialization.</p>
      <p>Consider a certain graph of semantic data G in which the connections between concepts are
built on the basis of descriptive logic. We will briefly describe the DL, which is the basis for all DL of
the family. ALC means «Attributive Language with Complements». It is defined in [16]. The
language is based on the previously introduced language AL (Attributive Language), to which the
addition constructor (negation) was added. Syntax describes a set of correctly constructed language
expressions, and semantics indicates their formal meaning.</p>
      <p>be finite, non-empty sets of atomic concepts</p>
      <p>Let CN = {A1, . . . , Am } і R N = {R1, . . . , Rn }
and atomic roles. The ALC syntax is defined as follows:
• M and L are concepts;
• an arbitrary atomic concept A is a concept;
• if C is an arbitrary concept, then ШC , Ch D and Cg D are concepts, corresponding
constructors are called addition, intersection and union;
• if C is a concept, R is an atomic role, then j R .C and i R.C are arbitrary concepts.</p>
      <p>ALC semantics is defined through the concept of interpretation. An interpretation is a pair of
I = (D, gI ) where Δ – is a non-empty set, called the domain of interpretation, aI is an interpreting
function that assigns the relation A I8 Δ to each atomic concept A and to each atomic role R as an
binary relation RI 8 ΔЧ Δ . Other formulas are interpreted as follows:
(y A)I = D \ AI , (ChD)I =  C I 1DI , (CgD)I = C I 2DI</p>
      <p>j R.C = {a9D |j b9D ((a,b)9R I ob9C I )}
i R.C = {a9D |i b9D ((a,b)9R I ® b9C I )}</p>
      <p>There are cases when it is necessary to describe specific characteristics of an object In order to
describe the real world, for example, the number of pages in an information resource. To solve this
problem, a specific area with a fixed set of predicates is created [17]. A concrete domain is a
pair D = (D, ), where D is a non-empty set and  is a set of predicates in D . It can be assumed that
given a set of predicate symbols PN where each predicate symbol P О PN  is associated with an n
arity and  maps an n - relation to it as P D8 Dn . It should be noted that  always contains a single
predicate D , that is PN always includes M symbol and is interpreted as MD = D . Also is always</p>
      <p>Next the essence of the (T Box ) terminology is revealed for DL ALC . However, all introduced
concepts are easily transferred to other DL. Terminologies describe general knowledge about
concepts and roles. To describe knowledge about specific individuals (their belonging to concepts and
roles), the DL offers a system of facts about individuals or A B ox . For this, a set of names of
individuals is entered into the DL. There are two types of facts: a statement about an individual's
belonging to a concept (written as C (a ) ); the statement about the belonging of a pair of individuals a
and b to role (written as R (a,b)).</p>
      <p>A system of facts or A B ox is a finite set of statements of form C (a ) and R (a,b) , where a and
b are individuals, C is an arbitrary concept and R is a role.</p>
      <p>Here are some ALC extensions that were used to fulfill the tasks of the dissertation work.</p>
      <p>R -follower is an individual who is the right part of the role R . We denote the set of R
followers for e that can be written as R I (e) , where e О D : R I (e) = {d О D | (e,d) О R I }. We denote
the power of such a set by |R I (e)| . The following constructors are called numerical role constraints. If
R is a concept, n and 0 is a natural number, then:
• Ј 1R is a concept for limitation of functionality;
• Ј nR and і nR is a concept for quantitative limitation;
• Ј nR.C and і nR.C is a concept for qualitative limitation.</p>
      <p>The following constructors are interpreted as follows:
(Ј 1R ) = {e О D | R I (e) Ј 1},</p>
      <p>I
(Ј nR ) = {e О D | R I (e) Ј n }</p>
      <p>I
(і nR ) = {e О D | R I (e) і n }</p>
      <p>I
(Ј nR.C ) = {e О D | R I (e) З C I Ј n }</p>
      <p>I
(і nR.C ) = {e О D | R I (e)1C I і n }</p>
      <p>I
,
,
,
closed with respect to the complement, that is for every n-predicate symbol P in PN there is an
npredicate symbol in P , which is interpreted as Dn \ P D .</p>
      <p>Let be a given concrete area D with a set of predicate symbols PN. Also let a finite set of
symbols be given: CN are atomic concepts, R N are atomic roles, A F8 RN are atomic abstract
attributes, CF are atomic concrete attributes. A sequence of f1 ј fkh з k і 1 with atomic abstract
attributes fi О A F and one concrete attribute h ОCF will be called a complex, concrete attribute.</p>
      <p>Concepts of ALC(D) logic are defined by grammar [17]:</p>
      <p>M|L | A |yCh D | Cg D |j R.C |i R .C |j клйu1 ј un ъыщ.P</p>
      <p>M|L | A |y C h D | C g D |j R .C |i R .C |Ј 1R |
Ј nR |і nR |Ј nR .C |і nR .C |j клйu1 ј un ыщъ.P .
additions:
•
•
where A ОCN , R О R N , u1, ..., un are arbitrary attributes, P О PN is the n-concrete predicate.
The semantics of ALC(D) logic is considered as I = (D,gI ) interpretation with the following
sets Δ and D must not intersect;
each atomic abstract attribute f О A F is assigned a partial function f I : D ® D ;
•</p>
      <p>each atomic abstract attribute h ОCF is assigned a partial function f I : D ® D .</p>
      <p>A composite concrete attribute u = f1 ј fkh is interpreted as a composition of partial
functions uI (x ) = hI (fkI (ј f1I (x )ј ). As a result, a partial function uI : D ® D is formed.</p>
      <p>The only new (compared to ) type of concept is interpreted as follows:
x1 oј ou I n (e) = xn o{x1,ј , xn }О P D }
(j клйu1 ј un ъыщ.P )I = {e О D |j x1 ј xn О D : u I 1 (e) =</p>
      <p>The set of points on which the attribute u is defined is expressed by the concept u., where 
is a specific predicate that is always present in the PN signature. The following equivalence is valid:
y $ клйu1,ј , un ыщъ.P є y $u1.Mgј h y $u1.Mg $ йклu1,ј , un ыщъ.y P .</p>
      <p>Indeed, the condition e О (y j клйu1,ј , un ъыщ.P )I means that either one of functions uiI is undefined
at point е or the tuple u1I (e),ј , unI (e) does not belong to the predicate P D , P, but belongs to its
complement (y P )D . So, the G graph we have is given by BD
(17)
(18)
(19)
(20)
(21)</p>
      <p>When building a materialization, rules are set according to which it should be built. Consider
the problem of excessive materialization, which can be caused by the following way of constructing
concepts. For example, let's take the computer components motherboard and RAM. The concept that
will determine the compatibility of these two components will be defined as follows:
R amDDR 4 _ 2_ MaimboardDDR 4 є
(" hasSlotT ype.DDR 4g Memory)h(Ј 1hasSlotT ype.DDR 4gMainboard)</p>
      <p>As a result of the materialization, we will get the next G graph that will be set C 62 = 15 possible
combinations that will determine the concept RamDDR 4_ 2_ MaimboardDDR 4 . If we take into
account that the motherboard also has limitations in terms of supporting the maximum size of RAM
and the real situation will become even more complicated.</p>
      <p>Table 2</p>
      <sec id="sec-5-1">
        <title>Example of configuration with combination RAM</title>
      </sec>
      <sec id="sec-5-2">
        <title>RAM Size</title>
      </sec>
      <sec id="sec-5-3">
        <title>Main Board</title>
      </sec>
      <sec id="sec-5-4">
        <title>RAM Slots</title>
      </sec>
      <sec id="sec-5-5">
        <title>Ram Model 1 32 MainBoard 2</title>
      </sec>
      <sec id="sec-5-6">
        <title>DDR4 Model 1 DDR 4</title>
      </sec>
      <sec id="sec-5-7">
        <title>Ram Model 2 12 MainBoard 4 128</title>
      </sec>
      <sec id="sec-5-8">
        <title>DDR4 Model 2 DDR 4</title>
        <p>Such dependence means that even with a small number of components, the knowledge base
representation system will have to store a huge number of relationships that will determine the
materialization. Accordingly, the inference on such a graph will work very slowly due to the huge
number of combinations that form nodes of the graph available for search, as stated in [17], such an
inference problem belongs to the P Space class. This means that the complexity depends on the size of
the input data and to solve the problems of inference and feasibility of concepts, it is necessary to
reduce the set of input data. To avoid such a problem, it is proposed to divide the knowledge base,
which traditionally consists of TBox and ABox into two components, so that the subject area is
described DL SHI F T and then ALC(D) (Fig. 2)</p>
      </sec>
      <sec id="sec-5-9">
        <title>Max Memory support 32</title>
        <sec id="sec-5-9-1">
          <title>Description logic</title>
        </sec>
        <sec id="sec-5-9-2">
          <title>Knowledge base</title>
        </sec>
        <sec id="sec-5-9-3">
          <title>Reasoning</title>
        </sec>
        <sec id="sec-5-9-4">
          <title>Description logic</title>
        </sec>
        <sec id="sec-5-9-5">
          <title>Knowledge base</title>
        </sec>
        <sec id="sec-5-9-6">
          <title>Reasoning</title>
        </sec>
        <sec id="sec-5-9-7">
          <title>TBox</title>
        </sec>
        <sec id="sec-5-9-8">
          <title>ABox</title>
        </sec>
        <sec id="sec-5-9-9">
          <title>Tbox(D)</title>
        </sec>
        <sec id="sec-5-9-10">
          <title>Abox(D)</title>
        </sec>
        <sec id="sec-5-9-11">
          <title>Dig data</title>
        </sec>
        <sec id="sec-5-9-12">
          <title>Rules</title>
          <p>K D is called executable and the I interpretation is called a K D model and written as I QK D .</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Estimate quality of metadata and an information object family in Big Data</title>
      <p>Metadata quality assessment is intended to find out to what extent certain metadata or metadata
schemas present in a BD meet the tasks that were set before the BD when it was designed. They
contribute to the quality functioning of the semantics in the BD. The quality of metadata affects many
processes related to the use of inference, building connections between the IO description, their input,
storage, identification, search and access.</p>
      <p>There are two aspects of quality related to metadata. The first of them refers to IO metadata
(what IO metadata is, how fully it describes IO, whether it meets a certain metadata schema standard).
The second aspect is related to the schema of metadata (is the schema of metadata standard, to what
extent the chosen schema meets the needs of the description of IS in a specific subject area). The
quality of both aspects is described below.</p>
      <p>Compliance with the standard. This characteristic indicates whether a standard IO metadata
description scheme is used. The use of a standard metadata scheme is a fundamental issue in the
consideration of the problem of the organization of search and retrieval of knowledge. The existence
of IOs in the DB, the metadata of which do not meet or do not fully meet the standard, significantly
reduces the resolution of fundamentally important issues facing the DB and reduces its quality. The
measure of compliance with the standard can be the ratio of the number of non-standard metadata to
the total number of metadata used in the description of the IO [13]:</p>
      <p>where n(IO(md)) is the total number of IO metadata, а is the number of metadata that does not
meet the standard adopted for this BD model</p>
      <p>The completeness of the description of the IO in relation to the metadata scheme. This
characteristic indicates the extent to which the metadata schema is fully used to describe the IO.
Please note that not all metadata of the selected scheme can be applied to some types of IOs. Several
metadata schemes can be used simultaneously in the DB network, but the completeness is determined
relative to only those metadata that participate in the construction of semantic links between IOs.</p>
      <p>Therefore, the degree of completeness of the description of the IO, according to the selected
MS metadata scheme, is determined as follows:</p>
      <p>S tandard (IO ) = 1
w(IO(m d))
n(IO(m d)) ,
Com pleteness (IO, MS ) =</p>
      <p>P resent (IO(md))
Required(IO(md)) ,
(22)
(23)
(24)
(25)
(26)
where: m d is metadata, MS is metadata schema, P resent (md) is the total number of metadata
required to describe the IO, which is actually present in the IO description, Required(IO(md)) is the
total number of MS metadata required to describe the IO.</p>
      <p>Compliance with metadata schema. A metadata schema can set certain properties to its
metadata. The characteristic of matching the metadata scheme determines how well the properties of
the metadata of the IO correspond to the properties of the corresponding metadata of the selected
scheme. Such properties include the type of data or attributes of relations between IOs, which in
general are also included in the quality model.</p>
      <p>Let n is the number of metadata in the MS scheme, m i – is the number of properties of
metadata mdi , Conformance (i, j ) is compliance of property j of metadata mdi ІО with the standard
specification of the MS schema. Conformance (i, j ) is calculated by the formula:</p>
      <p>Conformance (i, j ) = пнпм1 - iff i - metadata property belong to j - property from MS
ппп0 - ot herwise
о
The correspondence of the IO to the i MS schema metadata is calculated according to the formula:
Then the correspondence to the Conformance(MS) metadata schema is calculated using the formula:
Conform ance (m di ) = j = 1
n
е Conformance (mdi )
Conformance (MS ) = i= 1
mi
е Conform ance (i, j )
m i
n
е S tandard (IOi )
S tandard (MS ) = i= 1
n</p>
      <p>Metadata scheme quality characteristics. A set of specially selected metadata make up a
metadata schema. In the general case, such a set can be arbitrary, but this significantly reduces the
quality of the BD, because our BD environment becomes isolated from other data sets and will not be
able to take (at least fully) in the process of integration and reasoning information, in a sense the
system becomes isolated because even using mappings between data schemas will be inefficient due
to the scale of the data. In this regard, efforts are being made to develop and use standard metadata
schemas, which are usually aimed at describing IOs of a certain class. There are many metadata
schemes. In this connection, the question of choosing the most suitable for a certain subject area
arises. This task is facilitated by the evaluation of the quality of the metadata scheme.</p>
      <p>Compliance with standard metadata schema. This characteristic evaluates the extent to which
all DB information objects conform to the standard. For IO, the characteristic of compliance with the
standard is also significant, but it is at the IO level. In general, compliance with the standard scheme
is evaluated as the arithmetic mean of compliance with the IO standard
(27)
(28)
(29)
(30)</p>
      <p>The completeness (usage) of the metadata scheme. This characteristic provides an opportunity
to assess how much a certain scheme is used to describe the entire population of BD IOs. It is based
on the characteristic of the completeness of the description of the IO in relation to the metadata
scheme and is its arithmetic average for all IOs of the BD:</p>
      <p>This characteristic makes it possible to assess to what extent the decision to use a certain
metadata scheme is justified, and, if necessary, to make a decision to replace it.</p>
      <p>Let's introduce metrics for evaluating the IO family. A family is a systematized set of IOs that
are united into a single whole based on some meaningful or formal criteria of belonging, for example,
regarding the general content, sources, purpose, semantic independence, method of use, etc.</p>
      <p>Completeness of the family of IO. This characteristic establishes to what degree of
completeness the family contains those IOs that it should contain. Completeness can be measured
only when it is known what exactly the collection should contain, that is, when the original family,
which acts as a sample, is known [13]. As a rule, families are distinguished on the basis of IO
attributes.</p>
      <p>The formula for measuring family completeness is as follows:</p>
      <p>Com pleteness (F ) =
n
е IOi (F )
i= 1
n
е IOi (Foriginal )
i= 1</p>
      <p>Conformity of the collection to the standard. Determines the extent to which collection IOs
conform to the standard. Compliance with the standard of the family can be considered as the
arithmetic average of compliance with the standard of its IO:
where n is the number of IOs in the collection</p>
      <p>A variety of standards. It is believed that the family should be based on one standard metadata
scheme specified in the external ontology, as the use of many schemes deteriorates the operational
characteristics. The quality of this feature can be measured as the inverse of the number of metadata
schema standards used.</p>
      <p>Consistency. There are many different situations where a collection can be considered
inconsistent (conflicting). For non-limiting generalizations, we consider only one situation when there
are two IOs with absolutely identical values of their metadata.</p>
      <p>Let the function IdentMd (IOi , IOj ) acquire the following values:</p>
      <p>Then the family matching function is defined as follows:</p>
      <p>n n IdentMd (IOi , IO j )
Consistency (F ) = 1 - е е
i= 1 j = 1,j №i n Ч(n - 1)</p>
      <p>For modeling our approach, we are using Neo4j as a system for storing and managing big data
[18], [19]. Neo4j is a database whose data model is a graph, specifically a property graph. We took a
database for electronic components consisting of boxes, main boards, and memory modules. Our goal
is to find all available interpretations which will be models for our knowledge base. It means the need
to find all compatible components or find a list of components that are compatible with the selected.
This problem more detail describe in [20], [21], [22]. As specified in these works the quality of the
result depends on the quality of metadata. And another important characteristic for semantic networks
is the speed of reasoning for checking interpretation. It is related to time which needs to get answers
about the compatibility of electronic components.</p>
      <p>The metrics of quality data are allowing us to reveal a problem with missing required metadata
for interconnecting components. Due to this information and metrics like compliance with metadata
schema as a result of cleaning data, we built graph storage which consists of 44195 relations, we don't
have any nodes without missing important data. This graph has a relation between memories, main
boards, and cases. At first look, this graph does not belong to big data but if we take only 54 different
types of memories, 113 types of mainboards, and 119 types of cases the result of materialization gives
246912 available combinations for our system. This materialization is not included in the concrete
domain. Materialization in the concrete domain will bring an enormous quantity of available nodes
because if we have for example attribute which describes the count of ram slots on main board it
allows putting on these slots a different combination of memory modules. Our optimization also
includes checking only bi-directional dependencies between components.</p>
      <p>Our idea to split the knowledge database into two-part brings the possibility of extracting
information from a database with materialization without a concrete domain.</p>
      <p>We are build relation in our graph that it responsibility to DL SHI F T the main condition for
building relation avoid concrete domain. On Fig. 3 demonstrate relation between our components.
into account all possible variations, including the quantities of the selected components.</p>
      <p>The second approach q2 consisted in grouping components by common value of attributes in such
a way as to avoid building additional connections. And the last optimization q3 consisted in the fact
that first all compatible components were searched, and only then the conditions of quantitative
restrictions for a concrete domain were checked for satisfaction.</p>
      <p>One problem is that the same component can be reinstalled twice or more depending on the
number of previously selected components. That is, if the motherboard has 8 RAM sockets, then there
may be a situation when 8 identical memory modules are selected, and there may be 8 different
modules. Moreover, for the motherboard, we must check not only the quantitative limitation of the
number of occupied sockets, but also the limitation regarding the maximum amount of memory
supported by the motherboard
Table 3</p>
      <sec id="sec-6-1">
        <title>Example of configuration with combination Type optimization and query</title>
        <p>q1 list all mainboards
q1 list all mainboards for the specific memory modules
q2 list all mainboards (specification was grouped)
q2 list all mainboards for the specific memory modules (specification was
grouped)
q3 list all main boards (specification was grouped and quantity restriction
included) time for two query
q3 list all main boards for the specific memory modules (specification was
Execution</p>
        <p>time
5612 ms
grouped and quantity restriction included) time for two query</p>
        <p>As we can see, the simplification of requests gives a significant increase in the speed of execution.
But result BD systems depend on characteristics such as the completeness of the description,
compliance with the metadata scheme. It should be noted that according to the expert evaluation of
work with web resources, the response of the web service should be up to 600 ms.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusions</title>
      <p>The complexity of big data applications combined with the lack of standards for the representation
of information objects, processing and storage requires significant resources. Data quality is one of
the approaches that will allow achieving modeling of data that will require simpler algorithms for
analysis. Analysis of data quality allows increasing their accuracy in various aspects. Enrich data
semantics is a complex process of describing big data by ontological means. However, there is a
problem with the speed of inference, the article proposes a method of knowledge base materialization
in the environment of big data to optimize inference. The quality of the data plays a key role in this,
allowing to build of appropriate graphs of schematic data on the basis of metadata.</p>
      <p>Higher data quality levels can help produce better reasoning results but also help improve data
maintainability and reusability and integration.</p>
    </sec>
    <sec id="sec-8">
      <title>7. References</title>
      <p>[1] J. Stuart Ward and A. Barker, "Undefined By Data: A Survey of Big Data Definitions" 2013.
[2] D. Laney, «3D data management: Controlling data volume, velocity and variety» META
group, 2001.
[3] M. Schroeck, R. Shockley, J. Smart, D. Romero-Morales and P. Tufano, "Analytics: The</p>
      <p>Real-World Use of Big Data" IBM, 2012.</p>
      <p>O. Novytskyi, G. Proskudina та O. Ovdiy, «Development of an digital library quality
model» в Інформація, комунікація, суспільство 2014 : матеріали 3-ої Міжнародної
наукової конференції ІКС-2014, 21–24 травня 2014 року, Україна, Львів, Славське,
2014.</p>
      <p>O. M. Spirin, S. M. Ivanova, O. V. Novytskyi, Z. V. Savchenko, V. A. Reznichenko,
A. V. Yatsyshyn, N. M. Andriychuk, V. A. Tkachenko, M. A. Shinenko and Y. A.
Labzhynskyi, Collective monograph. Electronic library information systems of scientific and
educational institutions, В. Ю. Биков and О. М. Спірін, Eds., Kyiv: Pedagogical press, 2012,
p. 176.
2002.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [4]
          <string-name>
            <surname>I. C</surname>
          </string-name>
          . Intel IT Center,
          <article-title>"Centre. Big Data Analytics: Intel's IT Manager Survey on How Organizations Are Using Big Data"</article-title>
          <source>Santa Clara</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Suthaharan</surname>
          </string-name>
          , «
          <article-title>Big data classification: Problems and challenges in network intrusion prediction with machine learning» ACM SIGMETRICS Performance Evaluation Review, т</article-title>
          .
          <volume>41</volume>
          , № 4, pp.
          <fpage>70</fpage>
          -
          <lpage>73</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wilkinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dumontier</surname>
          </string-name>
          , I. Aalbersberg, G. Appleton,
          <string-name>
            <given-names>M.</given-names>
            <surname>Axton</surname>
          </string-name>
          , А. Baak,
          <string-name>
            <given-names>N.</given-names>
            <surname>Blomberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Boiten</surname>
          </string-name>
          , L. da Silva Santos,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bourne</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Bouwman</surname>
          </string-name>
          ,
          <article-title>"The FAIR Guiding Principles for scientific data management and stewardship" Scientific data</article-title>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ceravolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Azzini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Angelini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Catarci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Cudré-Mauroux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Damiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mazak</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Van Keulen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Jarrar</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Santucci</surname>
            and
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Sattler</surname>
          </string-name>
          ,
          <article-title>"Big data semantics"</article-title>
          <source>Journal on Data Semantics</source>
          , vol.
          <volume>7</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>85</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Amsler</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>Application of Citation-based Automatic Classification» Austin</source>
          ,
          <year>1972</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Woods</surname>
          </string-name>
          , «
          <article-title>What's in a link: Foundations for semantic networks</article-title>
          .
          <source>» Representation and understanding</source>
          , pp.
          <fpage>35</fpage>
          -
          <lpage>82</lpage>
          ,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>