<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluation of Network Reliability and Element Importance Metrics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandr Moshnikov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ITMO University</institution>
          ,
          <addr-line>49 Kronverksky Pr., St. Petersburg, Russian Federation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article considers an approach to assessing the importance metrics and reliability of networks. The Monte Carlo method is used to estimate Birnbaum metrics and failure probability with determination of the confidence interval. To conduct a computational experiment, the R software package is used. A description is given of the representation of the control system reliability model in the iGraph package, which provides visualization of the results. The model of a three-level network structure is considered as an example.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Monte Carlo method</kwd>
        <kwd>Reliability estimation</kwd>
        <kwd>R language</kwd>
        <kwd>Importance metrics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Currently, there is a rapid development of information technologies and their implementation
in various areas of human activity [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Control transmission networks have become an integral
part of people’s lives, without which information exchange is practically unthinkable. In such a
situation, the analysis of the technical characteristics of existing data transmission networks
and the design of new networks, taking into account the given characteristics, remains one of
the urgent tasks in the field of information technology.
      </p>
      <p>
        In addition to such technical characteristics of computer networks as: performance, latency,
security, scalability, extremely important characteristics are complex reliability indicators:
availability factor, average unavailability time per year [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The reliability of the network also
indirectly depends on the safety of the operation of control systems for any objects in which the
untimely response (due to failures and failures in the data transmission network) of the control
system to any critical changes in the control object can lead to serious consequences. In this
situation, the analysis of reliability indicators of distributed control systems is a particularly
relevant problem. Reliability is defined as the probability of a system or a sub-component
functioning correctly under certain conditions over a specified interval of time [3].
      </p>
      <p>Issues of reliability of systems with a network structure are still relevant [4, 5].</p>
      <p>For instance, the reliability of network nodes, termed as the terminal reliability, is the
probability that a set of operational edges provides communication paths between every pair of
nodes. Another closely related concept with reliability is availability, which can be defined as
the probability that a component will be available when demanded [3].</p>
      <p>Importance measures (IMs) are used to evaluate the efect of component reliability on system
reliability. IMs are useful tools in reliability engineering [14], risk analysis [15, 17], and system
reliability optimization. These measures can help reliability engineers to find a better solution
rapidly because they can identify the weakest links of the system, which are the premise and
foundation of system design, maintenance, and resource configurations. During the system
design period, component importance can help designers determine cost efective design ideas
with relatively high system reliability and low cost rapidly.</p>
      <p>Note that the significance of a system element according to Birnbaum  reflects the
degree of influence of changes in the element’s readiness coeficient on changes in the system’s
readiness coeficient. The significance of a system element according to Barlow-Proshan 
reflects the probability that a system failure that occurred at a certain point was caused by this
element. The significance of a Vesely-Fassel system element reflects the probability that this
element is one of the failed elements, provided that the system failed. Vesely-Fassel significance
  characterizes those elements that are most often involved in system failures.</p>
      <p>The cost of increasing the risk for a system element  reflects the importance of
maintaining the current level of reliability of this element. The cost of reducing the risk for a system
element  reflects the degree to which the system’s availability coeficient increases if this
element is replaced with a flawless element. The critical significance of a system element 
reflects the probability that this element is critical for the system at a given time. The potential
for improvement of a system element  reflects the gain in system reliability if this element is
replaced with a completely reliable one [16].</p>
      <p>The most widely used importance measures were acquired in the component allocation
problem. The optimal component allocation problem arises when a pool of available components
exists (in the market or in the company store) for each subsystem and a designer chooses a
set of components for each subsystem such that the system reliability is maximized subject to
constraints such as weight and cost.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Reliability and importance assessment</title>
      <sec id="sec-2-1">
        <title>2.1. Monte-Carlo method for reliability estimation</title>
        <p>The easiest way to estimate  is to use Crude Monte-Carlo simulation [19]. Let (1), ..., ()
be independent identically distributor random vectors with the same distribution as . Then
1
 =  ·

∑︁ ℎ(())
=1
( ) = 1 · ( − 2)
is an unbiased estimator for  , where ℎ() is reliability function. Its sample variance is given
by
(1)
(2)</p>
        <p>An important measure for the "eficiency" of any estimator is its relative error. Relative error
for  is given by
√︂ 1 − 
( ) = (3)</p>
        <p>To achieve acceptable accuracy, we assume an error less 0.01, and the number of iterations is
5143.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Element importance metrics</title>
        <p>It is natural to assume that diferent elements afect the system’s behavior in terms of reliability
in diferent ways. The ability of the researcher to quantify the nature of the elements influence
on the behavior of the system is of particular importance in the analysis of systems. This makes
it possible to identify system weaknesses, select optimal redundancy, and make a rational impact
on the reliability of the system as a whole.</p>
        <p>The importance of the element  in the system is defined as a private derivative of the
availability factor (the probability of) the system availability (the probability of) the element for
which an analysis of its significance:
 (, ) =
ℎ()</p>
        <p>This characteristic is called Birnbaum significance (BIM-significance) [ 6]. The significance
is estimated by the number of times the system availability coeficient increases when the
element availability coeficient increases. BIM-significance does not depend on the readiness
coeficient of the element , but depends only on  for all  =  in satisfies the inequalities
0 ≤  () ≤ 1.</p>
        <p>Birnbaum importance considers the relationships between the system performance when
component i is perfect, the system performance when component i fails, and the current system
performance.</p>
        <p>For the BIM-significance indicator, you can get an expression in the following form:
 (, ) = ℎ(1, ) − ℎ(0, )
where ℎ(, ) is reliability function, ℎ(1, ) for absolutely reliable component, ℎ(0, ) for
absolutely unreliable component.</p>
        <p>Other metrics are also used to analyze the significance of elements. The list is presented in
table 1.</p>
        <p>Birnbaum  () reflects the degree of influence of the change in the availability factor of
the element on the change in the availability factor of the system. Risk Decrease () reflects
the importance of maintaining the current level of reliability of this element. Fussell-Vesely
  () reflects the probability that this element is one of the failed elements, provided that the
system failed, the Vesely-Fassel significance characterizes those elements that are most often
involved in system failures. Risk Increase  () reflects the importance of maintaining the
current level of reliability of this element. Criticality Importance () reflects the probability
that this element is critical for the system at a given time.
(4)
(5)</p>
        <p>Many researches have been devoted to computational aspects of significance estimation,
including [7, 8].</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Modeling using R</title>
      <p>R is a programming language for statistical data processing and working with graphics, as well
as a free open-source computing environment for the GNU project. The R language contains
tools that allow you to create multiple parallel threads of calculations (due to simultaneous
loading of several processor cores) and reduce the time spent on modeling several times. To
assess the accuracy of the results obtained, the bootstrap method is proposed. The essence of
the method in this case is that on the basis of one available sample (obtained using the graph
traversal algorithm), a series of pseudo-samples of the same size is formed, consisting of random
combinations of the original set of elements. In this case, the "random selection with return"
algorithm is used, i.e. the extracted element is returned to the original set and has a chance to
be selected again. For each random sample to estimate the probability of failure (or probability
of failure) and thus formed the sample probabilities of system failure (or probability of
failurefree operation), which further evaluated the necessary statistical data (standard deviation or
confidence limits). To calculate the number of iterations and estimate the confidence interval,
a standard approach is used in accordance with [3]. Various techniques can also be used to
improve accuracy, with the most widespread sampling by significance [10].</p>
      <p>To search the graph for paths between certain vertices, use the width traversal algorithm (an
implementation of this algorithm in the iGraph library is used). To generate random numbers
with an exponential distribution law, the basic functions of the R language are used [9].</p>
      <p>Reliability modeling includes  iterations. At each iteration, a random operating time before
failure is generated for each system element (vertex) (this time is generated based on the
specified failure rate of the system element). After that, the elements are sorted in ascending
order of uptime to failure, and the element (vertex) with the lowest time is selected. This vertex
is removed from the graph and the existence of paths between certain vertices of the graph
is checked (between which vertices the presence of a path should be checked is listed in the
description of system failure criteria). If all necessary vertexes are found, the current iteration
continues and the next element in increasing time to failure is selected and the corresponding
vertex is removed from the graph. Next, it checks again whether there are paths between certain
vertexes. If no paths are found between the specified vertexes, the system is considered to have
failed. The failure time  is fixed and a new iteration begins.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Numerical examples</title>
      <sec id="sec-4-1">
        <title>4.1. The dodecahedron graph</title>
        <p>For this experiment, we take the dodecahedron graph (Fig. 1), the model that is widely considered
for a network reliability example. The dodecahedron graph with 20 vertices, 30 edges, and fault
criteria - loss connection of {1, 20}.</p>
        <p>In this model, it is assumed that the vertexes are reliable, and the edges (communication
lines) may fail. The probability of edge failure is constant over time and is equal to . Using the
Monte-Carlo method we estimate the dependence of the probability of failure of the graph on
the probability of failure of each edge. The obtained simulation results are in good agreement
with the results obtained [11].</p>
        <p>The obtained results show a sharp decrease in the probability of connectivity when the
probability of edge failure is less than 10− 1, which correlates well with the presence of one
path of length 3, and two paths of length 4 (by length is meant the smallest number of edges
lying between the vertices {1, 20}).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. SCADA system</title>
        <p>Automated process control system is a group of technical and software solutions designed to
automate technological processes in industrial enterprises. As a rule, automatic process control
systems are understood as a complete solution that ensures the automated execution of the basic
operations of the technological process of production. Components of automatic process control
systems can be separate automatic control systems and automated devices connected in a single
complex. Such as Supervisory control and data acquisition systems (SCADA), distributed control
systems (DCS), emergency protection systems.</p>
        <p>
          SCADA is a complex of equipment, distributed across three levels of the hierarchy, depending
on the functional purpose: upper level: process operator panels; mid-level: server racks, central
computing server, lower level: remote control terminal [12]. The architecture of the process
control system takes into account the requirements for the implementation of the principle of a
single failure and has structural redundancy [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], the structure of the process control system is
shown in Fig. 3.
        </p>
        <p>The system consists of the following units:
1. The Hardware of the main computing resources (S1, S2);
2. Control servers (CS1, CS2) is designed for collecting, processing and storing information
about the operation of system equipment, as well as information interaction.
3. Remote terminal unit (B1-B8) is designed for control field equipment.
4. Four independent Ethernet line;</p>
        <p>The representation in SCADA of the system in the form of a graph and its representation in
the package R is shown in Fig. 4. The standard reliability data (Failure rate: Central control
panel main, redundant - 50 · 10− 6ℎ− 1, Central control unit - 30 · 10− 6ℎ− 1, Remote control
terminal - 20 · 10− 6ℎ− 1, Commutator - 10 · 10− 6ℎ− 1) are considered as initial data [12]. It is
assumed that issues related to the process of ensuring computational reliability are provided by
the necessary capacities [13].</p>
        <p>A system failure is considered to be the loss of communication between the fictitious vertex
F1 (SCADA system operator) and the field equipment control subsystems B1-B8.</p>
        <p>The simulation results of failure probability are presented in Fig. 5
Table 2 shows the reference reliability importance metrics estimated for the SCADA system.</p>
        <p>Increasing the reliability of the elements with the biggest significance will allow achieving
the required failure probability. As a result of the BIM assessment, CPUs of CS unit make the
greatest contribution to system reliability. If further reliability improvements are needed, these
elements should be considered. Possible ways to improve reliability can be considered: the
introduction of continuous monitoring, the choice of more reliable components, using the of
redundancy by reserving.</p>
        <p>According to the results of the Monte Carlo simulation (fig. 5), it can be argued that the
probability of the SCADA functioning in 5000 hours will be no less than 0.992 with a confidence
probability of 0.90. To improve accuracy, methods of reducing the variance of a sample estimate,
for example, the Cross-Entropy Monte-Carlo method [11, 18], can be used.</p>
        <p>The reliability estimated values and importance metrics can be used for system modernization,
in particular when using one of the LKA-LKD algorithms-heuristics [16] for example.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>The reliability models of three-level networks based on the model systems with independent
elements are also considered, a method for assessing reliability is proposed. The Monte Carlo
method is used to estimate failure probability and reliability importance metric with
determination of the confidence interval. To conduct a computational experiment, the R software package
is used. A description is given of the representation of the control system reliability model
in the iGraph package, which provides visualization of the results. R language was primarily
created and is continuing to evolve as a statistical data processing tool. The value of the BIM
metric is determined for further system improvement.
Packet Segmentation (2019) 2019 Wave Electronics and its Application in Information
and Telecommunication Systems, WECONF 2019, art. no. 8840643. doi:
10.1109/WECONF.2019.8840643
[3] Gertsbakh B., Shpungin Y. Models of Network Reliability: Analysis, Combinatorics, and</p>
      <p>Monte Carlo// Boca Raton, FL, USA: CRC, 2009
[4] Andronov A., Jackiva I., Santalova D. Estimation of the Parameters of Continuous-Time
Finite Markov Chain. Distributed Computer and Communication Networks. DCCN 2019.</p>
      <p>Lecture Notes in Computer Science, vol 11965. Springer, Cham
[5] Nikiforov I. Detection and Detectability of Changes in a Multi-parameter Exponential
Distribution. Distributed Computer and Communication Networks. DCCN 2019. Lecture
Notes in Computer Science, vol 11965. Springer, Cham
[6] Birnbaum Z. W. “On the importance of diferent components in a multicomponent system,”
in Multivariate Analysis II. New York, NY, USA: Academic, 1969, pp. 581–592.
[7] Y. Du, S. Si, T. Jin. Reliability Importance Measures for Network Based on Failure Counting</p>
      <p>Process. IEEE Transactions on reliability, vol. 65, no. 1, pp. 267-279
[8] Kamalja K. K., Amrutkar K. P. Reliability and Reliability Importance of
Weighted-r-Within</p>
      <p>Consecutive-k-out-of-n. IEEE Transactions on reliability, vol. 67, no. 3, pp. 951-969
[9] Crawley MJ. The R Book. 2nd ed. Wiley Publishing; 2012.
[10] Blanchet J., Rudoy D. Rare event simulation and counting problems, in Rare Event
Simulation Using Monte Carlo Methods. 1st ed. New York, NY, USA: Wiley, 2009
[11] Vaisman R., Kroese D.P., Gertsbakh I.B. Improved Sampling Plans for Combinatorial</p>
      <p>Invariants of Coherent Systems. IEEE Transactions on reliability, vol. 65, no. 1, pp. 410-424
[12] Moshnikov, Aleksandr and V. S. Kolomoitcev. “Reliability Assessment of Distributed
Control Systems with Network Structure.” 2020 Wave Electronics and its Application in
Information and Telecommunication Systems (WECONF) (2020): 1-4.
[13] Bogatyrev V.A., Bogatyrev S.V., Golubev I.Yu. Optimization and the process of task
distribution between computer system clusters. Automatic Control and Computer Sciences
2012, 46(3), pp. 103-111
[14] Compare M, Bellora M, Zio E. Aggregation of importance measures for decision making
in reliability engineering. Proceedings of the Institution of Mechanical Engineers 2017,
Part O: Journal of Risk and Reliability, 231(3): 242–254
[15] Fang C, Marle F, Xie M. Applying importance measures to risk analysis in engineering
project using a risk network model. IEEE Systems Journal 2017, 11(3): 1548–1556
[16] Kuo W, Zhu X (2012). Importance Measures in Reliability, Risk and Optimization: Principles
and Applications. Chichester: John Wiley and Sons
[17] Moshnikov, A.; Bogatyrev, V. Risk Reduction Optimization of Process Systems
under Cost Constraint Applying Instrumented Safety Measures. Computers 2020, 9, 50.
https://doi.org/10.3390/computers9020050
[18] K-P. Hui, N. Bean, M. Kraetz and D.P. Kroese, The Cross-Entropy Method for Network
reliability estimation, The Annals of Opereations Research, 134:101-118, 2005.
[19] R.Y. Rubinstein and B. Melamed. Modern Simulation and Modeling. Viley series in
probability and Statistics, 1998.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Krieger</surname>
            <given-names>U.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Markovich</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Modeling</surname>
          </string-name>
          and
          <article-title>Reliability Analysis of a Redundant Transport System in a Markovian Environment. Distributed Computer and Communication Networks</article-title>
          .
          <source>DCCN 2019. Lecture Notes in Computer Science</source>
          , vol
          <volume>11965</volume>
          . Springer, Cham
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Bogatyrev</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bogatyrev</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bogatyrev</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          <string-name>
            <surname>Multipath</surname>
          </string-name>
          Redundant Transmission with
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>