<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Moscow, Russia
" nak@dscs.pro (N.A. Kharitonov); alt@dscs.pro (T. Alexander)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Local Parameter Training of Algebraic Bayesian Networks: Conjugate Distributions and Expert Knowledge With Uncertainty</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nikita A. Kharitonov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tulupyev Alexander</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Laboratory of Theoretical and Interdisciplinary Problems of Informatics, St. Petersburg Institute for Informatics and Automation of the Russian Academy of Sciences</institution>
          ,
          <addr-line>St. Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>St. Petersburg State University</institution>
          ,
          <addr-line>St. Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>In the work the local parametric training of Algebraic Bayesian networks is considered. The theorem about the change of Dirichlet distribution parameters during transition from a priori to a posteriori probability distribution on propositional quantum formulas is formulated and proved. The proof is based on the conjugation property of the multinomial and Dirichlet distributions. One of the main points in machine learning models research is the training of the model parametric synthesis and structural synthesis. There is a complex mathematical apparatus that determines the correctness of operations behind specific algorithms and calculations. This is also correct for Algebraic Bayesian networks. The object of the research is approach to local parametric training of Algebraic Bayesian network, what is training of network represented by knowledge pattern. The a priori distribution of probabilities presented in a knowledge pattern is studied and approaches to the formation of the a posteriori distribution resulting from learning are considered.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;machine learning</kwd>
        <kwd>probabilistic graphical models</kwd>
        <kwd>Algebraic Bayesian networks</kwd>
        <kwd>parametric training</kwd>
        <kwd>Dirichlet distribution</kwd>
        <kwd>multinomial distribution</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        Neural networks [
        <xref ref-type="bibr" rid="ref1">1, 2, 3, 4, 5</xref>
        ] and probabilistic graphical models are one of the machine
learning models. Belief Bayesian networks, [6, 7, 8], Markov networks [9, 10, 11] and Algebraic
Bayesian networks [12], which are the subject of the research, belong to the class of
probabilistic graphical models.
      </p>
      <p>Approach for generating the structure of Bayesian network basing on con-straint-based,
assessment-based and search-based methods is described in [6]. The use of Belief Bayesian
networks for prediction of students expulsion is the subject of research [7]. The detection and
prediction of emergencies on atomic power plants by Bayesian network was studied in [8].</p>
      <p>One of the main ideas during the work with probabilistic graphical models is their
decomposition in smaller parts, which describe some information about the object [13]. These parts are
knowledge patterns in Algebraic Bayesian network theory [12]. Fig. 1 presents an example of
one of the possible Algebraic Bayesian network graphical representation: the non-directional
graph with knowledge patterns in nodes.</p>
      <p>The representation of knowledge pattern in the form of proposal formulas-quantum set is
used in the context of this research [12], moreover, each propositional formula has scalar or
interval probabilities estimate in N.Nilsson interpretation [14]. Other possible representations
of knowledge pattern are ideals of conjuncts or disjuncts [12]. The work [15] has description
of transition matrix between the set of propositional-formulas-quanta, ideal of conjuncts and
ideal of disjuncts.</p>
      <p>Main operations of probabilistic-logical inference during the work with Algebraic Bayesian
networks are consistency maintaining, a priori and a posteriori inference [16].</p>
      <p>Consistency maintaining is verification of correctness of probability estimates which are
presented in the network and their changing if it is necessary and possible [17].</p>
      <p>A priori inference is receipt of probability estimates of some propositional formula based on
the estimates presented in the network [12].</p>
      <p>A posteriori inference is the process of changing estimates in network basing on some
information, which is presented as evidence. Also the evidence probability is calculated in this
process. The study [18] describes the sensitivity of local a posteriori inference.</p>
      <p>When working with Algebraic Bayesian networks it is necessary to receive the
probabilities of elements in network. The research [19] provides an approach for obtaining algebraic
Bayesian network with interval estimates basing on data set with missing values. The approach
for receiving quantum probabilities with scalar estimates is investigated in [13] is the subject
of this research.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Local parameter training of Algebraic Bayesian networks</title>
      <p>Tasking of local parametric training of algebraic Bayesian networks is proposed in the
paper [13].</p>
      <p>Examples of incomplete (with missing values) implementations(the missing value is marked
with ?):</p>
      <p>It is possible that line in input data set is presented as the set of atom variables with 0, 1 or
missing value. That case can be easily converted to the quanta. The example for several lines
is presented in table 1.</p>
      <p>When formalizing local parametric machine training with Dirichlet distribution (single and
conjugated distributions) the first step of the training is the creation of "a priori" knowledge
pattern with equal quanta probabilities. Application of such uniform distribution on quanta
corresponds to the implicit hypothesis that in conditions of complete uncertainty we can not
give preference to any quantum.</p>
      <p>Although this is the subject of another study, it should be noted that this hypothesis is not
always true. For example, expert knowledge expressed as incomplete, inaccurate,
nonnumerical information may afect the choice of a priori distribution. It may become known from an
expert that features of a subject area are such that  ( 1) ≤  ( 2),  ( 2) ≥ 0.75. The information
obtained (in fact, initially - knowledge obtained from the Expert Advisor with uncertainty) may
significantly afect the choice of a priori distribution.
in alphabet, under which knowledge pattern is built.</p>
      <p>According to the definition of quanta probability, the probability of each element will be 1/,
where  is the number of quanta in knowledge pattern,  = 2 , where  is a number of atoms</p>
      <p>During the first step of the algorithm the knowledge pattern is trained on a part of data set
without missing values.</p>
      <p>During the second step the part of data set with missing values is used. On each step the
quantum with missing values is converted to the set of quanta without missing values, which
are not contradictory to the given. That means that these quanta are acquired from given
by all possible replacements of missing value by 0 or 1. Basing on the received set of quanta
the probabilities in the knowledge pattern are changed: for those quanta, which are in set,
the probabilities increase with the preservation of their ratio, for other they proportionally
decrease.</p>
      <sec id="sec-3-1">
        <title>Theorem</title>
      </sec>
      <sec id="sec-3-2">
        <title>Theorem</title>
      </sec>
      <sec id="sec-3-3">
        <title>Proof</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Probability distribution</title>
      <p>After the training of knowledge pattern by proposed method the probabilities of quanta in
knowledge pattern will be described by the Dirichlet distribution.</p>
      <p>Approaches to the analysis of what happens at receipt of additional data with conjugated
distributions are considered deeply enough in modern probability and machine learning theories;
in the proof we will use receptions from[20].</p>
      <p>The probabilities of quanta in "a priori" knowledge pattern can be described by Dirichlet
distribution with parameter vector set as unit vector. The probability density function in that
case is a gamma function with knowledge pattern dimensional as parameter:
Dir(  ) =</p>
      <p>Γ(∑ =0   )</p>
      <p>∏ =0 Γ(  )  =0
∏ p(  )</p>
      <p>−1 =
=

Γ( )
∏ =0 Γ(1)  =0</p>
      <p>∏ p(  )0 = Γ( ),
distribution parameter vector.
where   is the quanta vector with  dimensional, Γ – gamma function and {  } = {1} is the</p>
      <p>Let us take a look at the data set. Every line of data set (which is presented as quantum 
with possible missing values) can be converted to the vector of  dimension by next rule:
• this is the 1 value on position , if quantum   is not contradict to the  ;
• this is the 0 value in other cases.</p>
      <p>Index of quantum is its binary representation converted to decimal number system.</p>
      <p>This interpretation of data set allows to consider each input line, corresponding to quantum
without missing values, as random experiment, the result of which is the one of the quanta.
The 1 value in the relevant vector shows, which quantum is the result.</p>
      <p>Note, that each of this experiments is independent on others, as data set has no information
about the relationship between trials. Thus, the part of data set without missing values can be
considered as    independent experiments with  possible results. That means that it can be
described by multinomial distribution.</p>
      <p>Rows corresponding to quanta with missed values need to be considered separately, since
for such rows the number of outcomes is more than one, what is impossible in a multinomial
distribution.</p>
      <p>Each such line  can be converted to the set of vectors { } of length , such so |  | = 1∀ and
∑   = . This set of vectors can be similarly presented as the set of independent experiments,
described by multinomial distribution.</p>
      <p>Thus, on each step the likelihood function on each step has multinomial distribution.
According to the Bayes theorem the  step of training is:
p ( | ) =</p>
      <p>p ( | )p ( )
∫ p ( | )p ( )
,
where p ( ) is a priori distribution, which is equal to the a posteriori distribution for each  ≥ 1.
For  = 0 the a priori distribution is Dirichlet distribution: p0( ) = Dir(  ).</p>
      <p>The Dirichlet distribution is conjugate prior to the multinomial distribution, so, a posteriori
distribution on each set is Dirichlet distribution.</p>
      <p>Thus, the result of training will be the knowledge pattern with quanta probabilities which
have Dirichlet distribution.</p>
      <sec id="sec-4-1">
        <title>The theorem is proved</title>
        <p>Let us consider the knowledge pattern   built over the atoms  1,  2. In case there is no
information about the subject area, the a posteriori distribution of quanta will set their
probabilities equal to 0.25. However, some information which is known before training, expressed,
for example, in expert estimates, may afect the a priori distribution. Let p( 1) ≤ 0.5. In this
case the a priori probability of quanta is: p(00) = p(10) = 0.375, p(10) = p(11) = 0.125.</p>
        <p>It should be noted that initial constraints have only limited the class of probability
distributions from the whole simplex of possible distributions and narrowed interval estimates. The a
priori distribution was chosen as the mass center of the resulting figure in hyperspace.
Generally speaking, when considering a priori distribution issues, there are tasks of constructing a
canonical representative — a knowledge pattern with scalar estimates — for an initial
knowledge pattern with interval estimates. In addition, the gamma distribution parameters can be
affected by the "uncertainty level" of the initial knowledge pattern — the extent to which possible
distributions of probabilities (compatible with the knowledge pattern) relative to the
canonical representative vary. However, both formalization of this remark and consideration of the
degree of uncertainty are tasks of other studies.</p>
        <p>Let us suppose that the training resulted in a knowledge pattern with a probability vector of
quanta   . In works[12, 13] the approach to obtaining the probabilities of conjuncts   and the
vector of probability of disjuncts   is described. This approach allows to convert the result of
training proposed above to the vector of conjuncts or disjuncts.</p>
        <p>In order to do this, we define the matrices   and   ∶</p>
        <p>1 = [01 11] ,   +1 = [     ] ,
 1 = [11 01] ,   +1 = [ ◦    ] ,
  =  ×   ;
  =  ×   .
where  is a matrix with all elements equal to 1,  ◦ is received from   by replacing first row
elements with 0.</p>
        <p>The probabilities of conjuncts and disjuncts can be calculated in the next way:</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The study covers local parametric learning of Algebraic Bayesian networks. The training
process is described and the theorem on the probabilities of quanta in knowledge pattern having
Dirichlet distribution is proved.</p>
      <p>Further researches in this direction are studying the distribution of probabilities for a
knowledge pattern represented as an ideal of conjuncts or disjuncts, as well as researching of the
families of distributions received in training with the expert’s knowledge, for example training
with obtaining interval estimates.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The work was carried out with the financial support of the Russian Foundation for Basic
Research (project No. 18-01-00626) and within the framework of the project on the state
assignment of SPIIRAS No. 0073-2019-0003.
[2] A. M. M.F. Kazemi, M.A. Pourmina, Novel neural network based ct-nsct
watermarking framework based upon kurtosis coeficients, Sensing and Imaging 21 (2020) 7.
doi:NovelNeuralNetworkBasedCT-NSCT.
[3] Y. Li, J. Xiang, Existence and global exponential stability of anti-periodic
solutions for quaternion-valued cellular neural networks with time-varying delays,
Advances in Diference Equations 2020 (2020) 47. doi: https://doi.org/10.1186/
s13662-020-2523-4.
[4] Y. Sidi, S. Selouani, B. Zaidi, A. Bouchair, Improving dysarthric speech recognition
using empirical mode decomposition and convolutional neural network, Eurasip Journal
on Audio, Speech, and Music Processing 2020 (2020) 1–7. doi:https://doi.org/10.
1186/s13636-019-0169-5.
[5] M. Stimberg, D. Goodman, T. Nowotny, Brian2genn: accelerating spiking neural network
simulations with graphics hardware, Scientific reports 10 (2020) 1–12.
[6] J. Dai, J. Ren, W. Du, V. Shikhin, J. Ma, An improved evolutionary approach-based
hybrid algorithm for bayesian network structure learning in dynamic constrained search
space, Neural computing &amp; applications 32 (2020) 1413–1434. doi:https://doi.org/
10.1007/s00521-018-3650-7.
[7] D. Delen, K.Topuz, E. Eryarsoy, Development of a bayesian belief network-based dss
for predicting and understanding freshmen student attrition, European journal of
operational research 281 (2020) 575–587. doi:https://doi.org/10.1016/j.ejor.2019.
03.037.
[8] Y. Zhao, J. Tong, L. Zhang, G. Wu, Diagnosis of operational failures and on-demand
failures in nuclear power plants: An approach based on dynamic bayesian networks, Annals
of nuclear energy 138 (2020) 107181.
[9] G. Masetti, L. Robol, Computing performability measures in markov chains by means of
matrix functions, Journal of Computational and Applied Mathematics 368 (2020) 112534.
[10] M. Shi, Y. Tang, X. Zhang, Y. Zhang, J. Xu, Modeling and simulation of packet delivery
rate in lte-v network based on markov chain, Tsinghua Science and Technology 25 (2020)
357–367. doi:https://doi.org/10.26599/TST.2018.9010142.
[11] Z. Wang, W. Yang, Markov approximation and the generalized entropy ergodic
theorem for non-null stationary process, Proceedings of the Indian Academy of
Sciences: Mathematical Sciences 130 (2020) 13. doi:https://doi.org/10.1007/
s12044-019-0542-4.
[12] A. Tulupyev, S. Nikolenko, A. Sirotkin, Bayesian Belief Networks: Probabilisticlogic
Approach [in Russian], SPb.: Nauka, Saint-Petersburg, Russia, 2006.
[13] A. Tulupyev, A. Sirotkin, S. Nikolenko, Bayesian Belief Networks [in Russian], SPbSU</p>
      <p>Press, Saint-Petersburg, Russia, 2009.
[14] N. Nilsson, Probabilistic logic, Artificial Intelligence 28 (1986) 71–87.
[15] E. Malchevskaya, A. Zolotin, A. Tulupyev, Algorithms of the a posteriori inference in the
algebraic bayesian networks: refining of the matrix-vector representation (in russian),
in: In: Fuzzy systems and soft computing. Industrial applications. (FTI-2017), 2017, pp.
376–388.
[16] A. Zolotin, E. Malchevskaya, N. Kharitonov, A. Tulupyev, Local and global
logicalprobabilistic inference in the algebraic bayesian networks: matrix-vector description and
the sensitivity questions (in russian), Fuzzy systems and soft calculations. Tver: TvGTU
(2017) 133–150.
[17] N. Kharitonov, E. Malchevskaia, A. Zolotin, M.Abramov, External consistency
maintenance algorithm for chain and stellate structures of algebraic bayesian networks:
Statistical experiments for running time analysis, in: Proceedings of the third international
scientific conference intelligent information technologies for industry (IITI’18), Springer,
Cham, 2019, pp. 23–30.
[18] A. Zolotin, A. Tulupyev, Sensitivity statistical estimates for local a posteriori inference
matrix-vector equations in algebraic bayesian networks over quantum propositions,
estnik St. Petersburg university-mathematics 51 (2018) 42–48. doi:https://doi.org/10.
3103/S1063454118010168.
[19] N. Kharitonov, A. Maximov, A. Tulupyev, Algebraic bayesian networks: Naïve frequentist
approach to local machine learning based on imperfect information from social media
and expert estimates, Communications in Computer and Information Science 1093 (2019)
234–244.
[20] M. DeGroot, Optimal Statistical Decisions, McGraw-Hill, 1970.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>R. D'souza</surname>
            , P.-Y. Huang,
            <given-names>F.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Yeh</surname>
          </string-name>
          ,
          <article-title>Structural analysis and optimization of convolutional neural networks with a small sample size</article-title>
          ,
          <source>Scientific reports 10</source>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          . doi: https: //doi.org/10.1038/s41598-020-57866-2.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>