<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Estimation of Informativeness of Recognition Signs at Extreme Information Machine Learning of Knowledge Control System</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Sumy State University</institution>
          ,
          <addr-line>Sumy</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1829</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The method of evaluating the informativeness of recognition signs is considered within the framework of information-extreme intellectual technology of machine learning. Evaluation for a student's response to a test task is accepted as a sign of recognition. Optimization of signs vocabulary is carried out directly in the process of machine learning of the knowledge control system. As a criterion for optimization of machine learning parameters, the modified Kulbak information measure is considered. Based on the optimal in the informational sense geometric parameters, that were received in the process of machine learning, of embedded hyperspherical containers of recognition classes, the unimodal deciding rules have been constructed. Verification of the effectiveness of the proposed method was carried out for the task of assessing the informativeness of tests on machine knowledge control of a particular discipline.</p>
      </abstract>
      <kwd-group>
        <kwd>extreme information intellectual technology</kwd>
        <kwd>knowledge control system</kwd>
        <kwd>machine learning</kwd>
        <kwd>information criterion</kwd>
        <kwd>unimodal classifier</kwd>
        <kwd>vocabulary of recognition signs</kwd>
        <kwd>informativeness of recognition sings</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Wide introduction of electronic means of machine knowledge control has conditioned
the relevance of the solution to the problem of adequacy of evaluation of the actual
level of knowledge of the person being examined [1]. One of the key tasks aimed at
solving this problem is assessment of tests informativity [2]. With the development of
computerized systems in education, intelligent information technologies that can
simulate the cognitive processes inherent in a person when making decision
classifications became widely used when solving this task [3]. At the same time, the methods
of artificial neural networks have become the most widespread [4] the content of the
discipline. The analysis of existing data mining methods shows that they are mostly of
a model nature and do not sufficiently take into account the abovementioned
conditions of the real educational process.</p>
      <p>One of the promising directions of increasing the functional efficiency of machine
knowledge control systems is application of ideas and methods of the so-called
extreme information intellectual technology (IEI-technology) data analysis, which is
based on maximizing the information capacity of the system in the process of
machine learning [5, 6].</p>
      <p>The article considers the method of information-extreme machine learning of
computerized knowledge control system with optimization (here and further in the
informational sense) of geometric parameters of hyperspherical containers of classes of
recognition with the built-in structure.</p>
      <p>2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem statement</title>
      <p>Let's consider in the framework of IEI-technology the formalized statement of the
problem of information synthesis of adaptive knowledge control system with
optimization of the recognition signs vocabulary. Evaluation of the response to a test, which
is calculated by the evaluation function on the scale of 0 to 100, is taken for a sign of
recognition. Let's imagine we have the character set {X mo | m  1, M } of classes of
recognition, which characterize different levels of student knowledge, and the
threedimensional learning matrix of student responses to test assignments
|| ym(j,)i ||, i  1, N, j  1, n , are given, where N is number of recognition signs, which
is determined by the number of tasks in the test; n is the number of
vectorimplementations of classes of recognition, which ensures the representativeness of the
training sample. With that, the line of the learning matrix {y m(j,)i | i  1, N} defines the
j vector-implementation of the recognition class, and column {y m(j,)i | j  1, n}
determines a training sample of the values of the i sign. In addition, the structured
vector of machine learning parameters of the knowledge control system is known
g m  R1m , R2m , , |N| , m  1, M ,
where R1m is inner radius of container of class X mo ; R2m is external radius of
container of class X mo ;  is a parameter that determines the lower control attribute for
recognition signs; |N | in the informational sense, is redundant initial vocabulary of
power N .</p>
      <p>
        In the process of machine learning of the knowledge control system, it is necessary
to determine the optimal values of the coordinates of the vector (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), which provide the
maximum average of the character set recognition classes of the information criterion
for learning parameters optimization.
      </p>
      <p>*
E 
1 M</p>
      <p>
         max E m(k ) ,
M m1 GE {k}
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
where E m(k ) is information criterion, calculated on the k -th step of machine learning,
for optimizing the parameters of the knowledge control system functioning when
recognizing implementation of the class X mo ; GE is working (admissible) area of
determining the function of the information criterion; {k} is an ordered set of steps in
machine learning. With functioning knowledge control system in the examination
mode, that is, directly in the machine knowledge control according to test tasks, it is
necessary to determine the affiliation of the vector-implementation,
developed-byoptimal-dictionary, to one of the recognition classes from given character set {X mo } .
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Mathematical model</title>
      <p>Since the process of knowledge evaluation is poorly formalized, the mathematical
model of machine learning of the knowledge control system with the optimization of
the vocabulary of recognition signs will be presented as a categorical model. In this
case, the structure of the input mathematical description of the knowledge control
system has the form
 B  G, T , , Z, Y , X ; f1, f 2 
where G is input factors space; T is a set of moments of information removal time;
 is recognition signs space; Z is space of given levels of knowledge; Y is
sampling population – input learning matrix; X is binary learning matrix; f1  is
operator of the input learning matrix Y formation; f2 is is operator of forming a working
binary learning matrix X .</p>
      <p>Fig. 1 in the form of a directed graph, shows a categorical model of machine
learning of the knowledge control system with optimization of the recognition signs of
vocabulary.</p>
      <p>
        In Fig. 1, Cartesian product G T    Z sets the universe of tests. Operator 
reflects the vectors-implementations of the binary learning matrix X partitioning
~ |M | of the space of signs into recognition classes, which is generally fuzzy. Operator
 checks the main statistical hypothesis about affiliation of the
vectorimplementation x m(j) to the corresponding recognition class. As a result of statistical
verification, a set of hypotheses is formed I |L| , where L is the number of statistical
hypotheses, and operator  forms a set of precise characteristics |Q| where Q  L2 .
Operator  calculates the value E of the information criterion for optimization of the
learning parameters, and operator r reconstructs the recognition class containers that
are built in the radial basis of the signs space at each step of machine learning.
Category model, shown in Fig. 1, has additional contours for optimizing machine learning
parameters (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ). The circuit, which is closed through a term-set D, a system of control
tolerances, optimizes control tolerances for recognition signs. The following contour
optimizes the vocabulary  of recognition signs. Operator 1 : E   removes
recognition signs from the vocabulary, creates new vocabularies and determines the
informative nature of the recognition signs. In this case, a sign is considered
informative if the value of the information criterion decreases at its removal. If the value of
the criterion does not change, then the sign is considered non-informative and after
the completion of machine learning the system is removed from the vocabulary. With
the help of the operator 2 :    the dimensionality of the recognition signs
varies and, accordingly, a new learning matrix Y is formed. Operator u regulates the
process of machine learning of the knowledge control system.
      </p>
      <p>Thus, the proposed categorical model is essentially a generalized structural scheme
of the information-extreme machine learning of the knowledge control system with
optimization of a set of test tasks.</p>
      <p>4</p>
    </sec>
    <sec id="sec-4">
      <title>Machine learning algorithm</title>
      <p>
        According to the categorical model (Fig. 1), information-extreme machine learning
with optimization of the vocabulary of recognition signs is carried out according to
the tricyclic procedure for finding the global maximum of the information criterion
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) for optimizing the machine learning parameters (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) in the form
*  arg max{max{ max
{k} G GE {d}
      </p>
      <p>
        E
(k )
(d )}},
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(k )
where E
      </p>
      <p>
        (d ) is the averaged by the character set of recognition classes of the
information criterion value (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), calculated on the а k -th step of machine learning;
G is the permissible area of determination of the parameter  of the control
tolerance fields for the recognition signs; {d} is a set of values of the radii of the
recognition classes containers.
      </p>
      <p>
        The internal cycle of the algorithm (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) implements the basic algorithm of machine
learning, which optimizes the geometric parameters of the recognition classes
containers. For an enclosed hyperspherical container, for instance, the recognition class
X mo the optimization parameters are the internal radius R1m , closest to the center of
the scattering of the recognition class vectors, and the outer radius R2m . The input
data for the basic machine learning algorithm is the three-dimensional array of the
input learning matrix {y m(j,i) | m  1, M}; i  1, N; j  1, n}, recognition signs of which
form the initial vocabulary |N| . Besides, the N is dimensional unit vector x0 is
given, the vertex of which determines the center of scattering of
vectorsimplementations of all recognition classes.
      </p>
      <p>Let's consider the main stages of implementation of the basic machine learning
algorithm:
1. Formation of an array of binary learning matrix {x m(j,)i } , the elements of which are
1, if
determined by the rule: xm(j,i)  
 0, if else.</p>
      <p>
        AHK ,i  y m(j,i)  ABK ,i ;
2. Initialization of the counter of recognition classes: m  0 ;
3. m  m  1;
4. Initialization of the counter of the external radius optimization steps R2m  0;
5. R2 m  R2 m 1;
6. Computing the information criterion (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) and finding its maximum value E m* the
working (admissible) function domain;
      </p>
      <p>*
7. If (Em  Em ) &amp; (R2m  N ) , then see point 5, otherwise - point 8;
8. If E m </p>
      <sec id="sec-4-1">
        <title>Em (R2), then R2 m assign the optimal value R2*m ;</title>
        <p>max</p>
        <p>
          GE GR2
9. If m  M , then see point 3, otherwise - point 10, since the external radii of the
containers of all classes of recognition are determined;
10. Initialization of the class of recognition counter: m  0 ;
11. m  m  1;
12. Initialization of the counter of internal radius optimization steps: R1m  0 ;
13. If m  1, then R11*  0, otherwise point 14 is implemented;
14. R1m  R1m 1;
15. Calculation and search for the maximum value of the information criterion (
          <xref ref-type="bibr" rid="ref3">3</xref>
          );
16. If (Em  E m* ) &amp; (R1m  R2 m ) , then see point 14, otherwise - point 17;
17. If E m 
        </p>
        <p>max</p>
        <p>GE GR1
18. If m  M , then point 11 is implemented, otherwise point 19;
19. STOP.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Em (R1), then R1m assign the optimal value R1*m ;</title>
        <p>As a criterion for optimizing machine learning parameters in IEI-technology, Shannon
entropy measure and the Kulbak information measure are mainly used, the
modification of which is as follows [6]:
</p>
        <p>
          
 D k   D k  10 r 
  k   k   10 r D k   D k    k   k  ,
E mk   0,5 log 2  1 2 
1 2
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
where D k  is the first and second authenticity of classification decisions, calculated
1
at k -th step of machine learning; D k  is second authenticity;  k  is error of the
2
first kind;  k  is error of the second kind; 10 r is a sufficiently small number to
avoid division into zero (in practice, the value of r s selected from the interval
1  r  3) .
        </p>
        <p>
          Thus, the basic function of the basic learning algorithm in the framework of the
IEI-technology is computation of the information criterion at each step of learning,
organization of finding a global maximum in the work area defining the function of
the criterion and determining the optimal geometric parameters of the classes of
recognition containers. To optimize the recognition vocabulary, let's apply a classical
scheme of sequential hereditary selection. Let's consider the main stages of the
algorithm implementation for evaluating the informativeness of recognition signs in the
process of information-extreme machine learning by parallel optimization of control
tolerances on recognition signs:
1. Initialization of the counter of changing steps of the control tolerances field
parameter on the recognition signs:   0 ;
2.     1 ;
3. Lower AHK control attestation for recognition of recognition: AHK  ABK  ,
where ABK is upper control tolerance equal to the maximum value of the rating
scale.
4. Initialization of the counter of recognition signs which are deleted from the
vocabulary in the current step and returned again in the next step: i  0 ;
5. The basic algorithm is implemented according to the above mentioned scheme
and the averaged maximum value of the information criterion (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) is calculated by
the character set of the recognition classes for the initial signs vocabulary;
*
6. If E i  E i , then point 2 is implemented, otherwise point 7;
7. i  i 1 ;
8. If i  N , then point 5 is implemented, otherwise point 9. Informative recognition
signs for which the condition is fulfilled, are identified
Pi  E ALL  Ei*  0
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
where E ALL is the maximum average value of the information criterion for
optimizing machine learning parameters for the initial signs recognition vocabulary;
9. STOP.
        </p>
        <p>Based on the optimal geometric parameters of embedded hyperspherical containers of
recognition classes, received in the process of machine learning, unimodal deciding
rules, which, unlike polymodal, are characterized by a single center of scattering of
vectors-implementations of all recognition classes, are built. In this case, the deciding
rules for determining the implementation of x ( j) , in the mode of examination, for
example, class X mo have the form
(x( j)  ~|M |)(X mo  ~|M |){if R1m  d[x0  x( j) ]  R2m
then
x ( j)  X mo</p>
        <p>else x ( j)  X mo },
where d[x0  x ( j) ] is code distance between implementation x ( j) and a single
vector-implementation x0 , the vertex of which defines the general center of scattering of
vectors-implementations of all recognition classes. Thus, information-extreme
machine learning consists in purposeful approximation of the information criterion for
optimizing the parameters of the system's operation to its maximum limit value.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Simulation results</title>
      <p>
        Implementation of the above algorithm of machine learning was carried out on the
example of estimating the informativeness of the tests in the discipline "Intelligent
Systems", which is offered to students of Sumy State University (Sumy, Ukraine) in
the specialty "Computer Science". The character set consisted of four classes of
recognition: class X1o characterized the level of knowledge "excellent", class X 2o –
“good”, class X 3o – “satisfactory” and class X 4o – “poor”. The answer for each of the
tests was evaluated on a 100 point scale. The complete initial vocabulary consisted of
28 recognition signs, which were equal to the number of tasks in the test. The number
of vector-implementations for each of the recognition classes was 40, that is, to form
the input educational matrix. The maximum average value of the information criterion
calculated in the process of machine learning according to formula (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) was used to
evaluate the informativeness of the tests.
      </p>
      <p>
        In Fig. 2, shaded sections of the graph denote the work areas in which the search
for a global maximum of the information criterion for optimizing machine learning
parameters is is carried out. Analysis of Fig. 2 shows that the optimal value of the
parameter of the control tolerances field is equal to *  36 gradations of the
100*
point grading scale at the maximum value of the averaged criterion. E
 1,35 . With
the optimal parameter of the control tolerances fields on the recognition signs in
accordance with procedure (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ), the following optimal radii of the recognition class
containers were obtained: for container recognition class X1o («excelent») internal radius
was equal to R11  0 and an external radius – R21  13 (here and further in code units
from the geometric center of the recognition class), for container recognition class
X 2o («good») the radii were equal respectively R12  12 and R22  15 , for container
recognition class X 3o («satisfactory») – R13  15 and R23  17 , for container
recognition class X 4o («poor») – R14  17 and R24  22.
E1*  1,93 at the first authenticity D1*,1  0,95 and error of the second kind 1*  0,23 ,
for class X 2o is E2*  1,18 ( D1*, 2  0,72 , *2  0,13 ), for class X 3o is E3*  0,87 (
D1*,3  0,78 , *3  0,25 ) and for class X 4o is E4*  1,43 ( D1*,4  0,80 , *4  0,15 ).
According to the algorithm (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ), the evaluation of the informative nature of the tasks in
the test was carried out by forming at each step of the machine learning new versions
of vocabularies for which the optimization of the learning parameters was conducted
and the maximum average value of the information criterion was carried out.
      </p>
      <p>
        Fig. 3 shows the dynamics of the change of the averaged information criterion (
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
in the process of evaluating the informativeness of the recognition signs.
      </p>
      <p>Analysis of Fig. 3 shows that for the knowledge control system on the relevant
discipline of twenty eight signs of information recognition, there is twenty one sign,
which are indicated by light grey color. Sign 1 is non-informative because its deletion
does not change the information criterion. Other signs (9, 11, 19, 22, 26 and 28) refer
to interfering (disinfecting) ones, indicating the need for the processing of appropriate
tasks in the test or their removal.</p>
      <p>The most informative is the sign 17. Deleting this sign from the vocabulary leads not
only to a decrease in the maximum value of the information criterion, but also to
change of the values of the optimal geometric parameters of the recognition classes
container. The most disturbing (disinfecting) is the sign 19, the informativeness of
which is equal to P19   0,15 . Removing this sign from the vocabulary also changes
the optimal values of the geometric parameters of the classes of recognition
containers and the exact characteristics of machine learning. But unlike in the previous case,
the tendency of these indicators is the opposite.</p>
      <p>
        Fig. 4 shows the graph of the averaged information criterion dependence (
        <xref ref-type="bibr" rid="ref6">6</xref>
        ) on the
parameter A of the control tolerance field on the recognition signs obtained in the
course of machine learning with removing the most informative sign of recognition
17 and the sign 19 that interferes.
maximum value of the information criterion, which is equal to E  1,45 . In
addition, we received a new optimal value of the control tolerance field parameter, which
was equal to *  32 . In this case, the radii of the classes of recognition containers
have not changed except for the inside radius of the class container X 2o , which
became equal to R12  13 .
      </p>
      <p>6
In this work, the difference between the maximum values of the information criteria
of machine learning functional efficiency with the full set of initial tests and a set with
the withdrawn test is considered as a measure of informativeness of the test. A test,
the removal of which led to a decrease in the maximum value of the information
criterion, was accepted as informative. As a result, 21 informative tests were selected out
of 28 initial tests, for which the input learning matrix was formed according to the
results of machine knowledge control. The use of informative tests set allowed to
increase the maximum value of the information criterion (Fig. 4) in comparison with
the use of the initial set (Fig. 2), which confirms the perspective of the proposed
method.</p>
      <p>The paper proposes an information-extreme method for evaluating the
informativeness of test tasks for machine knowledge control. By eliminating from the initial set
of non-informative and disinforming tests, the functional efficiency of machine
learning of the knowledge control system is improved. Further improvement of the
proposed method is to increase the depth of machine learning, which is determined by the
number of contours of optimizing the parameters of the knowledge control system
functioning in the categorical model.</p>
      <p>7
8</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Menqui</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yun</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ying</surname>
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Machine Learning and Modern Education</article-title>
          . In: 4th International Conference, eLEOT
          <year>2018</year>
          , Shanghai, China, pp.
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Lemaître</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Current developments in the French engineering education system</article-title>
          .
          <source>In: European Journal of Engineering Education</source>
          ,
          <volume>42</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>131</fpage>
          -
          <lpage>141</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Zgurovskii</surname>
            ,
            <given-names>M. Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pankratova</surname>
            ,
            <given-names>N. D.</given-names>
          </string-name>
          : Sistemnyi analiz: problemy, metodologiya, prilozheniya, Kiev: Izdatelstvo Naukova dumka (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>F. T. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>E. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          :
          <article-title>The Art of Learning: Neural Networks and Education</article-title>
          , CRC, PressINC (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dovbysh</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martynenko</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalenko</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budnyk</surname>
            ,
            <given-names>M. M.:</given-names>
          </string-name>
          <article-title>Informationextreme algorithm for recognizing current distribution maps in magnetocardiography</article-title>
          .
          <source>In: Journal of Automation and Information Sciences</source>
          ,
          <volume>2</volume>
          (
          <issue>43</issue>
          ), pp.
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dovbysh</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Budnyk</surname>
            ,
            <given-names>M. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moskalenko</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          :
          <article-title>Information-Extreme Algorithm for Optimizing Parameters of Hyperellipsoidal Containers of Recognition Classes</article-title>
          .
          <source>In: Journal of Automation and Information Sciences</source>
          , New York: Begell House Inc.,
          <volume>10</volume>
          (
          <issue>44</issue>
          ), pp.
          <fpage>35</fpage>
          -
          <lpage>44</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>