<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Digital Content Processing Method for Biometric Identification of Personality Based on Artificial Intelligence Approaches</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cherkasy State Technological University</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cherkasy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ukraine</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>t.utkina</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>k.rudakov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>s.mitsenko</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>m.chychuzhko}@chdtu.edu.ua</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>fedorovee</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@ukr.net</string-name>
          <email>ineks-kiev@ukr.net</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>E. O. Paton Electric Welding Institute</institution>
          ,
          <addr-line>Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1899</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The paper suggests a method for processing digital content for biometric identification based on artificial intelligence approaches. To get the goal the methods of forming digital content characteristics, creating a structure model of a system for processing digital content, the method of selecting the structure determination of parameter values of the mathematical model of digital content processing system are suggested. The suggested characterization of digital content automates the processing of digital content which increases the accuracy and speed of determining the values of signs. The suggested creation of a model structure of a digital content processing system provides knowledge in the form of easily accessible for human understanding rules that simplifies the process of determining the structure of the system and also allows parallel processing of information that allows increasing the learning speed. The suggested selection of structure method of determining values of model parameters of the processing system of the digital content based on the genetic algorithm uses a combination of directed and random search that decreases the probability of a hit in local extremum and provides an acceptable speed of determining values of the model parameters. The suggested method of digital content processing for biometric identification of a personality by voice can be used in various intelligent digital content processing systems.</p>
      </abstract>
      <kwd-group>
        <kwd>digital content processing</kwd>
        <kwd>biometric identification of personality</kwd>
        <kwd>artificial neural network</kwd>
        <kwd>fuzzy inference systems</kwd>
        <kwd>genetic algorithm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Human-machine interfaces are one of the directions of digital content processing. For
these interfaces, biometric identification of a person is important.</p>
      <p>
        Automated biometric identification of a person means decision making based on
acoustic and visual information, which improves the quality of recognition of the
person being studied [
        <xref ref-type="bibr" rid="ref1">1-3</xref>
        ]. Unlike the traditional approach, computer biometric
identification speeds up and improves the accuracy of the recognition process, which is
especially critical in limited time conditions.
      </p>
      <p>A special class of biometric identification of a person is formed by methods based
on the analysis of acoustic information [4-8].</p>
      <p>The methods of biometric identification of a person by voice include: dynamic
programming [9, 10]; vector quantization [11, 12]; artificial neural networks [13, 14];
decision tree [15]; Gaussian mixture models (GMM) [16-19]; their combination [20].</p>
      <p>Artificial neural networks are the most popular methods.</p>
      <p>The advantages of neural networks consist in: the possibility of their training and
adaptation; the ability to identify patterns in the data, their generalization, i.e.
extracting knowledge from data, therefore, knowledge about the object is not required (for
example, its mathematical model); parallel processing of information, which increases
the computing power.</p>
      <p>The disadvantages of neural networks include: the difficulty of determining the
network structure, since there are no algorithms for calculating the number of layers
and neurons in each layer for specific applications; the difficulty of forming a
representative sample; a high probability of a learning method and adaptation getting into a
local extremum; inaccessibility for human understanding of knowledge accumulated
by the network (it is impossible to present the relationship between output and output
in the form of rules), since it is distributed between all elements of the neural network
and is presented in the form of its weighting coefficients.</p>
      <p>Recently, neural networks have been combined with fuzzy inference systems.</p>
      <p>The advantages of fuzzy inference systems are the following: presentation of
knowledge in the form of rules that are easily accessible for human understanding; no
accurate assessment of variable objects is needed (incomplete and inaccurate data).</p>
      <p>The disadvantages of fuzzy inference systems include: the impossibility of their
training and adaptation (parameters of the membership functions cannot be
automatically configured); the impossibility of parallel processing of information, which
increases the computing power.</p>
      <p>Since genetic algorithms can be used instead of neural network learning algorithms
for training of membership function parameters, we note their advantages and
disadvantages.</p>
      <p>The advantages of genetic algorithms for neural networks training are the
following: the probability of getting into a local extremum decreases.</p>
      <p>The disadvantages of genetic algorithms for neural networks training are the
following: the speed of the solution search method is lower than that of neural network
training methods; in the case of binary genes, an increase in the search space reduces
the accuracy of the solution with a constant chromosome length; in the case of binary
genes, there are encoding/decoding operations that reduce the speed of the algorithm.</p>
      <p>In this regard, it is relevant to create a method of digital content processing for
biometric identification of a person, which will eliminate these drawbacks.</p>
      <p>The aim of the work is to increase the efficiency of digital content processing
system due to the artificial neuro-fuzzy network, which is trained on the basis of the
genetic algorithm.</p>
      <p>To achieve this goal, it is necessary to solve the following tasks:</p>
    </sec>
    <sec id="sec-2">
      <title>1. Generation of digital content attributes. 2. Creation of a model of digital content processing system. 3. Choice of the structure of the method for determining the parameter values of the mathematical model of digital content processing system.</title>
      <p>2</p>
      <sec id="sec-2-1">
        <title>Generation of digital content attributes</title>
        <p>The generation of digital content attributes in the case of biometric identification of a
person by voice provides for the following steps:
─ determination of vocal segments of a speech signal based on statistical estimation
of short-term energies;
─ definition of formants of the central frame of the vocal segment;
─ choice of vocal speech sound attributes based on formants of the central frame of
the vocal segment.
2.1</p>
        <sec id="sec-2-1-1">
          <title>Determination of vocal segments of a speech signal based on statistical estimation of short-term energies</title>
          <p>The paper proposes a method for determining vocal segments of a speech signal based
on statistical estimation of short-term energies, which includes the following steps:
1. Set a speech signal with one vocal sound y(n) , n 1, N f . Set the number of
quantization levels of a speech signal L (for an 8-bit sound sample L  256 ). Set
the length of the frame N , on which the short-term energy is calculated,
N  2b 1, where the integer parameter b is selected from the inequality
b 1  log2  fs fmin   b , fs is the sampling frequency of the speech signal in Hz,
8000  fs  22050 , fmin is the minimum frequency of the fundamental human
tone in Hz, fmin  50 . Set the parameter for adaptive threshold  , 0    1.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Calculate short-term energies</title>
      <p>N /2

mN /2
E(n) </p>
      <p>y2 (m  n) , n  N / 2 1, N f  N / 2 1 .
3. Calculate the mathematical expectation of short-term energies
 
1 N f N /21</p>
      <p>
N f  N 1 nN /21</p>
      <p>E(n) .
4. Calculate the standard deviation of short-term energies
 
1 N f N /21</p>
      <p>
N f  N 1 nN /21</p>
      <p>E 2 (n)   2 .
5. Calculate the adaptive threshold T    .
6. Determine the left and right borders of the vocal segment:
6.1. Set the sample number n  1 ;
6.2. If E(n)  T  E(n 1)  T , then N l  n 1 , go to step 6.1;
6.3. If E(n)  T  E(n 1)  T , then N r  n , proceed to completion;
6.4. If n  N f  N 1, then go to the next sample, i.e. n  n 1 , go to step 6.2,
else N r  n , proceed to completion.</p>
      <p>As a result, the left and right boundaries of the vocal segment are determined. For
the method of formants determining, the frame with the center in the sample with the
number N c  round  N l  N r  / 2 is selected as the central frame.
2.2</p>
      <sec id="sec-3-1">
        <title>Definition of formants of the central frame of the vocal segment</title>
        <p>The paper proposes a method for determining the formants of the central frame of the
vocal segment based on linear prediction coding, which includes the following steps:
1. Perform through the low-pass filter the balancing of the spectrum having a steep
decline in the high frequency region</p>
        <p>s (m)  s(m 1)  s(m) , m  N c  N / 2, N c  N / 2 ,
where  is the filtration parameter, 0   1.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>2. Calculate the autocorrelation function R(k)</title>
      <p>s (m)  s (m)w(m) , w(m)  0.54  0.46 cos</p>
      <p>Nc N /21k</p>
      <p>
mNc N /2
R(k) 
s (m)s (m  k) , k  0, p ,
2 m</p>
      <p>N
where</p>
      <p>w(m) is the Hamming window, p is the linear prediction order,
ceil( fd / 1000)  p  5  ceil( fd / 1000) , ceil( f ) is the function that rounds f to the
nearest integer.
3. Calculate the LPC coefficients a j in accordance with the Durbin procedure
[21, 22]:
3.1. E(0)  R(0) ;</p>
      <p> i1 
3.2. ki  R(i)   (ji1) R(i  j) E(i1) ;</p>
      <p> j1 
3.3. i(i)  ki ;
3.4.  (ji)   (ji1)  ki i(ij1) ,1  j  i 1 ;
3.5. E(i)  (1 ki2 )E(i1) ;
3.6. i  i 1;
3.7. if i  p , then go to step 2;
3.8. a j   (j p) ,1  j  p .</p>
    </sec>
    <sec id="sec-5">
      <title>4. Calculate the gain coefficient G</title>
      <p>G </p>
      <p>E </p>
      <p>p
R(0)   ak R(k ) .</p>
      <p>k1
5. Calculate the logarithmic energy spectrum using the gain and LPC coefficients
10 lgW (k )  10 lg</p>
      <p>G2
1  p am cos  2 km  2   p am sin  2 km  
 m1  N    m1  N  
2 , k  0, N 1
6. Calculate the frequency and amplitude of the formant in the logarithmic energy
spectrum of the central frame:
6.1. Set frequency number k  0 . Set the number of formants i  0 ;
6.2. If 10lgW (k)  10lgW (k 1) 10lgW (k)  10lgW (k 1) 10lgW (k)  0 ,
then fix the formant frequency, i.e. Fi1  k , and the formant amplitude, i.e.</p>
      <p>Ai1  10 lgW (k) , increase the number of local extremums, i.e. i  i 1;
6.3. If i  3 , then go to the next frequency, i.e. k  k 1 , go to step 6.2.
2.3</p>
      <sec id="sec-5-1">
        <title>Choice of vocal speech sound features based on formants of the central frame of the vocal segment</title>
        <p>The following vocal speech sound features have been chosen:
─ - the frequency of the first formant x1  F1 ;
─ - the frequency of the second formant x2  F2 ;
─ - the frequency of the third formant x3  F3 ;
─ - the amplitude of the first anti-formant x4  A1 ;
─ - the amplitude of the second anti-formant x5  A2 ;
─ - the amplitude of the third anti-formant x6  A3 .</p>
        <p>The total number of features is denoted as Q  6 .</p>
        <sec id="sec-5-1-1">
          <title>Creation of a model of digital content processing system</title>
          <p>The proposed digital content processing system that performs biometric identification
of a person by voice is the artificial neuro-fuzzy network, a graph model of which is
shown in Fig. 1.</p>
          <p>z1</p>
          <p>zM
x1
xN
…
…
…
…
…
…
…
…
…
y</p>
          <p>The input (zero) layer contains N (0)  Q neurons (corresponds to the number of
features). The first hidden layer implements the fuzzification and contains N (1)  MQ
neurons (corresponds to the number of values of linguistic variables). The second
hidden layer implements the aggregation of subconditions and contains N (2)  M
neurons (corresponds to the number of rules M ). The third hidden layer implements
the activation of conclusions and contains N (3)  M 2 neurons. The fourth hidden
layer implements the aggregation of conclusions and contains N (4)  M neurons.
The output layer implements the defuzzification and contains N (5)  1 neuron.</p>
          <p>All weighting coefficients are equal to 1.</p>
          <p>The creation of the mathematical model of digital content processing system
involves the following steps:
─ formation of a fuzzy rule base;
─ fuzzification;
─ aggregation of subconditions;
─ activation of conclusions;
─ aggregation of conclusions;
─ defuzzification.
3.1</p>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>Formation of a fuzzy rule base</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Imagine the j -th fuzzy rule in the form</title>
      <p>R j : IF x1 is 1j AND ... AND xQ is  Nj THEN y is  j ,
where xi is the name of the input linguistic variable, i 1, N ; y is the name of the
output linguistic variable;  ij is the fuzzy variable (the value of the linguistic variable
xi ), j 1, M , i 1,Q ;  j is the fuzzy variable (the value of the linguistic variable
y ), j 1, M .</p>
      <p>The fuzzy set Aij is the range of values of the fuzzy variable  ij , the fuzzy set B j
is the range of values of the fuzzy variable  j .
3.2</p>
      <sec id="sec-6-1">
        <title>Fuzzification</title>
        <p>Let’s determine the degree of truth of the i -th subcondition, i.e. let’s establish the
correspondence between the input variables xi of the j -th rule and the values of the
membership function  Aij (xi ) .</p>
        <p>Since a number of methods related to person identification by voice use the Gauss
function, we choose this function as  Aij (xi ) , i.e.</p>
        <p> 1  xi  mij 2  ,
 Aij (xi )  exp  2   ij  
where mij is the mathematical expectation,  ij is the standard deviation.
3.3</p>
      </sec>
      <sec id="sec-6-2">
        <title>Aggregation of subconditions</title>
        <p>The membership function of the condition for the j -th rule is defined as
 Aj (x)   A1j (x1)... Anj (xn ) , j 1, M .
3.4</p>
      </sec>
      <sec id="sec-6-3">
        <title>Activation of conclusions</title>
        <p>The membership function of the conclusion for the j -th rule is defined as
C j ( y)   Aj (x)B j ( y) , j 1, M ,
0, x  j  0.5
 x  ( j  0.5) 0.5, j  0.5  x  j
 B j ( y)  ( j  0.5)  x 0.5, j  x  j  0.5</p>
        <p>0, x  j  0.5
3.5</p>
      </sec>
      <sec id="sec-6-4">
        <title>Aggregation of subconditions</title>
        <p>The membership function of the final conclusion is defined as
is a triangular function.</p>
        <p>C ( y)  max(C1 ( y),...,CM ( y)) .
3.6</p>
      </sec>
      <sec id="sec-6-5">
        <title>Defuzzification</title>
        <p>To obtain the class number, the membership function maximum method is used.
y  arg max C (z j ) ; z j is the center of the fuzzy set C j .</p>
        <p>z j</p>
        <p>Thus, the mathematical model of digital content processing system (Fig. 1) can be
represented as</p>
        <p>Q
y  arg mzakx mj1a,Mx  B j (zk ) i1  Aij (xi ) , k 1, M .</p>
        <p>The determination of the parameters of this system is carried out on the basis of the
genetic algorithm.
4</p>
        <sec id="sec-6-5-1">
          <title>Choice of the structure of the method for determining parameter values of the mathematical model of digital content processing system</title>
          <p>The choice of the structure of the genetic algorithm, which allows to determine
parameter values of the mathematical model of digital content processing system,
involves the following steps:
─ identification of individuals of the initial population;
─ definition of fitness function;
─ choice of reproduction (selection) operator;
─ choice of crossing-over operator;
─ choice of mutation operator;
─ choice of reduction operator;
─ definition of a stop condition.
4.1</p>
        </sec>
      </sec>
      <sec id="sec-6-6">
        <title>Identification of individuals of the initial population</title>
        <p>Material genes have been selected for the following reasons:
─ - the ability to search in large spaces, which is difficult to do in the case of binary
genes, when an increase in the search space reduces the accuracy of the solution
with a constant chromosome length;
─ - the ability to configure solutions locally;
─ - the lack of encoding / decoding operations that are necessary for binary genes
increases the speed of the algorithm;
─ - proximity to the formulation of the most applied problems (each material gene is
responsible for one variable or parameter, which is impossible in the case of binary
genes).</p>
        <p>An ordered vector of parameters (mathematical expectations and standard
deviations) acts as the chromosome, which represents the i -th individual of the population
H  {hi}</p>
        <p>hi  (lx11  i *m11,lx12  i *m12 ,...,lx1n  i *m1n ,lxn2  i * mn2 ,
lx11  i * 11,lx12  i * 12 ,...,lx1n  i *  n1 , lxn2  i *  n2 ) , i 1,| H | ,
mkj  rxkj  lxkj ,  kj  rxkj  lxkj , j 1, M ,</p>
        <p>| H | | H |
where | H | is the population power, lxkj , rxkj are the left and right boundaries of the
values of the k -th feature, calculated experimentally.
4.2</p>
      </sec>
      <sec id="sec-6-7">
        <title>Definition of fitness function</title>
        <p>In the paper the following fitness function, which corresponds to the probability of
correct identification of a person by voice, is proposed</p>
        <p>F 
1 P 1, a  0</p>
        <p> I ( yp  d p )  max , I (a)  
P p1 mij ,ij 0, a  0
,
where d p is the response received from the object (person), y p is the response
obtained by the model, P is the number of test implementations.
4.3</p>
      </sec>
      <sec id="sec-6-8">
        <title>Choice of reproduction (selection) operator</title>
        <p>The following effective combination is used to select parameter vectors for crossing
and mutation as a reproduction operator</p>
        <p>P(hi ) </p>
        <p>1
| H |
exp(1 / g(t)) 
1  i 1 </p>
        <p> a  (2a  2) | H | 1  (1  exp(1 / g(t))) .</p>
        <p>| H | </p>
        <p>Thus, in the early stages of the genetic algorithm, an uniform selection is used to
ensure that the entire search space is studied (random selection of chromosomes), and
in the final stages, linearly ordered selection is used to make the search directed (the
current best chromosomes are preserved). This combination does not require scaling
and can be used to minimize fitness function.
4.4</p>
      </sec>
      <sec id="sec-6-9">
        <title>Choice of crossing-over (crossover, recombination) operator</title>
        <p>To combine the two options of the vector of parameters selected by the reproduction
operator, an uniform crossing-over is used as the crossing-over operator.</p>
        <p>Parents are selected through the following effective combination – in the early
stages of the genetic algorithm, outbreeding is used to provide an investigation of the
entire search space, and in the final stages, inbreeding is used to make the search
directed. This combination does not require scaling and can be used to minimize fitness
function.</p>
        <p>After the selection of parents, a cross is carried out, and two descendants are
produced.</p>
        <p>For a global search for the optimal vector of parameters, it is necessary to increase
the variety of options.
4.5</p>
      </sec>
      <sec id="sec-6-10">
        <title>Choice of mutation operator</title>
        <p>To ensure the variety of options for the vector of parameters after crossing-over, an
non-uniform mutation is used.</p>
        <p>The mutation step is defined as
 
(Max j  hij )r 1
 
  
 
(hij  Min j )r 1
 
t b</p>
        <p> , r  0.5
T 
t b</p>
        <p>
           , r  0.5
T 
,
where Max j , Min j are the maximum and minimum values of the j -th gene; t is the
iteration number; T is the maximum number of iterations; r is the random number,
r [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ] ; b is the parameter controlling the speed of step decrease, b  0 .
        </p>
        <p>To simulate annealing, the probability of mutation is defined as</p>
        <p>Pm  P0 exp(1/ g(t)) , g(t)   g(t 1) , 0    1, g(0)  T0 , T0  0 ,
where P0 is the initial probability of mutation.</p>
        <p>Thus, in the early stages of the genetic algorithm, a large step mutation occurs with
high probability, which provides an investigation of the entire search space, and in the
final stages, the probability of mutation and its step tend to zero, which makes the
search directed.
4.6</p>
      </sec>
      <sec id="sec-6-11">
        <title>Choice of reduction operator</title>
        <p>The reduction operator allows to create a new population based on the previous
population and parameter vectors obtained by crossing-over and mutation. As a reduction
operator, a scheme (   ) is applied that does not require scaling and can be used to
minimize fitness function.
4.7</p>
      </sec>
      <sec id="sec-6-12">
        <title>Definition of a stop condition</title>
        <p>The following condition is proposed in the work
1 max F(hi )    t  T .</p>
        <p>i</p>
        <p>The values of  and T are calculated experimentally.</p>
        <sec id="sec-6-12-1">
          <title>Numerical research</title>
          <p>1. To solve the problem of increasing the efficiency of digital content
processing system for biometric identification of a person by voice, the corresponding
speaker recognition methods have been investigated. These studies have shown that
today the use of artificial neural networks in combination with the fuzzy inference
system and the genetic algorithm is the most effective method.</p>
          <p>2. The proposed method of digital content processing for biometric
identification of a person by voice automates the process of generation of digital content
features, provides a representation of knowledge in the form of rules that are easily
accessible for human understanding, and simplifies the determination of the structure of
the model due to the fuzzy inference system; reduces the probability of falling into a
local extremum and provides an acceptable speed for determining the parameter
values of the model by choosing the effective structure of the genetic algorithm; allows
parallel processing of information due to the artificial neural network.</p>
          <p>3. As a result of a numerical study, it has been found that the proposed method
of digital content processing provides 0,98 probability of biometric identification of a
person by voice, which exceeds the probability obtained by the artificial neural
network such as a multilayer perceptron.</p>
          <p>4. The proposed method of digital content processing for biometric
identification of a person by voice can be used in various intelligent systems for digital content
processing.
2. Jain, A.K., Flynn, P., Ross, A. (Eds.): Handbook of biometrics. Springer, New York, NY
(2008).
3. Dunstone, T., Yager, N.: Biometric system and data analysis: design, evaluation, and data
mining. Springer, New York (2009).
4. Singh, N., Khan, R., Shree, R.: Applications of Speaker Recognition. Procedia
Engineering. 38, 3122–3126 (2012). doi: 10.1016/j.proeng.2012.06.363
5. Li, Q.: Speaker authentication. Springer-Verlag Berlin Heidelberg, Heidelberg (2012).
6. Keshet, J., Bengio, S.: Automatic speech and speaker recognition: large margin and kernel
methods. John Wiley &amp; Sons, Chichester (2009).
7. Herbig, T., Gerl, F., Minker, W.: Self-learning speaker identification: a system for
enhanced speech recognition. Springer, Berlin (2013).
8. Campbell, J.: Speaker recognition: a tutorial. Proceedings of the IEEE. 85, 1437–1462
(1997). doi: 10.1109/5.628714
9. Togneri, R., Pullella, D.: An Overview of Speaker Identification: Accuracy and
Robustness Issues. IEEE Circuits and Systems Magazine. 11, 23–61 (2011).
doi: 10.1109/MCAS.2011.941079
10. Beigi, H.: Fundamentals of speaker recognition. Springer, New York (2011).
11. Reynolds, D.A.: An overview of automatic speaker recognition technology. IEEE
International Conference on Acoustics Speech and Signal Processing. 4, 4072–4075 (2002).
doi: 10.1109/ICASSP.2002.5745552
12. Kinnunen, T., Li, H.: An overview of text-independent speaker recognition: From features
to supervectors. Speech Communication. 52, 12–40 (2010).
doi: 10.1016/j.specom.2009.08.009
13. Reynolds, D., Rose, R.: Robust text-independent speaker identification using Gaussian
mixture speaker models. IEEE Transactions on Speech and Audio Processing. 3, 72–83
(1995). doi: 10.1109/89.365379
14. Zeng, F.-Z., Zhou, H.: Speaker Recognition based on a Novel Hybrid Algorithm. Procedia</p>
          <p>Engineering. 61, 220–226 (2013). doi: 10.1016/j.proeng.2013.08.007
15. Jeyalakshmi, C., Krishnamurthi., V., Revathi, A.: Speech recognition of deaf and hard of
hearing people using hybrid neural network. 2010 2nd International Conference on
Mechanical and Electronics Engineering. (2010). doi: 10.1109/ICMEE.2010.5558589
16. Nayana, P., Mathew, D., Thomas, A.: Comparison of Text Independent Speaker
Identification Systems using GMM and i-Vector Methods. Procedia Computer Science. 115, 47–54
(2017). doi: 10.1016/j.procs.2017.09.075
17. Chauhan, V., Dwivedi, Sh., Karale, P., Potdar, S.M.: Speech to text converter using
Gaussian mixture model (GMM). International Research Journal of Engineering and Technology
(IRJET). 3, 160–164 (2016).
18. Reynolds, D.A.: Automatic speaker recognition using Gaussian mixture speaker models.</p>
          <p>IEEE Transactions on Speech and Audio Processing. 3, 1738–1752 (1995).
19. Fedorov, E., Lukashenko, V., Utkina, T., Rudakov, K., Lukashenko, A.: Method for
parametric identification of Gaussian mixture model based on clonal selection algorithm.</p>
          <p>CEUR Workshop Proceedings. 2353, 41–55 (2019).
20. Larin, V.J., Fedorov, E.E.: Combination of PNN network and DTW method for
identification of reserved words, used in aviation during radio negotiation. Radioelectronics and
Communications Systems. 57, 362–368 (2014). doi: 10.3103/S0735272714080044
21. Rabiner, L.R., Juang, B.-H.: Fundamentals of speech recognition. Pearson Education,
Delhi (2005).
22. Markel, J.D., Gray, A.H.: Linear prediction of speech. Springer-Verlag, Berlin (1976).</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bolle</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Connell</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pankanti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ratha</surname>
            ,
            <given-names>N.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Senior</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          : Guide to biometrics. Springer, New York (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>