<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Method for Parametric Identification of Gaussian Mixture Model Based on Clonal Selection Algorithm</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cherkasy State Technological University</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cherkasy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shevchenko blvd.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ukraine</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>k.rudakov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>t.utkina}@chdtu.edu.ua</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>fedorovee</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@ukr.net</string-name>
          <email>ineks-kiev@ukr.net</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>E. O. Paton Electric Welding Institute</institution>
          ,
          <addr-line>Kyiv, Bozhenko str., 11, 03680</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1899</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The problem of increasing the efficiency of parametric identification of Gaussian mixture model (GMM) is considered. A method for identifying GMM parameters based on clonal selection algorithm with preliminary formation of a learning set taking into account the structure of vocal sounds, which increases the likelihood of speaker recognition, is proposed. The parameter identification method is intended for software implementation on the GPU using the CUDA technology, which speeds up the process of parametric identification. This method has been studied on the TIMIT database and serves for intelligent systems of biometric personal identification.</p>
      </abstract>
      <kwd-group>
        <kwd>quasi-periodic signal</kwd>
        <kwd>speaker recognition</kwd>
        <kwd>gaussian mixture model</kwd>
        <kwd>learning set formation</kwd>
        <kwd>parametric identification</kwd>
        <kwd>clonal selection algorithm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Automated biometric identification of a person means making decisions based on
acoustic and visual information, which improves the quality of recognition of the
person under study. Unlike the traditional approach [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], computer biometric
identification speeds up and improves the accuracy of the recognition process, which is
especially critical in the conditions of limited time.
      </p>
      <p>
        A special class of biometric identification of a person is formed by methods based
on the analysis of acoustic information [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        The well-known voice identification methods, such as dynamic programming
[34], analyze the entire signal, which increases the likelihood of recognition, but require
entire storage of learning signals and long comparison of the analyzed signal with all
learning signals. Other well-known methods, such as vector quantization [
        <xref ref-type="bibr" rid="ref5 ref6">5-6</xref>
        ], neural
network method [
        <xref ref-type="bibr" rid="ref7 ref8">7-8</xref>
        ] and decision tree [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], store only generalized characteristics of
learning signals and carry out a quick classification of signal components, but analyze
these components without interconnecting them, which reduces the recognition
probability. Network methods, including those that use Gaussian mixture models (GMM)
[
        <xref ref-type="bibr" rid="ref10 ref11">10-13</xref>
        ], which quickly analyze the entire signal and store only generalized
characteristics of learning signals, are a compromise between the above groups of methods. At
the same time, GMM-based methods use the identification of its parameters based on
local search, which reduces the probability of recognition.
      </p>
      <p>The aim of the work is to increase the efficiency of parametric identification
method of Gaussian mixture model (GMM) due to a metaheuristic clonal selection
algorithm, with preliminary formation of a learning set.</p>
      <p>To achieve this goal, it is necessary to solve the following tasks:</p>
    </sec>
    <sec id="sec-2">
      <title>1. to develop a method for learning set formation; 2. to create a method of GMM parametric identification; 3. to develop an algorithm of GMM parametric identification; 4. to conduct a numerical study of the proposed method of parametric identification.</title>
      <p>2</p>
      <sec id="sec-2-1">
        <title>Formal problem statement</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Let a learning set be defined for a particular speaker. Then the problem of increasing the probability of speaker recognition by Gaussian mixture model (GMM) is presented as the problem of finding such a vector of parameters for this model that satisfies the criterion.</title>
      <p>3</p>
      <sec id="sec-3-1">
        <title>Literature review</title>
        <p>
          Biometric identification methods based on dynamic programming [
          <xref ref-type="bibr" rid="ref3 ref4">3-4</xref>
          ] compare the
signal representing the recognizable speaker with all signals representing the
wellknown speakers. The advantage of these methods lies in their ability to analyze the
entire signal and compare signals of different lengths, which increases the probability
of recognition. The disadvantage of these methods consists in the entire storage of
signals representing the well-known speakers and the duration of the procedure for
comparison of the recognizable signal with all available signals.
        </p>
        <p>
          Biometric identification methods based on vector quantization [
          <xref ref-type="bibr" rid="ref5 ref6">5-6</xref>
          ] form the
codebook vectors as averaged components of the signals representing the well-known
speakers and compare components of the signal representing the recognizable speaker
with these vectors. The advantage of these methods lies in the storage of only the
codebook vectors and quick comparison of the recognizable signal components with
these vectors. The disadvantage of these methods consists in the ability to analyze
only single components of the signal without interconnecting them, which reduces the
likelihood of recognition.
        </p>
        <p>
          Biometric identification methods based on artificial neural networks [
          <xref ref-type="bibr" rid="ref7 ref8">7-8</xref>
          ] identify
the neural network parameters, which store information about signal components
representing the well-known speakers and recognize signal components representing
the recognizable speaker through the neural network. The advantage of these methods
lies in the storage of only the neural network parameters and quick classification of
the recognizable signal components. The disadvantage of these methods consists in
the identification of neural network parameters based on local search and the ability to
analyze only single components of the signal without interconnecting them, which
reduces the likelihood of recognition.
        </p>
        <p>
          Biometric identification methods based on decision trees [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] automatically identify
threshold values (for example, formants) at decision tree nodes representing speech
characteristics of the well-known speakers and compare threshold values of the signal
component representing the recognizable speaker with threshold values of the
decision tree. The advantage of these methods lies in the storage of only threshold values
and quick classification of the recognizable signal components. The disadvantage of
these methods consists in the complexity and subjectivity of threshold values
formation and the ability to analyze only single components of the signal without
interconnecting them, which reduces the likelihood of recognition.
        </p>
        <p>Biometric identification methods based on Gaussian mixture models (GMMs)
[1013] identify GMM parameters, which store information about signal components
representing the well-known speakers, and recognize the signal representing the
recognizable speaker by means of these GMMs. The advantage of these methods lies in
the ability to quickly recognize the entire signal and store only GMM parameters. The
disadvantage of these methods consists in the identification of GMM parameters
based on local search, which reduces the likelihood of recognition.</p>
        <p>A common feature of all methods listed above consists in the following: they form
a learning set of signal patterns and analyze a signal without taking into account the
structure of vocal sounds, which leads to a decrease in the probability of recognition.</p>
        <p>Therefore, the increase of the efficiency of the method of GMM parametric
identification with preliminary formation of a learning set, taking into account the structure
of vocal sounds, is an urgent task.
4</p>
      </sec>
      <sec id="sec-3-2">
        <title>Method of learning set formation</title>
        <p>The method of learning set formation includes the following steps:
─ partitioning of a quasi-periodic signal into discrete patterns;
─ shifting of discrete patterns in time and amplitude;
─ interpolation of discrete patterns;
─ shifting and scaling of continuous patterns in time;
─ shifting and scaling of continuous patterns in amplitude;
─ sampling of continuous patterns.</p>
        <sec id="sec-3-2-1">
          <title>1. Partitioning of a quasi-periodic signal into discrete patterns</title>
          <p>Let’s define a finite set of discrete patterns of a quasi-periodic signal, described by
a family of integer bounded finite discrete functions X  {xi | i {1,..., I}} , in the
form:</p>
          <p> f (n), n {Nimin ,..., Nimax }
xi (n)  
0, n {Nimin ,..., Nimax }</p>
          <p>, i {1,..., I} ,
Amin  min f (n) , n {Nimin ,..., Nimax } , i {1,..., I} ,</p>
          <p>i n
Amax  max f (n) , n {Nimin ,..., Nimax } , i {1,..., I} ,</p>
          <p>i n
where Amin , Aimax – the minimum and maximum values of the function xi on the
i
compact {Nimin ,..., Nimax } .</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>2. Shifting of discrete patterns in time and amplitude</title>
          <p>Let’s define a finite set of discrete patterns shifted in time and amplitude, which are
described by a finite family of integer bounded finite discrete functions
X s  {xis | i {1,..., I}} , in the form:
xis (n)  xi (n  Nimin )  Aimin , n {0,..., Ni } , i {1,..., I} ,
0, n {0,..., Ni }</p>
          <p>Ni  Nimax  Nimin , Ai  Aimax  Aimin .</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3. Interpolation of discrete patterns</title>
          <p>A linear interpolation, which requires the least computational complexity, has been
chosen in the article. Let’s define a finite set of continuous patterns, obtained as a
result of linear interpolation and described by a finite family of real-valued bounded
finite continuous functions   { i | i {1,..., I}} , in the form:</p>
          <p>t [t,Ti  t]  i (t) 
   (tn ,tn1) (t)  xis (n)  xis (n  1)  xis (n) (t  tn )    {tn}(t)xis (n) ,
Ni  Ni 1
n1  t  n1
i {1,..., I} , t [t,Ti  t]  i (t)  0 , Ti  Ni t , tn  nt ,</p>
          <p>1, t  B
 B (t)  
0, t  B</p>
          <p>– indicator function,
where t – the quantization step in time.</p>
        </sec>
        <sec id="sec-3-2-4">
          <title>4. Shifting and scaling of continuous patterns in time</title>
          <p>Let’s define a finite set of shifted and scaled in time continuous patterns, described by
a finite family of real-valued bounded finite continuous functions
s  { is | i {1,..., I}} , in the form:
 t  Tmin 
t [Tmin ,Tmax ]  is (t)  i Ti Tmax  Tmin  ,</p>
          <p>
t [Tmin , Tmax ]  is (t)  0 ,
where [Tmin ,Tmax ] – the compact, set and single for all patterns.</p>
          <p>The article proposes to define Tmin ,Tmax as follows:</p>
          <p>Tmin  t ,
Tmax  round  fd  t ,</p>
          <p> fmin 
where fd – the sampling frequency in Hz (for vocal sounds 8 kHz is enough), fmin
– the minimum frequency of frequency range of speech sound (for vocal sounds 50
Hz is enough), round () – the function, rounding the number to the nearest integer.</p>
        </sec>
        <sec id="sec-3-2-5">
          <title>5. Shifting and scaling of continuous patterns in amplitude</title>
          <p>Let’s define a finite set of continuous patterns shifted and scaled in amplitude,
which are described by a finite family of real-valued bounded finite continuous
functions ss  { iss | i {1,..., I}} , in the form:
t [Tmin ,Tmax ]  iss (t)  Amin </p>
          <p>Amax  Amin
A max  Aimin  is t  ,</p>
          <p>i
t [Tmin ,Tmax ]  iss (t)  0 ,</p>
          <p>A max  max is t  ,
i t
A min  min is t  ,
i t
t [Tmin ,Tmax ] ,
where A min , Aimax – the minimum and maximum values of the function  is on the
i
compact [Tmin ,Tmax ] ; Amin , Amax – the minimum and maximum values, set and
single for all patterns, on a given compact [Tmin ,Tmax ] .</p>
          <p>The article proposes to define Amin , Amax as follows: Amin  0 , Amax  2b 1 ,
where b – the number of level quantization bits.</p>
        </sec>
        <sec id="sec-3-2-6">
          <title>6. Sampling of continuous patterns</title>
          <p>Let’s define a finite set of discrete patterns, obtained from continuous ones by
sampling in time and described by a finite family of integral bounded finite discrete
functions S  {si | i {1,..., I}} , in the form:</p>
          <p>si (n)  round ( iss (nt)) , n {N min ,..., N max } ,</p>
          <p>N min  Tmin / t , N max  Tmax / t , N  N max  N min .</p>
          <p>Each received discrete sample is considered as a feature vector and is located in a
single amplitude-time window.
5</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Method for parametric identification of Gaussian mixture model</title>
        <p>Each GMM associated with a particular speaker is defined on the basis of the total
probability formula as follows:</p>
        <p>K
P(s)   P(k ) p(s | k ) ,</p>
        <p>k 1
p(s | k ) </p>
        <p>1
p(s | k ) – densities of multi-dimensional Gaussian distribution, mk  (mk1,..., mkN )
– vector of mathematical expectations, Ck – covariance matrix, K – the number of
mixture components.</p>
        <p>In the work, for GMM parametric identification the target function, which means
the choice of such values of parameter vector  that deliver likelihood functions to
the maximum logarithm, is chosen:</p>
        <p>I K
F  ln P(S | )   ln  P(k) p(si | k)  max ,</p>
        <p>i1 k 1 
where   ((P(1),..., P(K )), (m1,..., mK ), (C1,..., CK )) – GMM parameters vector.</p>
        <p>Since the traditional EM algorithm used for GMM parametric identification
implements only local search, so currently for GMM parametric identification are
actively applied evolutionary, swarm and immune metaheuristics, which use the
parameter vector  as an individual of the population [14-16 ]. However, due to the
large dimensionality of the covariance matrix C , only a diagonal matrix C is used in
these metaheuristics, which reduces the likelihood of identification. Therefore, in this
work, the joint probability vector (P(s1,1),..., P(si , k),..., P(sI , K )) , where
P(si , k )  P(k ) p(si | k ) , is used as an individual of the population, which allows to
work with the off-diagonal matrix C .</p>
        <p>GMM parametric identification is based on clonal selection algorithm proposed in
[17-23] and is presented in the following form:
1. The number of mixture components K , the mutation parameter  , the maximum
number of iterations  max are set.
2. The intermediate population U  u of the power      is created, each
antibody of the population being represented as u  (u11,..., uIK ) .</p>
        <p>
          Each element uik is defined as uik  rand () , where rand () – a function that
returns a uniformly distributed random number in the range [
          <xref ref-type="bibr" rid="ref1">0,1</xref>
          ] .
        </p>
        <p>A finite set of discrete patterns S is given.
3. The iteration number   1 is set, the maximum value of the target function
4. A posteriori conditional probabilities for each antibody u are calculated:
7. Covariance matrices for each antibody u are calculated:</p>
        <p>I
 P(k | si )si
mk  i1I
 P(k | si )
i1</p>
        <p>, k {1,., K} .</p>
        <p>Ck  i1
I P(k | si )  si  m T  si  m </p>
        <p>k k
I
 P(k | si )
i1
, k {1,., K} .
8. The densities of multidimensional Gaussian distribution for each antibody u are
calculated:
p(si | k ) </p>
        <p>1
(2 )N det Ck</p>
        <p> 1 
exp   (si  m )T Ck1(si  mk )  , k {1,., K} , i {1,..., I} .</p>
        <p> 2 k</p>
        <p>K
9. The total probabilities for each antibody u are calculated P(si )   P(k) p(si | k) .
k 1
10. The value of the target function for each antibody u is calculated:</p>
        <p>I
F (u)   ln P(si ) .</p>
        <p>i1
11. The intermediate population U  u is ordered in descending of the target
function value.
12. The current population H  h of the power  is created by means of
reproduction and replacement operators, with each antibody of this population being
represented as h  (h11,..., hIK ) .</p>
        <p>As antibodies of the current population H  h , the first  (the best in the
target function) antibodies of the intermediate population U  u are taken. This
corresponds to the reduction operator with a selection scheme that provides the
search directionality (the best antibodies are preserved). Replacement (randomly
generated) antibodies can be among the best antibodies. This corresponds to the
replacement operator.
13. The value of the target function of the first antibody of the current population</p>
        <p>H  h is set by the minimum value of the target function a .
14. The value of the target function of the last antibody of the current population</p>
        <p>H  h is set by the maximum value of the target function b .
15. If | b  bold |  and    max , then    1 , bold  b , going to 16, otherwise
parametric identification is complete.
16. The affinity for each antibody h is calculated.</p>
        <p>The affinity determines the proximity of the current antibody to the best one and
is calculated based on the utility function in the form:
17. The mutation probability for each antibody h is calculated p(h)  exp  (h) .</p>
        <p>The larger  , the less likely is the mutation probability.
18. A set of clones C  c of the power  is created by means of a cloning operator,
each clone of this set being represented as c  (c11,..., cIK ) .</p>
        <p>The cloning operator plays a role similar to the reproduction operator of genetic
algorithm and is applied to the current population H  h . The number of clones
for each antibody h is defined as   .</p>
        <p>
19. A set of mutated clones C  c of the power  is formed by means of the
proposed mutation operator, each mutated clone of this set being represented in the
form c  (c11,..., cIK ) .</p>
        <p></p>
        <p>The mutation operator allows to obtain new antibodies with sharply different
properties from antibody clones.</p>
        <p>Each element cik is defined as:</p>
        <p>cik , rik  p(h)
rik  rand () , cik  
round (), rik  p(h)</p>
        <p>The features of the proposed variant of the mutation operator are the following:
it does not require the use of binary potential solutions, i.e. it doesn’t need to
convert probabilistic potential solutions into binary ones before the mutation and to
convert binary potential solutions into probabilistic ones after the mutation, which
reduces the computational complexity of the mutation operator and allows to
quickly find a solution.
20. A set of replacement antibodies D  d of the power  is created, each antibody
of the set is represented as d  (d11,..., dIK ) .</p>
        <p>Each element dik is defined as dik  rand () .
21. The intermediate population U  u of the power     is formed.</p>
        <p>The antibodies of the current population H  h and a set of mutated clones

C  c are taken as the first    antibodies of the intermediate population
U  u . The antibodies of a set of the replacement antibodies D  d are taken
as the last  antibodies of the intermediate population U  u . Going to 4.</p>
        <p>As a result of the method, GMM parameters will be defined.
6</p>
      </sec>
      <sec id="sec-3-4">
        <title>Creation of an algorithm for parametric identification of</title>
      </sec>
      <sec id="sec-3-5">
        <title>Gaussian mixture model based on clonal selection algorithm</title>
        <p>The algorithm for GMM parametric identification based on clonal selection
algorithm, which is intended for implementation on the GPU using the CUDA technology,
is presented in Fig. 1. This block diagram functions as follows.</p>
        <p>Step 1 – Set the number of mixture components K , the mutation parameter  ,
the maximum number of iterations  max .</p>
        <p>Step 2 – Set the intermediate population U and a finite set of discrete patterns S .</p>
        <p>Step 3 – Set the iteration number   1 , the maximum value of the target function
bold  0 .</p>
        <p>Step 4 – Calculate the total probabilities P(si ) for each antibody u using
I  K  (     ) threads that are grouped into I  (     ) one-dimensional
blocks. In each block, based on the reduction, a sum from K elements of the form
uik is calculated.</p>
        <p>Step 5 – Calculate a posteriori conditional probabilities for each antibody u using
I  K  (     ) threads that are grouped into I  K       N s
onedimensional blocks, where N s is the number of threads in the block. Each thread
computes P(k | si )  uik P  si  .</p>
        <p>Step 6 – Calculate the sum of a posteriori conditional probabilities qk for each
antibody u using I  K        threads that are grouped into K      
twodimensional blocks. In each block, based on the reduction, a sum from I elements of
the form P(k | si ) is calculated.</p>
        <p>Step 7 – Calculate a priori unconditional probabilities (weights of the mixture
components) P(k ) for each antibody u using K       threads that are
grouped into K        N s one-dimensional blocks. Each thread computes
P(k )  qk I .</p>
        <p>Step 8 – Calculate expectation vectors mk for each antibody u using
I  K  (     ) threads that are grouped into K  (     ) two-dimensional
blocks. In each block, based on the reduction, a sum from I elements of the form
P(k | si )si qk is calculated.</p>
        <p>Step 9 – Calculate the covariance matrices Ck for each antibody u using
I  K  (     ) threads that are grouped into K  (     ) two-dimensional
blocks. In each block, based on the reduction, a sum from I elements of the form
P(k | si )  si  mk T  si  mk  qk is calculated.</p>
        <p>Step 10 – Calculate the densities of Gaussian multidimensional distribution
p(si | k) for each antibody u using I  K  (    ) threads that are grouped into
I  K       N s one-dimensional blocks. Each thread computes:
p(si | k) </p>
        <p>1</p>
        <p>Step 11 – Calculate the total probabilities P(si ) for each antibody u using
I  K  (    ) threads that are grouped into I  (    ) one-dimensional
blocks. In each block, based on the reduction, a sum from K elements of the form
P(k ) p(si | k) is calculated.</p>
        <p>Step 12 – Calculate the target function values F (u) for each antibody u using
I  (    ) threads that are grouped into     two-dimensional blocks. In
each block, based on the reduction, a sum from I elements of the form ln P(si ) is
calculated.</p>
        <p>Step 13 – Order the target functions by the value based on parallel sorting by
merging into the intermediate population U using     threads that are grouped
into one one-dimensional block.</p>
        <p>1
2
3
23
4
5
6
7
8
9
10
11
12

13
14
15
16
17


18
19
20
21
22</p>
        <p>Step 14 – Create the current population H  h of the power  by means of
element-by-element  copying of the first antibodies of the intermediate population
U  u using I  K   threads that are grouped into I  K  N s one-dimensional
blocks.</p>
        <p>Step 15 – Set the value of the target function of the first antibody of the current
population H  h by the minimum value of the target function a .</p>
        <p>Step 16 – Set the value of the target function of the last antibody of the current
population H  h by the maximum value of the target function b .</p>
        <p>Step 17 – If | b  bold |  and   max , then    1 , bold  b , going to 18,</p>
        <p>Step 19 – Calculate the mutation probability for each antibody h using  threads
that are grouped into one one-dimensional block p(h)  exp  (h) .</p>
        <p>Step 20 – Create a set of clones C  c of the power  by means of
element-byelement copying of the antibodies of the current population H  h using I  K 
threads that are grouped into I  K  N s one-dimensional blocks.</p>
        <p></p>
        <p>Step 21 – Form a set of mutated clones C  {c} of the power  by means of the
proposed mutation operator, using I  K  threads that are grouped into I  K  N s
one-dimensional blocks. Each thread computes:</p>
        <p>cik , rik  p(h)
rik  rand () , cik  
round () rik  p(h)
.</p>
        <p>Step 22 – Create a set of random antibodies D  d of the power  using
I  K  threads that are grouped into I  K  N s one-dimensional blocks. Each
thread computes dik  rand () .</p>
        <p>Step 23 – Generate the intermediate population U  u of the power    
by means of element-by-element copying of the antibodies of the current population

H  h , a set of mutated clones C  c , a set of replacement antibodies D  d ,
using I  K  (     ) threads that are grouped into I  K       N s
onedimensional blocks. Going to 4.
7</p>
      </sec>
      <sec id="sec-3-6">
        <title>Experiments and results</title>
        <p>Numerical experiments were carried out using the CUDA technology of parallel
processing of information on the GeForce 920M video card with the number of threads in
the block N s = 1024. Parametric identification was carried out for each of the 100
GMMs, corresponding to 100 speakers. In the work it has been accepted that the
power of a set of patterns of vocal speech sounds I  1024 , the power of the current
population   20 , the power of a set of clones   1000 , the power of a set of
random antibodies   0.2,   4 , the mutation parameter   2.3 , the maximum
number of iterations  max  1000 , the parameter   106 .</p>
        <p>The dependence of the probability of incorrect recognition of the speaker on the
number of components of the mixture is shown in Fig.2. This dependence shows that
the probability of incorrect recognition of the speaker decreases with increasing
number of components of the mixture.</p>
        <p>Table 1 presents the speaker’s recognition probabilities obtained from the TIMIT
database using a neural network based on radial basis functions (RBFNN), trained on
the basis of error correction with the MFCC feature system, Gaussian mixture model
(GMM), trained on the basis of EM-algorithm with the MFCC feature system,
Gaussian mixture model, trained on the basis of clonal selection algorithm, with features
obtained by means of the proposed method of learning set formation.</p>
        <p>According to Table 1, the best results are given by GMM, trained on the basis of
clonal selection algorithm with features obtained by means of the proposed method of
learning set formation.</p>
        <p>This is due to the fact that error correction and the EM algorithm only perform a
local search, which increases the likelihood of falling into a local extremum, and
learning set formation based on the MFCC is done without taking into account the
structure of vocal sound.
8</p>
      </sec>
      <sec id="sec-3-7">
        <title>Conclusion</title>
        <p>The article deals with the problem of increasing the efficiency of parametric
identification of Gaussian mixture model (GMM). The method of learning set formation,
which uses shifting, scaling, interpolation and sampling of signal patterns that
correspond to quasi-periods to locate them in a single amplitude-time window, which
allows to take into account the quasi-periodic signal structure and increase the
probability of speaker recognition, has been improved. The method of GMM parametric
identification, which is based on clonal selection algorithm that allows to reduce the
likelihood of falling into a local extremum, has reached further development. The
proposed method allows probabilistic individuals of the population (potential solutions)
in the mutation operator, which accelerates parametric identification (it is not
necessary to perform the operations of converting real individuals into binary ones and vice
versa), and uses not the parameter vector, but the vector of joint probabilities as an
individual of the population, which allows to work with a non-diagonal covariance
matrix and increases the likelihood of speaker recognition. The algorithm of GMM
parametric identification, which is intended for software implementation on the GPU
using the CUDA technology that speeds up the GMM learning process, has been
created. The software that implements the proposed algorithm has been developed and
investigated on the TIMIT database. The experiments have confirmed the operability
of the developed software and allow to recommend it for use in practice. Prospects for
further research are to test the proposed methods on a wider set of test databases.
Procedia Engineering, Cochin, India (2017)
doi: doi.org/10.1016/j.procs.2017.09.075
12. Chauhan, V., Dwivedi, Sh., Karale, P., Potdar, S. M.: Speech to text converter
using Gaussian Mixture Model (GMM), International Research Journal of
Engineering and Technology (IRJET) 3(5), 160–164 (2016)
13. Reynolds, D. A.: Automatic Speaker Recognition Using Gaussian Mixture
Speaker Models. The Lincoln Laboratory Journal 8(2), 173–195 (1995)
14. Lin, L., Wang, Sh.: A New Genetically Optimized GMM for Speaker .</p>
        <p>In: Published in 6th World Congress on Intelligent Control and Automation,
pp. 10235–10239 (2006) doi: 10.1109/WCICA.2006.1714005
15. Zablotskiy, S., Pitakrat, T., Zablotskaya, K., Minker, W.:GMM Parameter
Estimation by Means of EM and Genetic Algorithms. In: Human-Computer Interaction.
Design and Development Approaches: 14th international conference on
Humancomputer interaction: design and development approaches, pp. 527–536.
Proceeding, Orlando, FL (2011) doi: 10.1007/978-3-642-21602-2_57
16. Saeidi, R., Mohammadi, H. R. S., Ganchev, T., Rodman, R. D.: Particle Swarm
Optimization for Sorted Adapted Gaussian Mixture Models. IEEE Transactions
on Audio, Speech, and Language Processing 17(2), 344–353 (2009)
doi: 10.1109/TASL.2008.2010278
17. de Castro, L. N., von Zuben, F. J.: The Clonal Selection Algorithm with
Engineering Applications. In: Proceedings of the Genetic and Evolutionary Computation
Conference (GECCO '00), Workshop on Artificial Immune Systems and Their
Applications, pp. 36-39. Las Vegas, Nv (2000)
18. de Castro, L. N., von Zuben, F. J.: Learning and Optimization Using Clonal
Selection Principle. IEEE Transactions on Evolutionary Computation 6(3), 239–251
(2002) doi: 10.1109/TEVC.2002.1011539
19. Babayigit, B. A., Guney, K., Akdagli, A.: Clonal Selection Algorithm for Array
Pattern Nulling by Controlling the Positions of Selected Elements. Progress in
Electromagnetics Research B 6, 257–266 (2008) doi: 10.2528/PIERB08031218
20. White, J. A., Garrett, S. M.: Improved Pattern Recognition with Artificial Clonal
Selection? In: Timmis J., Bentley P. J., Hart E. (eds). 2nd International
Conference on Artificial Immune Sysyems, vol. 2787, pp. 181–193. Springer, Berlin
(2003) doi: 10.1007/978-3-540-45192-1_18
21. Alba, E., Nakib, A., Siarry, P.: Metaheuristics for Dynamic Optimization,
Springer-Verlag, Berlin (2013)
22. Du, K.-L., Swamy, M. N. S.: Search and Optimization by Metaheuristics.
Techniques and Algorithms Inspired by Nature, Springer, Birkhäuser Basel (2016)
doi: 10.1007/978-3-319-41192-7
23. Brownlee, J.: Clever Algorithms: Nature-Inspired Programming Recipes,
Melbourne, Brownlee (2012)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shree</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Applications of Speaker Recognition</article-title>
          . In: Rajesh,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <surname>K</surname>
          </string-name>
          . (eds.).
          <source>International Conference on Modelling Optimization and Computing</source>
          , vol.
          <volume>38</volume>
          , pp.
          <fpage>3122</fpage>
          -
          <lpage>3126</lpage>
          . Elsevier Procedia Engineering (
          <year>2012</year>
          ) doi: 10.1016/j.proeng.
          <year>2012</year>
          .
          <volume>06</volume>
          .363
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          :
          <article-title>Speaker Recognition: a tutorial</article-title>
          .
          <source>Proceedings of the IEEE</source>
          <volume>85</volume>
          (
          <issue>9</issue>
          ),
          <fpage>1437</fpage>
          -
          <lpage>1462</lpage>
          (
          <year>1997</year>
          ) doi: 10.1109/5.628714
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Togneri</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pullela</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>An Overview of Speaker Identification: Accuracy and Robustness Issues</article-title>
          .
          <source>IEEE Circuits And Systems Magazine</source>
          <volume>11</volume>
          (
          <issue>2</issue>
          ),
          <fpage>23</fpage>
          -
          <lpage>61</lpage>
          (
          <year>2011</year>
          ) doi: 10.1109/
          <string-name>
            <surname>MCAS</surname>
          </string-name>
          .
          <year>2011</year>
          .941079
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Beigi</surname>
          </string-name>
          , H.:
          <source>Fundamentals of Speaker Recognition</source>
          , Springer, New York (
          <year>2011</year>
          ) doi: 10.1007/978-0-
          <fpage>387</fpage>
          -77592-0
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Reynolds</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          :
          <article-title>An Overview of Automatic Speaker Recognition Technology</article-title>
          .
          <source>In: 2002 IEEE International Conference on Acoustics, Speech, and Signal Processing</source>
          , vol.
          <volume>4</volume>
          , pp.
          <fpage>4072</fpage>
          -
          <lpage>4075</lpage>
          . IEEE, Orlando, FL, USA (
          <year>2002</year>
          ) doi: 10.1109/ICASSP.
          <year>2002</year>
          .5745552
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kinnunen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>An Overview of Text-Independent Speaker Recognition: From Features to Supervectors</article-title>
          .
          <source>Speech Communication</source>
          <volume>52</volume>
          (
          <issue>1</issue>
          ),
          <fpage>12</fpage>
          -
          <lpage>40</lpage>
          (
          <year>2010</year>
          ) doi: 10.1016/j.specom.
          <year>2009</year>
          .
          <volume>08</volume>
          .009
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Reynolds</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rose</surname>
            ,
            <given-names>R. C.</given-names>
          </string-name>
          :
          <article-title>Robust Text-Independent Speaker Identification Using Gaussian Mixture Speaker Models</article-title>
          .
          <source>IEEE Transactions on Speech and Audio Processing</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <fpage>72</fpage>
          -
          <lpage>83</lpage>
          (
          <year>1995</year>
          ) doi: 10.1109/89.365379
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Zeng</surname>
            , F.-
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
          </string-name>
          , H.:
          <article-title>Speaker Recognition based on a Novel Hybrid Algorithm</article-title>
          .
          <source>In: 25th International Conference on Parallel Computational Fluid Dynamics</source>
          , vol.
          <volume>61</volume>
          , pp.
          <fpage>220</fpage>
          -
          <lpage>226</lpage>
          . Elsevier Procedia Engineering (
          <year>2013</year>
          ) doi: 10.1016/j.proeng.
          <year>2013</year>
          .
          <volume>08</volume>
          .007
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jeyalakshmi</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krishnamurthi</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Revathi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Speech Recognition of Deaf and Hard of Hearing People Using Hybrid Neural Network</article-title>
          .
          <source>In: 2010 2nd International Conference on Mechanical and Electronics Engineering</source>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>87</lpage>
          . IEEE, Kyoto, Japan (
          <year>2010</year>
          ) doi: 10.1109/ICMEE.
          <year>2010</year>
          .5558589
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Rabiner</surname>
            ,
            <given-names>L. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jang</surname>
            ,
            <given-names>B. H.</given-names>
          </string-name>
          :
          <article-title>Fundamentals of speech recognition, Prentice Hall</article-title>
          , NJ, USA (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Nayana</surname>
            ,
            <given-names>P. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mathew</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Comparison of Text Independent Speaker Identification Systems Using GMM and i-Vector Methods</article-title>
          .
          <source>In: 7th International Conference on Advances in Computing &amp; Communications</source>
          , pp.
          <fpage>47</fpage>
          -
          <lpage>54</lpage>
          . Elsevier
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>