<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building of a Speaker's Identification System Based on Deep Learning Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Victor Solovyov</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleg Rybalskiy</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vadim Zhuravel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aleksandr Shablya</string-name>
          <email>alik.shablya@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yevhen Tymko</string-name>
          <email>e.tymko@kndise.gov.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kyiv Scientific Research Institute of Forensic Expertise of the Ministry of Justice of Ukraine</institution>
          ,
          <addr-line>Smolenska Str. 6, Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kyiv scientifically-research expertly-criminalistics center Ministries of Internal Affairs of Ukraine</institution>
          ,
          <addr-line>Vladimirska Str. 15, Kyiv, 01001</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>National Academy of Internal Affairs</institution>
          ,
          <addr-line>Solom'yanska Are. 1, Kyiv, 03035</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Odessa Research Institute of Forensic Expertise of the Ministry of Justice of Ukraine</institution>
          ,
          <addr-line>Uspenska Str. 83/85, Odesa, 65011</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Silentium Systems”</institution>
          ,
          <addr-line>Hornby Str 950 - 777, BC V6Z 1S4, Vancouver</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Paper discusses the main methodological and technological features of studying and building of speaker's identification systems built on the basis of deep learning neural networks. The problems and tasks arising in the process of creation of such systems are considered. The purpose of the research article is to show the suggested ways and methods for their elimination and solutions, which were found in the process of creation of an automated system for speaker identification and verification, built due to the use of such networks. In the process of its creation, the method of spectral analysis of speech signals at short time intervals was determined, which ensures high resolution. In addition, a solution of the problem of invariance of the system to different language groups and the duration of phonograms is proposed. This ensured the generality and efficiency of the obtained results of speaker's identification. As a result, an automated system for forensic identification and speaker's verification was developed on the basis of deep learning neural networks. In the process of the development of a system based on the comparison of the spectral characteristics of speech signals, several methods have been proposed and tested providing the possibility of identification (verification) of the speaker by speech messages of short duration.</p>
      </abstract>
      <kwd-group>
        <kwd>1 phonogram duration</kwd>
        <kwd>forensic identification of speaker</kwd>
        <kwd>deep learning neural network</kwd>
        <kwd>spectral analysis</kwd>
        <kwd>frequency domain</kwd>
        <kwd>efficiency</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The use of modern technologies of neural
networks for the examination of materials and
digital sound recording equipment allows, as a
rule, to obtain a more higher level of its efficiency
[
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ]. Usually, the efficiency of the system is
taken as a quantity determined by the probability
of errors of the first and second kinds, inherent in
mentioned type of examination.
      </p>
      <p>
        According to the materials of the SRE NIST
tests carried out recently, the point of cross of the
graphs of errors of both the first and second kinds
for the best automated speaker identification
systems is on average (3–10)%. At the same time,
several tests of systems on messages of less than
10 s duration are carried out relatively rarely [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Similar data are provided by the materials of
the working group of forensic speech and audio
analyzes of the European Network of Forensic
and Scientific Institutions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>It is generally accepted that the effectiveness
of automated speaker’s identification systems is
significantly lower than the effectiveness of, for
example, fingerprinting, video or DNA
recognition.</p>
      <p>Paper presents the results of research and
development of a speaker identification system
based on spectral characteristics of voice signals
and on deep learning neural networks, which, we
believe, can change this point of view.</p>
      <p>During past two decades, the number of
noteworthy publications related to technical
systems for voice identification numbers in the
thousands. Therefore, we will consider only the
main methodological and technological features
of research and design of such systems that are
directly related to the problems and tasks solved
by the conducted developments.</p>
      <p>Despite the hundreds of different methods and
algorithms for speaker’s identification based on
the physical characteristics of voice signals
various spectral parameters are generally
recognized as physical bases. This applies to both
classical methods and methods based on neural
networks. But the use of various characteristics
and parameters of signals with the use of classical
spectral methods allows only to improve slightly
the efficiency of identification, since any of them
is based on a discrete orthogonal Fourier
transform.</p>
      <p>
        Modern research in the field of
neurophysiology of hearing indicates the
feasibility of spectral analysis of speech signals at
short time intervals [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In particular, a number of
important applications of processing audio
information at short time intervals have already
become classics, for example, when compressing
audio files. Thus, the basic basis for the majority
of audio file compression formats is the
transformation of signals from the time domain to
the frequency domain at short time intervals (16–
20) ms [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. But the discrete orthogonal Fourier
transform at small time intervals (about 20 ms)
has insufficient frequency resolution from the
point of view of the neurophysiology of hearing.
So, for a time interval of 20 ms, the use of such a
transformation for the transition from the time to
the frequency domain, provides a frequency
resolution of 50 Hz. However, it is known from
the neurophysiology of hearing that the resolution
of human hearing is approximately 1 Hz [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        As it will be described below, such a low
frequency resolution at short time intervals
significantly reduces the efficiency of any speaker
identification systems and is one of the
determining factors in the spectral analysis of
audio information. And the modern practice of
expert examination points to serious problems of
speaker identification for phonograms of short
duration [
        <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
        ].
      </p>
      <p>Another important problem, in our opinion, is
that any methods and algorithms for conducting
an examination within the framework of this
methodology relate to certain parameters of sound
signals in the frequency domain, from which
some, as the most important, are selected by an
expert. For example, the frequency of the main
tone, the spectrum of specific sounds, etc. are
compared. In this case, all comparisons are made
for integrative assessments obtained as a result of
averaging the spectral parameters over the entire
duration of the phonogram. This approach, taking
into account numerous accompanying factors and
their variability, often does not provide a high
generality and efficiency of the obtained results.
An important property of almost all known
approaches, including those based on the use of
deep learning neural networks, is also the
difficulty in achieving common results for large,
gradually growing databases. Such databases
form the bases of Big Data arrays used to train
neural networks. But in most cases, the plots of
errors of the first and second kinds, used to check
the quality of network training, will be bound to
the DataSet obtained from a specific training
material. As a result, a large degree of data
generalization requires, as a rule, repetition of
research and calculations for a new training set.</p>
      <p>This raises the problem of the correct
quantitative assessment of the effectiveness of the
forensic identification of the investigated object.
The construction of graphs of the probability of
errors of the first and second kinds, in our opinion,
is the most informative and therefore the most
preferable option for determination of the
magnitude of such errors.</p>
      <p>The use of deep learning neural networks
allows us to consider the possibility of developing
an effective automated universal system designed
for forensic identification of a speaker. Under the
universality of the system, we consider its
suitability for work with the speech of speakers
speaking different languages, belonging to
different sexes and with phonograms of different
duration (including several seconds).</p>
      <p>Thus, in order to solve the problems that exist
in the construction of an automated system for
forensic identification of a speaker, we define the
following tasks:</p>
      <p>to determine the method of spectral analysis of
speech signals at short time intervals, providing
high resolution;</p>
      <p>to propose a solution to the problem of
invariance of the system to different language
groups and the duration of phonograms. This will
ensure the generality and effectiveness of the
results obtained for the identification of the
speaker.</p>
      <p>
        The purpose of this paper is to show the ways
and methods of solving these problems by use of
the example of the results obtained during
research and development of an automated system
for identification and verification of a speaker (the
“Avatar” system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Ways and methods of solving the tasks</title>
      <p>Let us consider a discrete non-orthogonal
time-frequency conversion for a 20 ms time
window with a signal of the audiofrequency
range. To be specific, we will use the Morlet
wavelet with the basis
(1)
Cmor t  Fb 12e j2 Fct e
t2</p>
      <p>
        Fb
,
where
j 1 – imaginary unit,
t – time,
Fb – wavelet width parameter,
Fc – wavelet center frequency [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        In this case, we will consider redundant
transformations, in which the number of samples
in the time domain falling on the selected area is
less than the number of samples in the frequency
domain. So, for example, let's take an arbitrary 20
ms fragment of the speech signal of sound [A]
with a sampling rate of 44100 Hz. Then the
number of discrete samples falling on a 20 ms
segment is N = 882. Let us construct and compare
two types of time-frequency conversion in the
frequency range from 0 to 2500 Hz for the same
signal segment. The first of them is
nonorthogonal based on the Morlet wavelet with a
frequency step DFc = 1 Hz [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The total maximum
possible number of frequency steps in the selected
range is 2500. The second is orthogonal with a
discreteness of DFc = 50 Hz (in accordance with
the duration of the time window). The total
maximum possible number of frequency steps in
the selected range is 50.
      </p>
      <p>Fig. 1 shows an illustration of a comparison of
the spectra of one signal fragment obtained for
two types of transformations.</p>
      <p>
        Visually, these graphs are very close.
However, the difference in the positions of the
local maxima of the spectra for applied
examination problems is very significant, since in
most methods for identifying speakers an
important factor is the value of the frequencies of
such maxima [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As it is seen in Fig. 1, the values
of the frequencies of the local maxima for
orthogonal and non-orthogonal transformations of
the same signal differ by more than 20 Hz. This
circumstance significantly affects the accuracy of
the assessment of the spectral parameters of
speech.
      </p>
      <p>
        It is known that when averaging any function
over a large number of time windows with a
duration of T = 20 ms, the calculation accuracy is
proportional to the square root of the number of
window transformations [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>But this means that to achieve an accuracy of
1 Hz, obtained with a non-orthogonal
transformation on an interval of 20 ms, with
orthogonal transformations taking into account
averaging, 400 × 20 ms = 8 sec are required.</p>
      <p>
        This shows the practical impossibility of
analyzing phonograms of short duration (several
seconds) by use of conventional methods of
timefrequency transformations, which is confirmed by
the modern practice of expertise [
        <xref ref-type="bibr" rid="ref3 ref4">3,4</xref>
        ].
      </p>
      <p>At the same time, as it will be shown below,
the use of non-orthogonal time-frequency
transformations with a higher resolution in the
frequency of localization of maxima (of the order
of 1 Hz) significantly increases the accuracy and
efficiency of speaker’s identification by use of
phonograms of short duration.</p>
      <p>
        The general concept of the developed speaker
identification system is based on the data of
classical studies in the neurophysiology of
hearing [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. One of the important factors in the
identification of a speaker by human auditory
analyzers are the individual characteristics of
vowel sounds [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Therefore, when designing the
system, a separate basic module was developed
for automatic extraction of vowel sounds from
speech phonograms. The methodology for its
development is based on deep study of neural
networks.
      </p>
      <p>The speaker’s identification technology used
in the system is based on the automatic
determination of the proximity of the spectral
characteristics of two vowel sounds – [A] and [I],
isolated from two different phonograms. At the
same time, the proximity of the characteristics of
two fragments of vowel sounds in phonograms is
determined on the basis of a special model created
on the basis of a deep learning neural network.
Let's consider some fundamental features of this
technology.</p>
      <p>In the vowel extraction module for each of the
two phonograms, arrays of fragments of sounds
[A] and [I] with a duration of 20 ms are formed.
Further, for all fragments, a non-orthogonal
Morlet wavelet transform is implemented with a
frequency resolution of local maxima of 1 Hz. The
fragments are converted in the frequency range
from 0 to 12000 Hz. Changing the value of the
upper frequency of the range makes it possible to
study the performance of neural networks for their
various configurations and structures.</p>
      <p>The obtained spectra are normalized by
dividing the amplitude of each spectral peak by
the sum of the amplitudes of all spectral peaks for
the entire frequency range</p>
      <p>AiN </p>
      <p>Ai  fi 
 Ai  fi 
,
(2)
where
AiN – normalized amplitude of the spectral
component selected at the i-th scanning step,
Ai(fi) – the amplitude of the spectral component
selected at the i-th scanning step.</p>
      <p>Then one of the classic technological
approaches is used for further forming of the
DataSet. Arrays of fragments consisting of
various combinations of two spectra are formed
from a set of phonograms with the speech of
different speakers. In this case, for fragments of
the spectra of the same speaker, marking "he" is
used, and for different speakers – “not he”. Arrays
of a combination of spectra and a separate array
with a priori known labeling are the basis of the
DataSet for training the neural network.</p>
      <p>Thus, to transform fragments in the frequency
range from 0 to 12000 Hz, the training DataSet
will be composed of fragments containing 24000
spectral amplitudes in various combinations. It
should be noted that for orthogonal
transformations, the number of frequencies is only
480, and in this version, training of a neural
network is a more simple task. Our studies of the
learning process of neural networks with different
structures for a DataSet with fragments containing
24000 frequency amplitudes, in practice, showed
the problematicness of obtaining effective results
for most of the known structures of neural
networks. As the analysis showed, the main
reasons are known factors – gradient attenuation
and retraining. In this the most general statement
the structure of a deep learning neural network
based on convolutional networks was used to
solve the problem in effective way. It should be
noted that parallel studies are conducted on the
basis of fully connected networks. For these
structures a preliminary selection of 50 local
maxima of the spectrum (normalized and then
ranked according to the magnitude of the
amplitude) was applied. For this variant of neural
structures, results were obtained with less
efficiency.</p>
      <p>The keras library (bakend tensorflow) was
used to train the neural network. The created
speaker identification model is actually intended
to solve the binary classification problem.
Fragments of the neural network training process
are shown in Fig. 2</p>
      <p>By use of the developed model, it was possible
to obtain a high efficiency of speaker
identification from the point of view of expert
examination practice. But the practical
implementation of the identification process
requires much time to calculate phonograms. In
addition, the identification model built by the
neural network is a “black-box”. It is not possible
to establish the causal relationships that determine
the speaker's identification. At the same time, the
implementation of software systems into the
practice of examination requires an “internal
conviction of an expert” in the correctness of
made decisions. This conviction can only be based
on well-known classical ideas about the
characteristics of sounds and speech.</p>
      <p>
        Therefore, in the process of designing a
speaker’s identification system on the basis of the
parameters of voice signals (the “Avatar” system
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]), an approach was formed on the basis of
classical concepts associated with a neural
network model.
      </p>
      <p>Analysis of various phonograms showed that
for fragments of spectra with a high probability of
speaker’s identification (above 0.999), the spectra
of 20 ms fragments are very close. So, Fig. 3
shows the spectra in the range 0 – 2500 Hz for the
same speaker (sound – [A]) with the probability
of correct identification 0.9991.</p>
      <p>As it can be seen from the spectra of the same
two fragments, considered in the frequency band
from 0 to 4000 Hz, both local maxima and
formant features of the [A] sound practically
coincide (Fig. 4).</p>
      <p>In accordance with the concept of the adopted
approach to the practical implementation of the
“Avatar” system with a good approximation in
terms of efficiency, a classical heuristic criterion
for the proximity of two spectra was introduced
the sum of the absolute values of the difference
between the normalized amplitudes of the spectra
of the sounds [A] and [I].</p>
      <p>N (3)
B   A  A ,</p>
      <p>1 2
i1
where</p>
      <p>A1 – the value of the averaged amplitude of the
spectrum of the first phonogram,</p>
      <p>A2 – the value of the averaged amplitude of the
spectrum of the second phonogram,</p>
      <p>N – the number of frequencies in a given
frequency range.</p>
      <p>The smaller this value, the closer the
characteristics of the voice of the speaker and
almost any function of the spectrum accordingly.
It includes the frequency of the main tone and
formant features.</p>
      <p>The classical approach to the assessement of
the effectiveness of forensic identification of any
object, including the speaker, is to determine the
magnitude of errors of the first and second kinds.
At the same time, plotting such errors is the most
preferable option for finding out the levels of such
errors. From the physical prerequisites of the tasks
of identification (and verification) of a speaker, it
is known that identification errors significantly
depend on the duration of sound phonograms. In
addition, it is generally accepted that the specific
characteristics of a language and language groups
should affect the effectiveness of identification.
Both of these factors were taken into account in
the studies and in the implementation of the
system under consideration. In particular, the
curves of errors of the first and second kind were
built for a mixture of speech messages of different
speakers, made in different languages – English,
Chinese, Russian and Ukrainian. As well as
messages made separately in English, Russian and
Ukrainian. The dependence of the magnitude of
errors on the duration of phonograms was also
determined. Thus in Fig. 5 and Fig. 6 graphs of
errors of the first and second kinds for several
variants of their construction have been shown.</p>
      <p>Important “technological” factors should be
noted, which, due to the insufficient mass use of
systems for automatic identification of speakers,
are practically not represented in scientific
publications today. At the same time, these factors
play a very significant role in the practice of
expertise (including in its legal aspects).</p>
      <p>The first factor is that the probability of
identification error in the studies under
consideration can be arbitrarily small. Including
less than 0.00001 (0.001%). It is possible under
provided that the fragments of the spectra of
sounds in two phonograms are very close on
average. And this can be observed in real practice
of examination. So, in Fig. 7 shows an illustration
of identification based on two phonograms with a
duration of approximately 60 seconds from the
same speaker.</p>
      <p>The most problematic from the point of view
of making decisions on identification (for any
systems) are phonograms, in which the analysis
results give close values for errors of the first and
second kinds. So in Fig. 8 a similar illustration is
shown. The decision made on the basis of
information about errors of the first and second
kinds when the calculated values are close to the
point of cross of the curves is less justified.</p>
      <p>Obviously, at the point of cross of the curves
of errors of the first and second kinds, the
hypotheses “he” and “not he” are equally
probable. At the same time, the probabilities of
errors of the first and second kinds on the graphs,
although the same, are not great. Such a
contradiction is a consequence of the incorrect
methodology for evaluating errors from the point
of view of the practice of interpreting the
probabilities of errors in these variants. But, due
to the experience of experts, it should be noted
that the expert practice is the area of most
common errors.</p>
      <p>When using integral estimates (for example,
when comparing the averaged spectra of the
fundamental tone for the sound [A], allocated
along the entire length of each of the
phonograms), the method of setting the error
probability threshold for decision making is
usually used. As a rule, this threshold is
determined by the assessment of the probability at
the point of cross of the graphs of errors.</p>
      <p>
        The peculiarity of its application is that,
regardless of the characteristics of such graphs
obtained in this case, the decisions made on the
basis of a given threshold are always subjective.
However, this methodology is generally
recognized in the practice of the probabilistic
approach to decision making [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We believe that
the use of such a threshold, if the obtained value
of the probability of an identification error is close
to it, has a too high degree of subjectivity. It is
advisable to use it in practice only for contrasting
values of the error probability. There is a
significant difference between the probability of
error for a specific measure of the proximity of
two spectra, and the probability of errors of the
first and second kinds at the point of intersection
of their graphs.
      </p>
      <p>At the same time, in the presence of large
arrays of fragments with a duration of 20 ms, used
to identify the speaker, provided that the above
mentioned binary approach to decision-making is
applied (“he” is “not he”), it is possible to build a
less subjective approach.</p>
      <p>In the developed system, a slightly different
approach is applied to the calculation of the
probabilities of identification errors for
phonograms of short duration. It is due to the
technological features of the adopted model.</p>
      <p>An array of fragments of speaker identification
by two phonograms is considered. For
phonograms with a duration from 1 sec. to several
minutes these are arrays of spectra ranging in
number from several hundred to tens of
thousands.</p>
      <p>It should be mentioned, that in the model under
consideration, identification is carried out by
separate fragments, consisting of combinations of
vowel sound spectra extracted from two
phonograms. The output of such a model is the
probabilities of correct identification (“he” – “not
he”) for each pair of compared fragments,
determined by the measure of the proximity of
their spectra. These probabilities are the
collections of discrete random variables.</p>
      <p>Then the statistical average of the probability
of correct identification can be calculated as the
average over the entire array using the
wellknown formula
mx 
1 N</p>
      <p> xi ,
N i0
(4)
where
xi – the assessment of the value of the
probability of correct identification for each i-th
fragment,</p>
      <p>
        N – the number of averaged fragments [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>Due to this approach, the probability of an
error in decision-making is determined by the
statistical averaged mx and the statistical standard
deviation (RMS) from the averaged one, defined
as</p>
      <p> xi  m 2 (5)
S  x ,</p>
      <p>N 1</p>
      <p>Since averaging uses a huge number of
fragments, which are used to determine the
statistical averaged and statistical standard
deviation when comparing two phonograms, they
are subject to the normal distribution law, and we
can use the corresponding probability density
distribution graph to determine the probability of
errors when making a decision.</p>
      <p>The decision error for the totality of the array
of fragments with this approach is the sum over
the probability density distribution in the range
from</p>
      <p>0 ≤ P ≤ 0.5 (Fig. 9). The illustration in Fig. 9
is given for the variant of the statistical average
number of the probability of correct identification
mx &gt; 0.5. For mx &lt; 0.5, the decision error is the
sum over the probability density distribution in
the range from 0.5 ≤ P ≤1.</p>
      <p>In Fig. 9 the average probability of
identification is P = 0.6, the number of fragments
is 894, the standard deviation of the probability of
identification by fragments is S = 0.21.</p>
      <p>The accepted approach to the assessment the
probabilities of identification errors contains only
calculated (experimental) parameters of
fragments of the identification array for two
phonograms. In particular, this mx is the average
value of the identification probability for the
entire array of identification fragments, N is the
number of fragments over which the averaging
was carried out, S is the standard deviation of the
statistical mean of the probabilities of correct
identification. These parameters completely
determine the identification errors. In this case,
the dependence of the identification efficiency on
the duration of the phonograms is automatically
taken into account, which is determined by the
value of the parameter N.</p>
      <p>Depending on the language of the speaker's
speech recorded on the phono-gram, the
parameters mx and S will change. So, for tonal
languages (for example, Chinese), due to the
greater variability of the characteristics of vowel
sounds, S will increase, which, in turn, will
increase the errors identification (this statement is
true for the characteristics of any tonal speech).</p>
      <p>Significant computational complexity is the
disadvantage from the point of view of the
implementation of this probabilistic approach in
an automatic identification system is. This
approach requires, for example, twice more
computational time than the described above
heuristic approach based on the comparisons of
the averaged spectra of vowel sounds. The second
important factor is the “non-standard approach”
when making expert decisions on speaker
identification.
2.1.</p>
    </sec>
    <sec id="sec-3">
      <title>Results and discussion</title>
      <p>In the process of the development of any
speaker’s identification system, the issue arises of
the applicability of the system and the
corresponding methodology to various language
groups. The account of the dependence of the
identification efficiency on the duration of
phonograms, as well as the dependence of the
identification efficiency on specific algorithms.</p>
      <p>From the point of view of eliminating
dependence on various algorithms, the approach
based on models of deep learning neural networks
is the most general one. But it works under the
condition that all information is supplied to the
input of the neural network. At the same time, an
issue arises that determines the completeness of
the model's coverage of various factors. In
particular, whether a particular DataSet can cover
most of the listed above factors.</p>
      <p>Our studies indicate a high probability of
covering most of the factors in the developed
approach. In particular, with the number of
speakers over 15 (male, female voices) and
several language groups, the results practically do
not change with an increase in the number of
speakers in the DataSet.</p>
      <p>Another important factor is the methodology
for work with a set of 20 ms phonogram
fragments, which ensures that there are practically
no parameters in the system that are selected by
an expert and therefore have a subjective
character. This makes it possible to uniformly
solve identification (and verification) problems
regardless of the duration of phonograms and
language groups.</p>
      <p>We believe that the developed identification
methodology has high versatility in the above
mentioned sense.</p>
    </sec>
    <sec id="sec-4">
      <title>3. Conclusions</title>
      <p>An automated system for forensic
identification and verification of the speaker
based on deep learning neural networks has been
developed. In the process of developing a system
based on the comparison of the spectral
characteristics of speech signals, methods have
been suggested and tested that provide the
possibility of identification (verification) of the
speaker by voice messages of short duration.</p>
    </sec>
    <sec id="sec-5">
      <title>4. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Solovyov</surname>
            <given-names>V. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rybalskiy</surname>
            <given-names>O. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhuravel</surname>
            <given-names>V. V.</given-names>
          </string-name>
          <article-title>Verification of fun-damental fitness of neuron networks of the deep educating for the construction of the system of exposure of editing of digital phonograms</article-title>
          .
          <source>Cybernetics and Systems Analysis</source>
          , Vol.
          <volume>56</volume>
          , No. 2,
          <string-name>
            <surname>March</surname>
          </string-name>
          ,
          <year>2020</year>
          , pp.
          <volume>326</volume>
          ¬-
          <fpage>330</fpage>
          . doi:
          <volume>10</volume>
          .1615 / 10.1007 / s10559-020-00249-2.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Solovyov</surname>
            <given-names>V.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rybalskiy</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhuravel</surname>
            <given-names>V.V.</given-names>
          </string-name>
          <article-title>Method of exposure of signs of the digital editing in phonograms with the use of neuron networks of the deep learning</article-title>
          .
          <source>Journal of Automation and Information Sciences</source>
          ,
          <year>2020</year>
          , Vol.
          <volume>52</volume>
          , No.
          <issue>1</issue>
          , pp.
          <fpage>22</fpage>
          -
          <lpage>28</lpage>
          . doi:
          <volume>10</volume>
          .1615 / JAutomatInfScien.v 52.
          <year>i1</year>
          .
          <fpage>30</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>NIST</surname>
            <given-names>USA</given-names>
          </string-name>
          , SRE. Available at: https://www.nist.gov/itl/iad/mig/nist-2019
          <string-name>
            <surname>-</surname>
          </string-name>
          speaker
          <article-title>-recognition-evaluation (accessed 24 August</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] ENFSI working group for forensic speech and audio analysis</article-title>
          . Availa-ble at: http: //www.Enfsi.eu/aboutenfsi/structure/working-groups/speech-andaudio
          <source>(accessed 24 August</source>
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Alexandrova</surname>
            <given-names>Yu.I.</given-names>
          </string-name>
          <string-name>
            <surname>Psychophysiology. M.- S.P .: Nauka</surname>
          </string-name>
          ,
          <year>2006</year>
          , 463 p.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Solovyov</surname>
            <given-names>V.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rybalskyi</surname>
            <given-names>O. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shablya</surname>
            <given-names>A. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhuravel</surname>
            <given-names>V. V.</given-names>
          </string-name>
          <article-title>System of automated search for votes</article-title>
          .
          <source>Informatics and Mathematical Methods in the Model</source>
          .
          <year>2015</year>
          , Vol.
          <volume>5</volume>
          , No.
          <issue>4</issue>
          , pp.
          <fpage>302</fpage>
          -
          <lpage>307</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Mallat</surname>
            <given-names>S.</given-names>
          </string-name>
          <article-title>A wevlet tour of signal processing</article-title>
          .
          <source>Асаdemic Ргеss</source>
          , New Yогk.
          <year>1999</year>
          ,
          <volume>671</volume>
          p.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bondarenko</surname>
            <given-names>M.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Driuchenko</surname>
            <given-names>A.Ya.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shabanov-Kushnarenko Yu</surname>
          </string-name>
          .P.
          <article-title>Vowel sounds in theory and experiment</article-title>
          .
          <source>Kh .: Khark. nat. un-t radioelektroniki</source>
          ,
          <year>2002</year>
          , 348 p.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Rybalskiy</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solovyov</surname>
            <given-names>V.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cherniavskyi</surname>
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhuravel</surname>
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheleznyak</surname>
            <given-names>V.K.</given-names>
          </string-name>
          <article-title>A probabilistic approach to making expert decisions on the analysis of complex objects</article-title>
          .
          <source>Bulletin of the National Academy of Sciences of Belarus. A Series of Physical and Technical Sciences. 2019</source>
          , Vol.
          <volume>64</volume>
          , No.
          <issue>3</issue>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>352</lpage>
          . doi:
          <volume>10</volume>
          . 29235 /
          <fpage>15</fpage>
          - 8358-2019-64-3-
          <fpage>346</fpage>
          -352.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>Handbucher der industriellen Messtechnik / 3 Auflage/ Herasgeber prof</article-title>
          ., Dr., P. Profos / 1 Auflage. Vulkan-Verlag,
          <year>Essen</year>
          .
          <year>1984</year>
          ,
          <year>491p</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>