<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Voice command recognition for noisy environments by means of cross-correlation portraits</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>A.I. Armer</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E.Yu. Galitskaya</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>N.A. Krasheninnikova</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ulyanovsk State Technical University</institution>
          ,
          <addr-line>Severny Venets St., 32, Ulyanovsk, 432027</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>22</lpage>
      <abstract>
        <p>Methods of voice command (VC) recognition in heavy noise environments are required for precise work of speech information systems on the factory floor and in transport. The paper considers a speaker-dependent way of VC recognition for VCs belonging to a limited vocabulary and being recognized in heavy noise environments. For this purpose, VCs are transformed into cross-correlation portraits (CCPs), i.e. special images. The VC under recognition is referred to a class with a minimal distance (metric) between CCP of this command and model CCPs of the class. The authors elaborated algorithms for VC transformation into CCPs, a method for defining VC boundaries, ways of model command optimization and metric choice. As a result, a rather precise VC recognition in heavy noise environment was obtained.</p>
      </abstract>
      <kwd-group>
        <kwd>voice command</kwd>
        <kwd>intensive noise</kwd>
        <kwd>recognition</kwd>
        <kwd>cross-correlation portrait</kwd>
        <kwd>metric</kwd>
        <kwd>precise definition of boundaries</kwd>
        <kwd>model command</kwd>
        <kwd>optimization of VC library</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The growth of production and transport intensity leads to increase in operator burden. To reduce such workload,
speech information systems are used. However, these systems often have to recognize VC precisely, especially for
noisy environments. At present, a large number of speech recognition systems functioning in nearly noiseless
environment have been developed. They include, for example, IBM Via Voice, its recognition accuracy is reported to
be 97% and its recognition vocabulary includes up to 2,000 VCs; Dragon NaturallySpeaking or Dragon for PC, this
software package accurately recognizes 70% of the vocabulary, which includes nearly 60,000 words; L&amp;H Voice
XPress, its accuracy is in the range of 90%-98% and its vocabulary size is nearly 1,000 words, etc. There are also
user-friendly systems of continuous speech understanding and processing, such as VocalIQ, Siri, Google Now and
Cortana. To compare VocalIQ with Siri, Google Now and Cortana the systems were given multiaspect requests in
a natural language [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The correct recognition rate was more than 90% for VocalIQ, while Google Now, Siri and
Cortana showed only 20% accuracy. Among home-grown technologies it is necessary to mention VoiceCom STC. It
is reported to recognize 100-200 VCs in a speaker-dependent version and 30-50 VCs in a speaker independent one
with accuracy 98%. However, these systems do not accurately work even in low loise environment. Recognition
systems for VCs from a limited vocabulary under acoustic noise are currently being developed mainly for aviation
and are used in voice control and flight control devices. Performance quality of such systems today is from 90 up
tp 98% of accurate VC recognition, depending on the test conditions and vocabulary size. Almost all tested systems
are speaker-dependent. According to the Air Force Research Laboratory - Wright-Patterson Air Force Base, flight
tests of an ITT VRS-1290 speaker dependent, continuous speech recognition system and a Verbex VAT31 showed the
following results: average word accuracy for VRS-1290 was 92-98%, if the vocabulary consisted of 50 commands;
average word accuracy for VAT31 was up to 97% (no information on vocabulary size is available). In 1997, flight test
results of the VC recognition system produced by National Research Council (Canada) were obtained. The system
was integrated into Bell 412HP Avionics Management System and showed an average 95% accuracy for vocabulary
consisting of 80 words, which were divided into 24 groups. According to the Smiths Industries Speech
Recognition Module system built into the CAMU of the Eurofighter, the accuracy of VC recognition in a standard aircraft
flight is at least 95% for a vocabulary consisting of 250 words, 25 of which can be simultaneously active. Currently,
Thales Avionics develops a VC recognition system for Rafale fighters. The VC recognition accuracy is required to
be above 95% for a vocabulary of 50-300 words. A 5-th generation jet fighter F-35 was equipped with DynaSpeak
      </p>
      <p>VC recognition system developed by SRT International. The developers report the recognition accuracy to be 98%.
A multipurpose 4-th generation Eurofighter is equipped with a voice control system developed by Logica. The
vocabulary consists of 250 words, and the average VC accuracy is not less than 95%. The developers declare, that for
the export version of the Rafale Block 05t, Thales Avionics has developed a speech control system with recognition
accuracy not less than 95% for a 300 VC vocabulary, but no information on its implementation is available. Patent US
6529866 B1, 4 March 2003, The United States of America as Represented by the Secretary of the Navy, describes a
method and system for transformation of an audio signal into speech. Audio signals are said to contain both VC units
and noise, but test and implementation information is not available. Patent WO 1999040571 A1, 3 February 1999,
Qualcomm Incorporated, describing a system and method for improving speech recognition accuracy in noisy
environment also provides no test or implementation data. Among home-grown technologies the following ones should
be noted. First of all, it is a VC recognition system tested on the Mikoyan MiG-29 (Fulcrum). Recognition
accuracy is reported to be 56-81%, no information on the vocabulary is available. Patent RF 2267820 1, 25 April 2006,
Ulyanovsk State Technical University. Recognition accuracy is reported to be 92%, vocabulary size is 23 VCs, and
noise level is 3dB. No information on implementation is available. Patent RF 2271578 2, 10 March 2006, Speech
Technology Center. The invention relates to speech analysis under adverse environmental conditions, e.g. in moving
transport or high level noisy workplaces. No test information is available. Despite the available developments, there
is no information on the actual application of VC recognition systems in avionics, since in real flights the systems
developed showed substantially less efficiency than anticipated. Thus, developing VC recognition systems for noisy
environments remains a challenging task. This paper examines a speaker dependent technique of VC recognition for
a limited vocabulary. A method of VC transformation into portraits, i.e. images, is used.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods of VC recognition</title>
      <p>
        The problems of speech recognition, in particular VC recognition, are widely discussed in modern literature.
The first methods of automatic sound recognition were obtained in the first half of the 20-th century [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Among
speech recognition techniques one can distinguish the following approaches: spectral methods [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6 ref7">3, 4, 5, 6, 7</xref>
        ], wavelet
transform [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], statistical methods [
        <xref ref-type="bibr" rid="ref10 ref11 ref5 ref8 ref9">5, 8, 9, 10, 11</xref>
        ], and neural networks [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ].
      </p>
      <p>
        This paper deals with VC recognition based on their transformation into portraits, i.e. flat images, and further
implementation of image processing techniques [
        <xref ref-type="bibr" rid="ref14 ref15 ref16 ref17 ref18">14, 15, 16, 17, 18</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Autocorrelation portraits</title>
      <p>Let S = s0, s1, s2, s3, ..., sN−1 be digital VC readouts. Then, a two-dimensional image X(i, k) = {xik : i =
1, 2, 3, ...; k = 1..K} will be its autocorrelation portrait (ACP). This image is obtained in the following way. Let
us divide VC S into M segments and perform the following transformations</p>
      <p>X(i, k) =</p>
      <p>Cov(S n, S n+k) ,
σnσn+k
(1)
where Cov(S n, S n+k) is a sample covariation of signal S intervals S n, S n+k, which are spaced kΔt apart, σ2, σ2n+k are
n
sample dispersions of segments S n, S n+k respectively. Thus, the k−th element of the i−th ACP line is equal to the
correlation coefficient between the i−th segment S i and the segment shifted left with respect to S i on k readouts. Fig.
1 shows ACP examples.</p>
      <p>
        Let us note some ACP characteristics, which make them favorable for VC recognition. VC portraits are unique,
i.e. ACPs of different VCs are unlike, whereas ACPs of the same VCs pronounced at different time intervals are
the same. Autocorrelation transformation normalizes a signal, as a result ACPs are nearly insensitive to noisiness
and slowly varying additives. If we consider additive white noise with dispersion σθ2, then its ACPs and VC ACP
readouts distorted by noise will differ by a constant factor. However, ACPs also have some negative characteristics,
e.g. the dependence of element brightness on the differences in the tone of VC pronunciation, as well as geometric
ACP distortions due to variations in speech rate. These distortions can be steadied by modifying ACP development,
e.g. taking into account loudness extremum. VC recognition by their ACPs is conducted in the following way. ACPs
of model VCs are stored in the memory. VC under recognition is transformed into ACP and it is referred to the class
with a minimal distance between its model portrait and ACP of a recognized VC. This distance (metric) between
two ACPs (i.e. images) is calculated as follows. At first, two images are aligned, i.e. for each line of one image a
corresponding line of another image is found. The average distance (e.g. Euclidean) between the corresponding lines
is considered to be the distance between the portraits. Such a correspondence for ACP of one and the same command
means the proximity of VC fragments, so the distance is relatively small, since it only occurs from the difference in
pronunciation and surrounding noise. If ACPs of different VCs are compared, then this distance is usually much more
visible due to the larger difference in sounds. In the process of command alignment dynamic programming based on
minimum distance criterion was applied. While testing the accuracy of VC recognition, commands were pronounced
by the speaker in real time. The vocabulary used consisted of ten VC groups, and there were 4-23 aviation commands
in each group. In total, the vocabulary included more than 100 VCs. Aircraft engine noise recorded in a flight mode
was used as a background and reference noise, the signal-to-noise ratio was 5-0 dB. Four male speakers took part in
the tests. Before the experiment each speaker recorded model VCs, each VC belonging to the given vocabulary was
pronounced twice. During the experiment on VC recognition each speaker pronounced all the commands from the
given vocabulary three times, all in all, more than 1,200 VCs were recorded during the experiment. Average command
accuracy was more than 95%. However, further processing has shown that the probability of accurate VC recognition
can be significantly reduced in the course of time. This problem is connected with model aging, i.e. speaker’s voice
pattern can change with time, and previously pronounced command models will not reflect the peculiarities of the
speaker’s voice at the very time of VC recognition. Therefore, it is required to update the commands from time
to time (e.g. before the flight), which, of course, has certain inconveniences. One VC model does not reflect all
the possible variants of its pronunciation, so the number of VCs was increased, i.e. the speaker pronounced each
VC more than once at different periods of time. The totality of all these patterns somehow reflected pronunciation
diversity. However, the increase in model number complicates and slows down the recognition algorithm, but it is
permissible only to a certain extent. Therefore, the model number should be limited. Besides, these models should
reflect the pronunciation diversity as much as possible. It turns out, that recognition accuracy depends greatly on the
correctness of model choice, and recognition deviations can be more than 10%. Thus, among several pronunciations
it is necessary to choose a certain number of VCs as model ones, so that the obtained model library contributed to the
best VC recognition accuracy. This problem of model library optimization was examined in [
        <xref ref-type="bibr" rid="ref19 ref20 ref21">19, 20, 21</xref>
        ]. Technically
it is impossible to conduct complete enumeration of all library patterns. That is why, a method of direct enumeration
giving an almost optimal result has been developed. Sometimes it is possible to change the VCs themselves, using
their synonyms. This problem was also considered and its solution was found while analyzing the synonym rings.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Cross-correlation portraits</title>
      <p>
        Another way to decrease the impact of VC pronunciation variability is to use a different kind of portraits instead
of ACPs. In the process of ACP development correlation coefficients between the segments of the same VC
(autocorrelation) are found. When ACPs are used for recognition, the distances between the ACP of a recognized command
and the ACP of a model command are found. If the distance between the ACP of the command under recognition and
the ACP of its model is found, the ACPs of two different pronunciations of this command will be compared. These
ACPs can significantly differ from each other (the distance will be large). Therefore, when comparing portraits it is
desirable to minimize the difference in pronunciation. For this purpose, it is necessary for pronunciation variability
to be somehow reflected in portraits. Let us consider a cross-correlation portrait (CCP), which consists of correlation
coefficients between segments of two VCs (cross-correlation) [
        <xref ref-type="bibr" rid="ref15 ref16 ref17 ref22">15, 16, 17, 22</xref>
        ]. Let there be two VCs S 1 and S 2.
Let us segment each command into M segments of the same length and determine the sample correlation coefficients
xik between the i−th segment of VC S 1 and a VC segment S 2, beginning with the k−th readout of the VC S 2 i−th
segment. As a result, we get a two-dimensional array (image) X = {xik}, called a CCP of VCs S 1 and S 2. Let us
consider CCP development in detail. As an example, let us consider the CCP development of two pronunciations of
one avionics VC, the first pronunciation is S 1 and the second pronunciation is S 2. Let us divide each VC into equal
segments, whereas N1 is the length of each interval for signal S 1, and N2 is the length of each interval for signal
S 2. Let N = min{N1, N2} be the minimal of these lengths. While specifying the number of intervals for each
command M it should be taken into account that if the segment length is too small it will not include the whole phoneme;
otherwise, if the segment length is big enough it will include several phonemes. Such segmentation will negatively
affect the correlation coefficient between separate phonemes in different VCs while developing CCPs. Let’s determine
correlation coefficients of signal S 1 i-th segment and signal S 2 i-th segment, shifted k = 0..K readouts right.
xik = N1 PNj=−01 S 1i·N1+jS 2i·N2+j+k−μ1iμ2i,k ,
      </p>
      <p>σ1iσ2i,k
μ1i = N1 P Nj=−01 S 1i·N1+ j,
μ2i,k = N1 P Nj=−01 S 2i·N2+ j+k,
σ1i2 = N1 P Nj=−01 S 1i2·N1+ j − μ1i2,
σ2i2,k = N1 P Nj=−01 S 2i2·N2+ j+k − μ2i2,k.
(2)
(3)
(4)
(5)
(6)
While choosing parameter K, it is necessary to take into account the following fact: if its value increases, than value
xik decreases. It is connected with correlation reduction of VC readouts along the line. This property proves the
inadvisability of using large values K while developing CCPs (large K means that K &gt; N).</p>
      <p>Obviously, if CCPs of the same pronunciation (S 1 = S 2 = S ) are developed, we get the ACP of a VC S . It is
desirable to examine the CCP of two pronunciations of one and the same command. It depends on two pronunciations,
so the pronunciation variability affects the portrait form. Fig. 2 shows CCPs of several VCs. For example, in the
picture Manevr3 + Manevr4 ’plus’ means that this very CCP was obtained from the third and fourth pronunciations
of the VC ”Manevr”.</p>
      <p>Note, that CCP characteristics are similar to those of ACP. But CCPs are less pronunciation dependent, as they
combine two different pronunciations. VC recognition by means of CCPs is carried out in the same way as recognition
by means of ACPs. For each VC, a model CCP made of two pronunciations of this VC is developed. These model
CCPs are stored in the memory. For the VC under recognition CCPs are developed with one of pronunciations of each
”Navigatsiya1+Navigatsiya2”</p>
      <p>”Navigatsiya3+Navigatsiya4”
”Noised
Navigatsiya1+Navigatsiya2”</p>
      <p>”Noised
Navigatsiya3+Navigatsiya4”
command group, then the distance between this CCP and the model CCP is found. The recognized VC is related to
the group with the smallest distance.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Optimization of voice command recognition by means of their CCPs</title>
      <p>The CCPs used have a number of characteristics, which come from both the properties of the speech signals
themselves and the structure of their CCP development. Let us consider some techniques increasing the recognition
accuracy by means of CCPs.</p>
      <sec id="sec-5-1">
        <title>5.1. Noisy models</title>
        <p>If a VC under recognition is too noisy, it increases the distance from its CCP to its model portrait formed by
means of noiseless pronunciations. Therefore, ’noisy models’ were used in the experiment, i.e. artificial noise was
added to the model commands. It came from the microphone placed far from the operator. As a result, the distorted
models and the command under recognition contained approximately the same noise, which significantly increased
the recognition accuracy.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Precise definition of boundaries</title>
        <p>While developing portraits, it is desirable for the VC time boundaries to be defined as precisely as possible. Then a
more accurate portrait alignment can be attained. Among several known techniques of useful signal detection, the one,
which shows the most accurate recognition results on the background of noise, was chosen. Besides, after definition
of VC boundaries by means of this technique some boundary adjustments were made, which resulted in recognition
accuracy.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Pause removal</title>
        <p>In some VCs, e.g. those consisting of two words, there are micro-pauses between speech units. These pauses can
differ in duration, but they do not contain any information. So, a special method for their removal was developed.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Optimization of portrait width</title>
        <p>Portrait width, i.e. the line length, can be chosen arbitrary. So, it is desirable to choose the optimal length, which
contributes to the best recognition accuracy. It turned out, the line length in the portrait of a VC under consideration
should be equal to K = D/(5M), where D is the length of the recognized VC, M is the number of lines in a portrait.
The line lengths of model CCPs are a bit longer, but they are no less than K.</p>
      </sec>
      <sec id="sec-5-5">
        <title>5.5. Choice of metric</title>
        <p>VC recognition by means of their CCP is based on detection of the portraits, which are as similar as possible.
Hence, there appears a problem to define the distance between two CCPs, i.e. metric defined on CCP. This distance
is considered to be equal to the average distance between the corresponding CCP lines. Moreover, any metric defined
on the lines, i.e. on finite sequences or vectors, can be used. Twelve known metrics (namely, Euclidean, Hilbert,
Zhuravlev, etc.) and their variants were tested. For the purpose of the problem under consideration, five metrics
showed the best results: Zhuravlev method (for ε = 10, ε = 20 and ε = 30), the Ruzicka distance and the Bray-Curtis
distance. Besides, analyzing the recognition results obtained while using these metrics it was found out that certain
recognition errors corresponded to certain metrics. Therefore, it is possible to improve recognition accuracy by using,
for example, two metrics. If the recognition results coincide, then the command is considered to be recognized; if
the recognition results differ, the command should be considered unrecognized. In such a case, the speaker should
pronounce the command once again.</p>
      </sec>
      <sec id="sec-5-6">
        <title>5.6. Optimization of a model library</title>
        <p>As in the case of VC recognition by means of CCPs, the words included in the model library significantly affect
the recognition accuracy. Therefore, it is required to optimize the model portrait library while recognizing VCs by
means of CCPs.</p>
      </sec>
      <sec id="sec-5-7">
        <title>5.7. Fourier analysis</title>
        <p>Each CCP line is a sequence of correlation function sample values. Because of speech signal quasi-periodicity,
the correlation function turns out to be similar to periodic. This quality was used to improve the portrait quality by
removing insignificant harmonics from each CCP line spectrum. This operation was performed by means of FFT. The
isolation of the most fundamental harmonics for each CCP line reduced the influence of speech signal pronunciation
variability.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <p>The experiments showed that using CCPs with the described above modifications significantly reduced the effect
of VC pronunciation variability and model aging. The recognition accuracy was nearly the same as in the ACP
recognition on newly-pronounced models. Thus, to evaluate the efficiency of the suggested method, an experiment
was conducted. The recognition was tested on two groups of VCs consisting of 10 commands each. The first group
included single-word commands, the second group of VCs included both single-word and two-word commands.
Each VC was pronounced 100 times by a woman-speaker. The experimental results are represented in Table 1. The
maximum VC recognition accuracy was 95.6%.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion References</title>
      <p>The present work suggests and examines a speaker-dependent method for recognizing voice commands from
a limited vocabulary in conditions of intense acoustic noise, e.g. on the background of an aircraft engine. This
method implies transformation of digitized commands into certain images and further application of image processing
methods. The method underwent various modifications in order to increase the recognition accuracy. Tests on a large
number of voice commands showed rather high efficiency of the suggested method.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Businessinsider</surname>
          </string-name>
          .
          <article-title>How apples vocaliq ai works [Electronic resource]</article-title>
          . ”-
          <year>2017</year>
          . ”- URL: http://uk.businessinsider.
          <article-title>com/how-applesvocaliq-ai-</article-title>
          <string-name>
            <surname>works-</surname>
          </string-name>
          2016-5.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Rabiner</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <article-title>Tsifrovaya obrabotka rechevykh signalov [Digital processing of speech signals]: translated from English</article-title>
          . Edited by
          <string-name>
            <given-names>M.V.</given-names>
            <surname>Nazarov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yu.N.</given-names>
            <surname>Prokhorov</surname>
          </string-name>
          / L.R. Rabiner,
          <string-name>
            <given-names>R.V.</given-names>
            <surname>Shafer</surname>
          </string-name>
          . ”- Moscow, Russia. : Nauka,
          <year>1981</year>
          . ”- P.
          <year>495</year>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Boykov</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>Primenenie veyvlet-analiza signala v sisteme raspoznavaniya rechi [Wavelet analysis in speech recognition] / F</article-title>
          .G. Boykov, Starozhilova T.K. // Trudy mezhdunarodnoy konferentsii Dialog 2003
          <source>[Proceedings of the international conference Dialogue</source>
          <year>2003</year>
          ]. ”- Zvenigorod, Russia. ”-
          <year>2003</year>
          . ”- Pp.
          <fpage>12</fpage>
          -
          <lpage>19</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Gudonavichyus</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Raspoznavanie rechevykh signalov po ikh strukturnym svoystvam [Speech signal recognition by means of their structural characteristics] /</article-title>
          <string-name>
            <given-names>R.V.</given-names>
            <surname>Gudonavichyus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.P.</given-names>
            <surname>Kemeshis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.B.</given-names>
            <surname>Chitavichyus</surname>
          </string-name>
          . ”-Leningrad, USSR. : Energiya,
          <year>1977</year>
          . ”-P.
          <year>64</year>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Myasnikova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <article-title>Ob”ektivnoe raspoznavanie zvukov rechi [Objective recognition of speech sounds] /</article-title>
          <string-name>
            <given-names>E.N.</given-names>
            <surname>Myasnikova</surname>
          </string-name>
          . ”- Leningrad, USSR. : Energiya,
          <year>1967</year>
          . ”- P.
          <year>148</year>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Pikone</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <article-title>Metody modelirovaniya signala v raspoznavanii rechi [Signal modeling methods in speech recognition] / D</article-title>
          . Pikone. ”-Kemerovo, Russia,
          <year>2000</year>
          . ”- P.
          <year>79</year>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Potapova</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Rech</surname>
          </string-name>
          <article-title>': kommunikatsiya, informatsiya, kibernetika [Speech: communication, information</article-title>
          , cybernetics] / R.K. Potapova. ”- Moscow, Russia.:
          <source>Radio i svyaz'</source>
          ,
          <year>1997</year>
          . ”- P.
          <year>568</year>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Sorokin</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Skrytye markovskie modeli v raspoznavanii rechi [Hidden markov models in speech recognition] /</article-title>
          <string-name>
            <given-names>V.N.</given-names>
            <surname>Sorokin</surname>
          </string-name>
          , V.A. Sukhanov // Rechevaya informatika [Speech informatics].
          <source>Collected papers</source>
          edited by V.V. Zyablov. ”- Moscow, Russia. ”-
          <year>1989</year>
          . ”- Pp.
          <fpage>104</fpage>
          -
          <lpage>118</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Peinado</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Discriminative codebook design using multiple vector quantization in hmm-based speech recognizers / A</article-title>
          . Peinado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Segura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rubio</surname>
          </string-name>
          [et al.] // IEEE Trans. Speech and
          <string-name>
            <given-names>Audio</given-names>
            <surname>Processing</surname>
          </string-name>
          . ”-
          <year>1996</year>
          . ”- Vol. IV, No.
          <volume>2</volume>
          . ”- Pp.
          <fpage>89</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Jelinek</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>Statistical Methods for Speech Recognition /</article-title>
          F Jelinek. ”- Cambridge. : MIT Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Shahshahani</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>A markov random field approach to bayesian speaker adaptation / B</article-title>
          . Shahshahani // IEEE Trans. Speech and
          <string-name>
            <given-names>Audio</given-names>
            <surname>Processing</surname>
          </string-name>
          . ”-
          <year>1997</year>
          . ”- Vol. V, No.
          <volume>2</volume>
          . ”- Pp.
          <fpage>183</fpage>
          -
          <lpage>191</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Fedyaev</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          <article-title>Neyrosetevoy interpretator rechevykh komand dlya upravleniya programmnymi sistemami [Neural network interpreter of voice commands for program system processing] /</article-title>
          <string-name>
            <given-names>O.I.</given-names>
            <surname>Fedyaev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            <surname>Gladunov</surname>
          </string-name>
          <article-title>// Proceedings of the 7th All-Russian conference ”Neural computers and their usage”, eduted by A.I. Galushkin</article-title>
          . ”- Moscow, Russia. ”-
          <year>2001</year>
          . ”- Pp.
          <fpage>298</fpage>
          -
          <lpage>301</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Lippmann</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>Neural classifiers useful for speech recognition</article-title>
          / R. Lippmann, B. Gold // in.
          <source>Proc. IEEE First Int. Conf. Neural Net</source>
          . ”-Vol. IV. ”
          <article-title>- 1987</article-title>
          . ”- Pp.
          <fpage>417</fpage>
          -
          <lpage>422</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Krasheninnikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Raspoznavanie rechevykh komand na fone intensivnykh shumov s pomoshch'yu avtokorrelyatsionnykh portretov [Speech command recognition on the background of noise using autocorrelation portraits] /</article-title>
          <string-name>
            <given-names>V.R.</given-names>
            <surname>Krasheninnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.I.</given-names>
            <surname>Armer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Krasheninnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.V.</given-names>
            <surname>Khvostov</surname>
          </string-name>
          // Naukoemkie tekhnologii. ”
          <article-title>- 2007</article-title>
          . ”- 9. ”- Pp.
          <fpage>65</fpage>
          -
          <lpage>76</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Krasheninnikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Cross-correlation portraits of voice signals in the problem of recognizing voice commands according to patterns / V.</article-title>
          <string-name>
            <given-names>R.</given-names>
            <surname>Krasheninnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.I.</given-names>
            <surname>Armer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.V.</given-names>
            <surname>Kuznetsov</surname>
          </string-name>
          , E.Yu Lebedeva // Pattern Recognition and
          <string-name>
            <given-names>Image</given-names>
            <surname>Analysis</surname>
          </string-name>
          . ”-
          <year>2011</year>
          . ”- Vol.
          <volume>21</volume>
          , No.
          <volume>2</volume>
          . ”- Pp.
          <fpage>192</fpage>
          -
          <lpage>194</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Krasheninnikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Variatsiya granits rechevykh komand dlya uluchsheniya raspoznavaniya rechevykh komand po ikh krosskorrelyatsionnym portretam [Voice command variability for voice command recognition accuracy by means of their cross-correlation portraits] /</article-title>
          <string-name>
            <given-names>V.R.</given-names>
            <surname>Krasheninnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yu. Lebedeva</surname>
          </string-name>
          , V.K. Kapyrin // Izvestiya Samarskogo nauchnogo tsentra RAN. ”
          <article-title>-2013</article-title>
          . ”-Vol.
          <volume>4</volume>
          (
          <issue>4</issue>
          ). ”-Pp.
          <fpage>928</fpage>
          -
          <lpage>930</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Krasheninnikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Povyshenie veroyatnosti pravil'nogo raspoznavaniya signalov po ikh krosskorrelyatsionnym portretam [Improvement of signal recognition accuracy by means of their cross-correlation portraits] /</article-title>
          <string-name>
            <given-names>V.R.</given-names>
            <surname>Krasheninnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Krasheninnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yu</surname>
          </string-name>
          . Galitskaya // Radiotekhnika. ”-
          <year>2014</year>
          . ”- Vol.
          <volume>7</volume>
          . ”- Pp.
          <fpage>107</fpage>
          -
          <lpage>110</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Vasil</surname>
          </string-name>
          <article-title>'ev, K. Statisticheskiy analiz izobrazheniy [Statistical image analysis] / K.K</article-title>
          . Vasil'ev,
          <string-name>
            <given-names>V.R.</given-names>
            <surname>Krasheninnikov</surname>
          </string-name>
          . ”- Ulyanovsk, Russia. : UlSTU,
          <year>2014</year>
          . ”- P.
          <year>216</year>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Armer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Ispol'zovanie ontologii dlya formirovaniya naborov etalonov rechevykh komand v zadache raspoznavaniya rechevykh komand na fone shumov [Using ontologies to generate a set of voice commands in the problem of speech recognition of voice commands in background noise] /</article-title>
          <string-name>
            <given-names>A.I.</given-names>
            <surname>Armer</surname>
          </string-name>
          , V.S. Moshkin // Radiotechnika. ”-
          <year>2016</year>
          . ”- Vol.
          <volume>9</volume>
          . ”- Pp.
          <fpage>72</fpage>
          -
          <lpage>77</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Armer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Podkhod k formirovaniyu naborov etalonov rechevykh komand s ispol'zovaniem ontologii [Formation of voice command model groups with ontology] /</article-title>
          <string-name>
            <given-names>A.I.</given-names>
            <surname>Armer</surname>
          </string-name>
          , V.S. Moshkin // Ontologiya proektirovaniya. ”
          <article-title>- 2016</article-title>
          . ”- Vol.
          <volume>6</volume>
          . ”- Pp.
          <fpage>270</fpage>
          -
          <lpage>277</lpage>
          . (in Russian)
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Krasheninnikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Optimization of dictionary and model library for recognition of speech commands / V.R</article-title>
          . Krasheninnikov,
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Krasheninnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.V.</given-names>
            <surname>Kuznetsov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yu</surname>
          </string-name>
          . Lebedeva // Pattern Recognition and
          <string-name>
            <given-names>Image</given-names>
            <surname>Analysis</surname>
          </string-name>
          . ”-
          <year>2011</year>
          . ”- Vol.
          <volume>21</volume>
          , No.
          <volume>3</volume>
          . ”- Pp.
          <fpage>505</fpage>
          -
          <lpage>507</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Krasheninnikov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>Optimization of dictionary and model library for recognition of speech commands based on cross-correlation portraits / V.R</article-title>
          . Krasheninnikov,
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Krasheninnikova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.V.</given-names>
            <surname>Kuznetsov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yu</surname>
          </string-name>
          . Lebedeva // Pattern Recognition and
          <string-name>
            <given-names>Image</given-names>
            <surname>Analysis</surname>
          </string-name>
          . ”-
          <year>2013</year>
          . ”- Vol.
          <volume>23</volume>
          , No.
          <volume>1</volume>
          . ”- Pp.
          <fpage>80</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>