<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Testing of the Speech Recognition Systems Using Russian Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Konstantin Aksyonov</string-name>
          <email>bpsim.dss@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Kalinin</string-name>
          <email>igor_kalinin@hotmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey Karavaev</string-name>
          <email>karav197@yandex.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dmitry Antipin</string-name>
          <email>diantigers@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilya Evdokimov</string-name>
          <email>psp720@mail.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Uriy Chiryshev</string-name>
          <email>iurii.chiryshev@mail.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tamara Afanaseva</string-name>
          <email>t.afanaseva42@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aleksandr Shevchuk</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Egor Talancev</string-name>
          <email>i.spyric@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LLC "UralInnovation"</institution>
          ,
          <addr-line>Ekaterinburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ural Federal University</institution>
          ,
          <addr-line>Ekaterinburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article includes the results of testing an existing speech recognition systems, using the Russian language model and dictionary. Objects of the research are question-answering systems, call-center processes, methods and systems for language processing. Several problems with Russian language speech recognition were identified and studied.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Introduction
In the research, algorithms for speech recognition are considered. These algorithmsare used by the following systems
for speech recognition: Sphinx [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Yandex.Speech Kit [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ], Google Cloud Speech-to-Text [
        <xref ref-type="bibr" rid="ref1 ref5 ref6">1, 5-6</xref>
        ].
CMU Sphinx [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] – is a major open-source cross-platform project for speech recognition, developed in Carnegie Mellon
University. It includes series of the Sphinx systems and program for studying acoustic model, which is called Sphinx
Train.
      </p>
      <p>Currently, CMU Sphinx includes language models for the several languages: English, Russian, German, Chinese and
others. It gives an opportunity to create acoustic models for other languages. The project uses BSD license, which
allows distributing the product commercially, the system also includes tools for speech recognition (keyword definition,
pronunciation evaluation). The development is still in progress.</p>
      <p>
        Yandex.Speech Kit – is a complex of speech technologies, developed by Yandex Company, it includes speech
recognition and synthesis. Speech Kit is used as the cloud service Speech Kit Cloud and Speech Kit Mobile [
        <xref ref-type="bibr" rid="ref7 ref8">7-8</xref>
        ].
Speech Kit Cloud SDK – is a program, which allows developers to use instruments for speech recognition made by
Yandex. Infrastructure of the service’s design based on probability of high loads, to provide access and trouble-free
system operations.
      </p>
      <p>Interaction with Speech Kit Cloud is performs HTTP API and includes different functions:
1. Interactive Voice Response
2. Automatic calls to transmit information about new services, to confirm an order or delivery, to
remind about the record, the collection of meter readings.
3. Inquiries by phone without operator's participation, recording for reception and maintenance.
4. The voice interface of "smart house" systems.
5. The voice interface of robots.</p>
      <p>6. Site management, using voice.</p>
      <p>
        Speech Kit Cloud is designed to recognize small speech fragments about 30 seconds long [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Speech Kit Mobile SDK [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is a program that allows to embed speech recognition and synthesis into a mobile
application on iOS, Android or WindowsPhone. Speech Kit Mobile SDK is used in the following Yandex services:
Yandex Search, Browser, Taxi, Maps, Navigator, Translator, Market, Music, Launcher, Auto, Keyboard.
Google Cloud Speech-to-Text is a technology, which allows developers to convert audio into text by applying neural
network models, using the API, which Google Cloud Speech provides. The API supports 120 languages and options
for supporting the user base. The system allows using voice commands, managing, rewriting an audio from call centers
and more. It can handle streaming or pre-recorded audio in real time using Google's machine learning technology.
1
      </p>
      <p>Deploying process of the speech recognition systems for testing the methods and
algorithms
Deploying CMU Sphinx and Google Speech Recognition.</p>
      <p>To deploy the Pocket Sphinx system and Google Speech Recognition, a solution from a third-party developer was
used, it is represented as a Speech Recognition Python package. It can be installed using the Pip tool:
pip install SpeechRecognition
pip install pocketsphinx
The installation requires Python 3.3 and later, Pip. The Python wrapper package for the Pocket Sphinx system must
be installed as well:
We used Python 3.6. Installation was performed in the Windows terminal from PyCharm (using virtual environment)
(Figure 1).
https://sourceforge.net/projects/cmusphinx/files/Acoustic%20and%20Language%20Models/Russian/
Files of one of the downloaded models should be placed in the pocket sphinx-data folder in the folder of the
speech_recognition package (Figure 2).
After that, the launch of the following Python script will display the result of speech recognition in both systems (the
data from the pre-recorded audio file is recognized - Figure 3).</p>
      <p>The GitHub repository of the SpeechRecognition package and installation instructions:
https://github.com/Uberi/speech_recognition
In order to work with the system, a key must be obtained on the site: https://developer.tech.yandex.ru/. The test key
(Figure.4) only works for a month and has a limit on the number of requests.</p>
      <p>
        Speech Kit Cloud API is a WebAPI, which means that in order to work with the system, a POST request
should be sent, and then an XML file with the recognition results will processed (Figures 5-6). Instructions for
interacting with Speech Kit Cloud API are presented in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>import urllib.request as req
from xml.dom import minidom
key = '1e692527-ad23-4fdb-b463-b34e545f9a13'
uuid = 'aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaab'
file = 'english.wav'
defread_file():
with open(file, 'rb') as f:
return f.read()
defget_text_xml(data):</p>
      <p>request = req.Request(url='https://asr.yandex.net/asr_xml?uuid=' + uuid +
'&amp;key=' + key + '&amp;topic=queries',
headers={'Content-Type': 'audio/x-wav', 'Content-Length': len(data)})</p>
      <p>response = req.urlopen(url=request, data=data)
return response.read()
defparse_xml(xml_string):
xmldoc = minidom.parseString(xml_string)
if
xmldoc.getElementsByTagName('recognitionResults')[0].attributes['success'].value
== '1':
return xmldoc.getElementsByTagName('variant')[0].childNodes[0].nodeValue
binary = read_file()
xml = get_text_xml(binary)
best_variant = parse_xml(xml)
print(best_variant)
2. Research and testing of the speech recognition methods and algorithms
This section presents the results of testing speech recognition methods and algorithms using existing systems Sphinx
Speech Recognition, Yandex Speech Kit and Google Speech Recognition.</p>
      <p>Fragments of the testing results for the Sphinx Speech Recognition system are presented in Tables 1-4.It includes only
cases with errors.</p>
    </sec>
    <sec id="sec-2">
      <title>3. Conclusion</title>
    </sec>
    <sec id="sec-3">
      <title>4. Acknowledgments References</title>
      <p>Пылесос
Тостер</p>
      <p>да
вы отсос
то вздыхать
Весы напольные
везде на вольная
пылесос
весы
напольные
0
пылесос
весы напольные
0
Based on the information presented above, we can conclude that the Yandex system is good at recognizing short
expressive phrases, as well as numerals. On contrary, the Google API is good at recognizing long phrases and terms.
The Sphinx system is struggling with recognition of Russian speech. The results of this research are being used in
development of the automatic system that calls the customers of TWIN.</p>
      <p>This work is supported by Act 211 Government of the Russian Federation, contract № 02.A03.21.0006.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cloud</surname>
          </string-name>
          Speech-to-Text https://cloud.google.com/speech-to-text/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>2. YandexSpeechKithttps://tech.yandex.ru/speechkit/</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>CMU</given-names>
            <surname>Sphinx -OPEN SOURCE SPEECH RECOGNITION</surname>
          </string-name>
          <article-title>TOOLKIT https://cmusphinx</article-title>
          .github.io
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. Features of TWINhttps://twin24.ai/#features</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cloud</surname>
          </string-name>
          Speech-to-Text API https://cloud.google.com/speech-to-text/docs/reference/rest/v1/speech/recognize
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>6. Speech Recognition using Google Speech APIhttps://pythonspot.com/speech-recognition-using-googlespeech-api/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. Speech Kit Cloud https://tech.yandex.ru/speechkit/cloud/</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. Speech Kit Mobile SDK https://tech.yandex.ru/speechkit/mobilesdk/</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>