<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic recognition of Slovak-English bilingual speech</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matus Pleva</string-name>
          <email>Matus.Pleva@tuke.sk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuan-Fu Liao</string-name>
          <email>yfliao@ntut.edu.tw</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Hladek</string-name>
          <email>Daniel.Hladek@tuke.sk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Stas</string-name>
          <email>Jan.Stas@tuke.sk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martin Lojka</string-name>
          <email>Martin.Lojka@tuke.sk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jozef Juhar</string-name>
          <email>Jozef.Juhar@tuke.sk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stanislav Ondas</string-name>
          <email>Stanislav.Ondas@tuke.sk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electronic Engineering, College of Electrical Engineering &amp; Computer Science National Taipei University of Technology</institution>
          ,
          <country country="TW">Taiwan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Electronics and Multimedia Communications, Faculty of Electrical Engineering and Informatics, Technical University of Košice</institution>
          ,
          <country country="SK">Slovakia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article describes the progress of a joint project on Multilingual Automatic Speech Recognition using Deep Neural Networks, in which the Technical University works together with National Taipei University of Technology in Taiwan. During the last year, we managed to train multilingual models of combinations of Slovak - English and Slovak - English - Chinese/Mandarin languages. In this paper, we are presenting the results of the Slovak - English model based on deep learning with and without language detection. Furthermore we present new bilingual Slovak-English code-switching database for bilingual systems training and testing. The results indicate that the use of the language detection module can lower the error rate of the multilingual model to the result similar to monolingual models that are generally better for the monolingual tasks.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Thanks to globalization, open culture and easy access to
information on the Internet, users are more exposed to the
multi-language environment than they were in the past. As
a result, foreign words began to appear in spontaneous
spoken language with foreign pronunciation (rather than
being adapted to Slovak pronunciation). This forced
developers to test multiple language models that would be
able to recognize multiple languages at the same time they
often mix.</p>
      <p>
        Multilingual LVCSR (Large Vocabulary Continuous
Speech Recognition) has made great progress in recent
years, notably by introducing Deep learning in Neural
Networks (DNNs), [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. In these works, DNNs were
taught separately to recognize many different languages
or to perform one primary role of speech recognition by
using several helper functions.
      </p>
      <p>
        Shared-hidden layer (SHL) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] architecture where
hidden layers are shared between languages, but the
output layer is language-dependent. Another approach is
used for Multi-component Recurrent Neural Networks
(MRNNs) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] implementation, where bilingual automatic
speech recognition systems with a large LVCSR
dictionary and LIDs (Language IDentification) have been
combined and run in parallel to assist each other. In this work,
an alternative was used when using a Linguistically
Universal / independent End-to-End model (LUE) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our
proposed method uses a language-specific gate
mechanism that allows the internal representation of a network to
be modulated in a language-specific manner. Similar
approach was presented for Cantonese/Turkish/Vietnamese
language specific gate units described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Training bilingual speech models</title>
      <sec id="sec-2-1">
        <title>Monolingual databases and acoustic models</title>
        <p>
          In our previous work we used Julius for English [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and
Slovak [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] local automatic speech recognition tasks. We
have decided to use the Kaldi tool 1 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] (open source
Apache License v2.0) to create an English and Slovak
language recognizer based on Deep Neural Networks.
TaipeiTech Lab already has experience in implementing
and testing Mandarin / English DNN bilingual recognizer
[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>The acoustic corpuses used for the training process of
the acoustic model of the English and Slovak languages is
shown in Table 1. In addition, a multilingual hidden layer
sharing recognizer (SHL) has been used as shown in Fig.
1.</p>
        <p>The acoustic corpuses were divided to 3 parts: Training,
Development and Evaluation. More details are depicted in
Table 2</p>
        <sec id="sec-2-1-1">
          <title>1http://kaldi-asr.org/doc/about.html</title>
          <p>To train a bilingual speech recognizer, databases with
sentences where words in a foreign language are spoken with
a foreign accent are required. It means that commonly
the most words of the sentence is in the speaker’s native
language but he use also different language words as
nonnative speaker. But it is not necessary, because for example
for children raised in bilingual family the native language
is difficult to find out.</p>
          <p>The so called Code Switching occurs when a speaker
alternates two or more languages or language dialects within
a single conversation. This is a well-known meeting
behavior in global technology-oriented companies where
English or German is the official language. During
meetings, the native language is mixed with English words or
phrases used in the company for specific tasks, processes,
equipment, etc.</p>
          <p>
            Other examples of bilingual speech can be found in
air traffic control when part of the communication
between Slovak flight dispatchers with Slovak pilots is
according to international English standards and then a few
polite phrases are heard in Slovak. Another example is
technology-oriented lectures or product and car tests, and
reviews and tutorials on English software. However, we
do not have to stick to a technology sub-group during
lectures, but English phrases often also appear in finance,
social sciences or climate conferences. [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ]
          </p>
          <p>
            From such sources we create the first Slovak English
bilingual database, [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ], which is currently annotated to
increase the accuracy of the language recognition system
(LID) but also to test the resulting bilingual system.
Current state of the database is depicted in Table 3.
2.3
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Textual corpuses and language models</title>
        <p>
          The language models were built to test the monolingual
and bilingual LVCSR engines on the Development (Dev
Set) and Evaluation sets (Test Set). The language model
for this research was provided by the project entitled
"Automatic Subtitling of Audiovisual Content for Hearing
Impaired"2 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. But this model was finally not used in the
tests presented in this paper, but it will be used in further
tests until the end of 2019.
        </p>
        <p>In this work we presented results of the Slovak language
model trained using neural networks and from the train
part of the KEMT-BN corpus (112 thousand utterances)
not used in testing. That is also a reason the results are</p>
        <sec id="sec-2-2-1">
          <title>2http://access.kemt.fei.tuke.sk/</title>
          <p>
            worse then our broadcast news engine baseline [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ], but
the important contribution is the comparison of
monolingual and bilingual systems results.
          </p>
          <p>For English language the same approach was chosen so
132 thousand utterances from LibriSpeech database were
used for English language model training and also for
phonetic vocabulary generating.
2.4</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Bilingual LVCSR</title>
        <p>
          For Slovak/English bilingual speech recognition we built
a bilingual LVCSR (as shown in Fig. 2). For this purpose
we shared the input and hidden layers across languages,
while the output layer is grouped, i.e., [OeLnVCSR; OsLkVCSR].
For better discrimination of similar phonemes all neurons
in the output layer are tuned at the same time (i.e., single
task goal) in this work [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>
          To help discriminate difference languages, an
LIDbased gating mechanism was used to control the bilingual
LVCSR’s outputs [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Finally, the multilingual LVCSR
output scores are multiplied by their corresponding
Language IDentification - LID module scores.
        </p>
        <p>
          For LVCSR training we used 43 MFCC feature vector
together with 100 dimensional i-vector. The TDNN with
850 neurons for each of the 6 hidden layers following the
AiShell/nnet3 recipe [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. The time splicing was chosen
(-2,-1,0,1,2), (-1, 0, 2), (- 3,0,3), (-7,0,2) and (-3,0,3)).
        </p>
        <p>For Language Identification training the same features
and time splicing was used to have even longer-term
language cues. The only difference is in number of neurons,
where we used 425 for 6 hidden layers.</p>
        <p>English senones</p>
        <p>Slovak senones
...</p>
        <p>...</p>
        <p>Layer N
Layer 4
Layer 3
Layer 2</p>
        <p>Layer 1
Input Layer
feature frames
...
...
...
...
...
...
...
...</p>
        <p>...</p>
        <p>En</p>
        <p>
          Sk
In Tables 4. 5. and 6 we can see the results achieved with
the proposed bilingual speech recognition system. The
results are presented in WER (Word Error Rate), which is
the ratio of successfully recognized words to all words in
the test database and is expressed as a percentage [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>It can be stated that thanks to the proposed
multilingual LVCSR and LID running in parallel and integrated
with the gate mechanism, it was possible to achieve results
comparable (see Table 6) with a single-language
recognizer (see Table 4.). These types of engines are generally
better at recognizing single-language data than bilingual
recognizer without the language identifier and gate
mechanism (see Table 5). It should be noted that the bilingual
test database is still under development, so it was not
possible to assess the results in a bilingual test when the words
of both languages and pronunciations are in one sentence.</p>
        <p>
          This article describes the ongoing work of our two-year
project (2018 to 2019). We have now conducted the first
tests of bilingual and monolingual large speech dictionary
recognition (LVSCR) based on Kaldi [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] (TDNN, ResNet
or even DenseNet) and we are also planning to test new
technologies such as DeepSpeech3 [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] or TensorFlow4.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgment</title>
      <p>This work was partly supported by Slovak Research and
Development Agency under contract no. APVV
SK-TW2017-0005, APVV-15-0517, APVV-15-0731, partly
Cultural and educational grant agency from project KEGA
009TUKE-4/2019 and partly Scientific grant agency by
realization of research project VEGA 1/0511/17 both
financed by the Ministry of Education, Science, Research
and Sport of the Slovak Republic and finally by the
Taiwan Ministry of Science and Technology MOST-SRDA
contract No. 108-2911-I-027-501 and
108-2221-E-027067.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>J. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Crosslanguage knowledge transfer using multilingual deep neural network with shared hidden layers</article-title>
          ,
          <source>in Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <source>2013 IEEE International Conference on. , Vancouver</source>
          , pp.
          <fpage>7304</fpage>
          -
          <lpage>7308</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D:</given-names>
          </string-name>
          <article-title>Multi-task recurrent model for true multilingual speech recognition</article-title>
          ,
          <source>in 2016 AsiaPacific Signal and Information Processing Association Annual Summit and Conference (APSIPA)</source>
          , Jeju, Korea,
          <year>Dec 2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Seltzer</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          :
          <article-title>Towards language-universal end-to-end speech recognition</article-title>
          ,
          <source>2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , Calgary. IEEE, pp.
          <fpage>4914</fpage>
          -
          <lpage>4918</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <surname>P.</surname>
          </string-name>
          <article-title>Multilingual speech recognition training and adaptation with language-specific gate units (</article-title>
          <year>2019</year>
          )
          <article-title>2018 11th International Symposium on Chinese Spoken Language Processing</article-title>
          , ISCSLP 2018 - Proceedings, art. no.
          <issue>8706584</issue>
          , pp.
          <fpage>86</fpage>
          -
          <lpage>90</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Pleva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cizmar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hudson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carruth</surname>
            ,
            <given-names>D. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bethel</surname>
            ,
            <given-names>C. L.</given-names>
          </string-name>
          :
          <article-title>Implementing English speech interface to Jaguar robot for SWAT training</article-title>
          ,
          <source>In proceedings: Applied Machine Intelligence and Informatics (SAMI)</source>
          ,
          <source>2017 IEEE 15th International Symposium on, Herlany</source>
          , Slovakia, IEEE, pp.
          <fpage>105</fpage>
          -
          <lpage>110</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Ondas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pleva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lojka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiktova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sulir</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cizmar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holcer</surname>
          </string-name>
          , R.:
          <article-title>“Speech technologies for advanced applications in service robotics</article-title>
          ,
          <source>” Acta Polytechnica Hungarica</source>
          ,
          <volume>10</volume>
          (
          <issue>5</issue>
          ), pp.
          <fpage>45</fpage>
          -
          <lpage>61</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Povey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghoshal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boulianne</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burget</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glembek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hannemann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motlicek</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qian</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwarz</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silovsky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stemmer</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vesely</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>The Kaldi speech recognition toolkit</article-title>
          .,
          <source>In Proceedings of IEEE 2011 Workshop on Automatic Speech Recognition and Understanding - ASRU</source>
          <year>2011</year>
          ,
          <article-title>Hilton Waikoloa Village, Big Island, Hawaii</article-title>
          .
          <source>IEEE Signal Processing Society</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C. T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>Y. F.</given-names>
          </string-name>
          :
          <article-title>A preliminary study on cross-language knowledge transfer for low-resource Taiwanese Mandarin ASR,</article-title>
          <year>2016</year>
          <article-title>Conference of The Oriental Chapter of International Committee for Coordination and Standardization of Speech Databases and Assessment Techniques (O-COCOSDA), Bali</article-title>
          , IEEE, pp.
          <fpage>33</fpage>
          -
          <lpage>38</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Panayotov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Povey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khudanpur</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Librispeech: an ASR corpus based on public domain audio books</article-title>
          .
          <source>In Proceedings of Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <source>2015 IEEE International Conference on, South Brisbane</source>
          , pp.
          <fpage>5206</fpage>
          -
          <lpage>5210</lpage>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Pleva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
          </string-name>
          , J.:
          <string-name>
            <surname>TUKE-BNews-SK</surname>
          </string-name>
          :
          <article-title>Slovak Broadcast News Corpus Construction and Evaluation</article-title>
          ,
          <source>In: LREC 2014 : Ninth International Conference on Language Resources and Evaluation : May 26-31</source>
          ,
          <year>2014</year>
          , Reykjavik, Iceland. - Paris : ELRA,
          <year>2014</year>
          , pp.
          <fpage>1709</fpage>
          -
          <lpage>1713</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hudak</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Methods for bilingual automatic speech recognition, 2019 Bachelor thesis in Slovak, KEMT, FEI</article-title>
          ,
          <string-name>
            <surname>TUKE</surname>
          </string-name>
          , Kosice,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Stas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viszlay</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lojka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koctur</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hladek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Automatic Transcription and Subtitling of Slovak Multi-genre Audiovisual Recordings</article-title>
          .
          <source>In In Human Language Technology. Challenges for Computer Science and Linguistics: 7th Language and Technology Conference, LTC</source>
          <year>2015</year>
          , Poznan´,
          <source>Poland, November 27-29</source>
          ,
          <year>2015</year>
          ,
          <string-name>
            <given-names>Revised</given-names>
            <surname>Selected Papers</surname>
          </string-name>
          (Vol.
          <volume>10930</volume>
          , p.
          <fpage>42</fpage>
          ). Springer. pp.
          <fpage>42</fpage>
          -
          <lpage>56</lpage>
          ).
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Lojka</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viszlay</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hladek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
          </string-name>
          , J.:
          <article-title>Slovak Broadcast News Speech Recognition and Transcription System</article-title>
          .
          <source>International Conference on Network-Based Information Systems</source>
          . In: Barolli L.,
          <string-name>
            <surname>Kryvinska</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Enokido</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takizawa</surname>
            <given-names>M</given-names>
          </string-name>
          .
          <article-title>(eds) Advances in Network-Based Information Systems</article-title>
          .
          <source>NBiS 2018. Lecture Notes on Data Engineering and Communications Technologies - LNDECT</source>
          , vol
          <volume>22</volume>
          . Springer, Cham, pp.
          <fpage>385</fpage>
          -
          <lpage>394</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>Y. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pleva</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hladek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viszlay</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lojka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
          </string-name>
          , J.:
          <source>”Gated Module Neural Network for Multilingual Speech Recognition,“ 2018 11th International Symposium on Chinese Spoken Language Processing (ISCSLP)</source>
          ,
          <source>Taipei City, Taiwan</source>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>135</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Bu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Na</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
          </string-name>
          , H.:
          <article-title>AIShell-1: An open-source Mandarin speech corpus and a speech recognition baseline</article-title>
          .
          <source>In 2017 20th Conference of the Oriental Chapter of the International Coordinating Committee on Speech Databases and Speech I/O Systems</source>
          and
          <string-name>
            <surname>Assessment (O-COCOSDA</surname>
            <given-names>)</given-names>
          </string-name>
          , Seoul, pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . IEEE.
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Mackova</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cizmar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A study of acoustic features for emotional speaker recognition in i-vector representation</article-title>
          .
          <source>Acta Electrotechnica et Informatica</source>
          ,
          <volume>15</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>15</fpage>
          -
          <lpage>20</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Amodei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananthanarayanan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anubhai</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Battenberg</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Case</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casper</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catanzaro</surname>
            , B., Cheng,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          et al.:
          <article-title>Deep speech 2: End-to-end speech recognition in English and Mandarin</article-title>
          .
          <source>In: International Conference on Machine Learning</source>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          ,
          <year>2016</year>
          , June.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>