<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dialectal Speech Recognition and Translation of Swiss German Speech to Standard German Text: Microsoft's Submission to SwissText 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuriy Arabskyy</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aashish Agarwal</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Subhadeep Dey</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oscar Koller Microsoft - Munich</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germany</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>yuarabsk</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>t-aagarwal</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>subde</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>oskoller} @microsoft.com</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper describes the winning approach in the Shared Task 3 at SwissText 2021 on Swiss German Speech to Standard German Text, a public competition on dialect recognition and translation. Swiss German refers to the multitude of Alemannic dialects spoken in the German-speaking parts of Switzerland. Swiss German differs significantly from standard German in pronunciation, word inventory and grammar. It is mostly incomprehensible to native German speakers. Moreover, it lacks a standardized written script. To solve the challenging task, we propose a hybrid automatic speech recognition system with a lexicon that incorporates translations, a 1st pass language model that deals with Swiss German particularities, a transfer-learned acoustic model and a strong neural language model for 2nd pass rescoring. Our submission reaches 46.04% BLEU on a blind conversational test set and outperforms the second best competitor by a 12% relative margin.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        While general speech recognition has matured to a
point where it surpasses human performance on
specific datasets
        <xref ref-type="bibr" rid="ref22">(Xiong et al., 2016)</xref>
        , dialectal
recognition as in the case of Swiss German
        <xref ref-type="bibr" rid="ref13 ref16">(Nigmatulina et al., 2020)</xref>
        or Arabic dialects
        <xref ref-type="bibr" rid="ref12 ref12 ref2 ref3">(Ali et al.,
2021; Hussein et al., 2021; Ali, 2018)</xref>
        still
represents a major challenge. Swiss German refers to
the multitude of Alemannic dialects spoken in the
German-speaking parts of Switzerland. It hence
represents dialects that differ significantly from
standard or high German in pronunciation, word
inventory and grammar. Moreover, it lacks a
standardized writing system. High German is
commonly used for the large majority of written
communication between people and in the media of
German-speaking Switzerland, while in rather
informal chats and short messaging Swiss people
may use transliterated non-standardized Swiss
dialect. The process of transcribing Swiss German
into high German text therefore requires speech
recognition with an inherent translation step.
Moreover, the task can be considered low-resource as
available data remains extremely scarce.
      </p>
      <p>
        In previous studies, Garner et al. (2014) tackled
this challenge by training hybrid models
(HMMGMM, HMM-DNN, and KL-HMM) to transcribe
Walliserdeutsch, a Swiss-German dialect spoken
in the south-western alpine canton of
Switzerland and further used a phrase-based machine
translation model to translate it to standard
German. Following this, other researchers explored
techniques to add the translation step in the
lexicon by directly mapping Swiss-German
pronunciation to standard German. Stadtschnitzer and
Schmidt (2018) estimated Swiss German
pronunciations from a standard German speech
recognition model using a data-driven technique, and
trained stronger TDNN-LSTM based acoustic
models. Kew et al. (2020) and Nigmatulina et al. (2020)
trained transformer-based G2P models from
standard German to Swiss pronunciations and trained a
Kaldi-based TDNN+ivector system using the WSJ
recipe1. Yet a third approach is to directly
apply end-to-end deep learning models. Büchi et al.
(2020) and Agarwal and Zesch (2020) at
SwissText 2020 used the Jasper architecture
        <xref ref-type="bibr" rid="ref15">(Li et al.,
2019)</xref>
        and Mozilla DeepSpeech
        <xref ref-type="bibr" rid="ref10">(Hannun et al.,
1https://github.com/kaldi-asr/kaldi/
tree/master/egs/wsj
2014)</xref>
        , respectively. In both cases the system was
first trained on high German data and then
transferlearned to Swiss German.
      </p>
      <p>In this paper, we describe our proposal to solve
the challenging task of transcribing Swiss German
speech to standard German text. It won the
competition at SwissText 2021 with a large margin to
other competing systems.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System Overview</title>
      <p>
        In this section, we propose our changes to a
conventional hybrid
        <xref ref-type="bibr" rid="ref5">(Bourlard and Dupont, 1996)</xref>
        automatic speech recognition (ASR) system, which
relies on lexicon and alignments for good
performance. We present details in order to enable it for
dialect speech recognition and translation.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Data</title>
        <p>To train our proposed model, we utilized a
selection of publicly available and internal datasets. Our
starting point was the Swiss Parliament Corpus V2
dataset (Plüss et al., 2020) shared as part of the
SwissText 2021 competition. It covers 293 hours
and contains recordings from the local parliament
of the Kanton Bern. Its transcripts are in standard
German while the audio covers Swiss German
(predominantly in the Bernese dialect). The dataset
has been preprocessed by the publishers with the
purpose of cleaning its annotations and ensuring a
good match between audio content and
transcription. It is provided with a choice of different
preprocessing flavors. We used the train_all split. In
addition, we used a 493-hour internal dataset
representing a media domain encompassing
conversational speech from interviews, discussions,
podcasts and others. A subset (around 50 hours) of the
data is annotated with both Swiss transliterations as
well as standard German. The remaining data has
only been annotated with standard German.
Additionally, we used an internal high German dataset
encompassing around 10k hours to pre-train our
model.</p>
        <p>In terms of test data, the SwissText 2021
competition was accompanied by a 13 hours
conversational test set covering Swiss German speakers
from all German-speaking parts of Switzerland.
The encountered dialectal distribution is claimed
to closely match the real distribution in
Switzerland. The set was not disclosed to the participants.
Hence, for the analysis in this paper, we report our
numbers on a publicly available test set part of the
dataset from the Bernese parliament (Plüss et al.,
2020). It comprises 6 hours of dialectal speech.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Lexicon</title>
        <p>We propose to incorporate the translation from
Swiss German to standard German as part of the
lexicon. However, this leads to a complex and often
ambiguous mapping from phoneme to grapheme
sequences, which is very different from languages
with a direct relation between writing scheme
and pronunciation (e.g. English or standard
German). Subsequently, statistical models that map
graphemes to phonemes (G2P) trained on Swiss
German data incorporating such translations yield
much noisier output with significantly higher phone
error rates as compared to G2Ps for standard
languages. To mitigate this problem, we construct the
lexicon in several stages.</p>
        <p>
          In a first step, we make use of parallel corpora
encompassing Swiss and standard German
annotations to extract word mappings between Swiss and
standard German. Sophisticated filtering methods
help to ensure a high quality of these mappings. We
opt for frequency filtering and filtering based on
vicinity in a word embedding space
          <xref ref-type="bibr" rid="ref4">(Bojanowski
et al., 2017)</xref>
          of Swiss German words taking the
most frequent mapping as center point.
        </p>
        <p>
          In a second step, a standard German G2P model
is applied to convert Swiss German transliterations
into corresponding phone sequences. This results
in a dictionary that maps standard German words to
Swiss pronunciations. Jointly with existing Swiss
German lexicon resources
          <xref ref-type="bibr" rid="ref19">(Schmidt et al., 2020)</xref>
          ,
the previously generated mappings are then used
to train a dedicated Swiss German G2P model.
        </p>
        <p>We evaluate the quality of the resulting G2P
model on a manually labeled test set. Those cover
mappings from standard German words to Swiss
German phone sequences and encompass a variety
of relevant categories such as diminuation,
shortening or translation. Refer to Table 1 for samples of
the assessed categories.</p>
        <p>
          The Swiss G2P model allows to find suitable
pronunciations for the relevant word inventory present
in the acoustic and language model training
corpora. However, to further increase the quality of
the given pronunciations, data-driven lexicon
learning techniques
          <xref ref-type="bibr" rid="ref23">(Zhang et al., 2017)</xref>
          are applied.
Those help to identify and correct noisy lexicon
entries.
2nd person plural 2nd person sing
        </p>
        <p>fragt
f hr a_ g ax t</p>
        <p>riecht
sh m oe k c ax t</p>
        <p>fragst
f hr a_ k sh</p>
        <p>riechst
sh m oe k sh</p>
        <p>Diminuation
erdmännchen
e_r t m eh n l i_
gläschen
g l e_ s l i_</p>
        <p>Shortening
gymnasium</p>
        <p>g ih m i_
schwimmbad
b a_ d ih</p>
        <p>Translation</p>
        <p>
          Variability
kopf
g hr ih n t
kneipe
b ai ts
kannst
k a sh
zweites
ts v ai t
Incorporating the translation from Swiss German to
standard German as part of the lexicon introduces
significant ambiguity in the decoding process. To
counteract, we suggest using a strong standard
German language model (LM) which helps to produce
accurate hypotheses. We employ a first pass
countbased LM to output up to 100 sentence hypotheses
and a second pass neural LSTM (long short-term
memory) LM
          <xref ref-type="bibr" rid="ref21">(Sundermeyer et al., 2012)</xref>
          for
rescoring
          <xref ref-type="bibr" rid="ref8">(Deoras et al., 2011)</xref>
          . The first pass model is a
5-gram LM trained on large amounts of standard
German text corpora totalling to over 100 billion
words. We apply Kneser-Ney smoothing
          <xref ref-type="bibr" rid="ref14">(Kneser
and Ney, 1995)</xref>
          .
        </p>
        <p>Furthermore, we make some adjustments to
better deal with Swiss German particularities, as
described in the following paragraphs.</p>
        <p>Compounds: German is a compounding
language and tends to compose words (particularly
nouns) of several smaller subwords. The
resulting chains of word stems can lead to an infinitely
large vocabulary size with words that occur very
infrequently throughout the corpus. This spreading
of probability mass weakens the LM. We hence
decompound all compounded words in the training
corpus and split them into subwords.</p>
        <p>
          Clitics: Swiss German tends to merge words
beyond compounding, not preserving word
stems
          <xref ref-type="bibr" rid="ref11 ref9">(Hollenstein and Aepli, 2014)</xref>
          . For instance,
the Swiss German ‘hemmer’ is the translation of
‘haben wir’ in standard German (English: ‘have
we’). We identified approximately 8000 clitics in
our corpus. We incorporate them in the decoding
process by updating lexicon and LM. Following
the example above the translated clitic ‘haben#wir‘
with the corresponding Swiss pronunciation is
added to the lexicon. As for the LM, we merge
occurrences of relevant word pairs and interpolate
with the unmerged LM.
2.4
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Acoustic model</title>
        <p>
          The acoustic model is trained with 80 dimensional
log-mel filterbank features, with a processing
window of 25ms and 10ms frame shift. The feature
vector from the previous frame is concatenated with
the current frame to obtain a 160 dimensional
vector. We used a LC-BLSTM (latency controlled
bidirectional long short-term memory) based acoustic
model, that is popularly applied in speech
recognition for controlling decoding latency to a few
frames
          <xref ref-type="bibr" rid="ref7">(Chen and Huo, 2016)</xref>
          . The model was
trained with alignments from a feed-forward
network with context-dependent tied states (senones).
The model has 9k senone units. The LC-BLSTM
is trained with 6 hidden layers with 512 units each.
The hidden vectors from the forward and backward
propagation were concatenated and then projected
to a 512 dimensional vector. The model is trained
with a cross entropy loss function. The decoding
lexicon is extended with Swiss German words for
training. The transliterations are used during forced
alignment whenever possible. This helps to reduce
the pronunciation ambiguity in the alignment phase
and is especially helpful in the early training phases
when no strong model is available for alignment.
        </p>
        <p>
          The results are reported in terms of BLEU
          <xref ref-type="bibr" rid="ref17">(Papineni et al., 2002)</xref>
          and word error rate (WER) on
the Swiss Parliament test set described in Section
2.1.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussions</title>
      <p>An ablation study of the proposed approaches is
presented in Table 2. All of the performance gains
in this section will be reported as relative
percentage improvements, while aforementioned table
contains absolute numbers.</p>
      <p>We first evaluate the effect of transfer learning
on the results with the Swiss Parliament training
set. It can be observed that it significantly helps to
improve both WER and BLEU. In particular, the
transfer-learned model (row 2, Table 2) improves
over the model trained from scratch (row 1, Table 2)
by 8:4% WER and 11:6% BLEU. The result shows
1
2
3
4</p>
      <p>Swiss Parliament
+ Transfer Learning
+ Internal Data
+ 2nd Pass Rescoring</p>
      <p>Further adding additional internal training data
shows additional gains in performance. As such,
we observe that the WER improves by 2:3% and
BLEU by 2:5%.</p>
      <p>Finally, 2nd Pass rescoring is applied as
described in Section 2.3 to reorder the top 100
hypotheses. It can be observed from row 4, Table 2
that rescoring helps to improve the performance by
5:8% WER and 8:9% BLEU.</p>
      <p>Our submission to SwissText 2021 achieves
46:04% BLEU on the official SwissText blind test
set. This leads to a 12% relative margin in BLEU
with respect to the second best competitor which
was 40:99%.</p>
      <p>The acoustic models have been trained using 8
GPUs for 25 epochs. This results in a total training
time of around 400 GPU-hours when training on
Swiss Parliament only and about 1200 GPU-hours
when adding the internal data.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>In this paper, we described a speech recognition
system that achieves strong results on the task of
recognizing Swiss German dialect and translating
it into standard German text. We proposed a
hybrid ASR system with a lexicon that incorporates
translations, a 1st pass language model that deals
with Swiss German word compounding and
clitics, an acoustic model that is transfer-learned from
standard German resources and a strong neural
language model for 2nd pass rescoring to smoothen
translation artifacts. Furthermore, we provided
an ablation study that allows to infer the effect of
adding training data, performing transfer learning
and 2nd pass rescoring. Our submission reached
46.04% BLEU on a challenging conversational test
set and outperformed all competing approaches by
a large margin.</p>
      <p>In terms of future work, we would like to
investigate word re-orderings as part of the
translation, which our current model does not actively
support. For instance, Swiss German frequently
moves verbs in relative clauses to different
positions with respect to the standard German word
order. Furthermore, sequence discriminative
training is a promising route for exploration as well as
using unsupervised data for acoustic model
training.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Aashish</given-names>
            <surname>Agarwal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Torsten</given-names>
            <surname>Zesch</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Ltl-ude at low-resource speech-to-text shared task: Investigating mozilla deepspeech in a low-resource setting</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Ahmed</given-names>
            <surname>Ali</surname>
          </string-name>
          , Shammur Chowdhury, Mohamed Afify, Wassim El-Hajj, Hazem Hajj, Mourad Abbas, Amir Hussein, Nada Ghneim, Mohammad Abushariah, and
          <string-name>
            <given-names>Assal</given-names>
            <surname>Alqudah</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Connecting Arabs: Bridging the gap in dialectal speech recognition</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>64</volume>
          (
          <issue>4</issue>
          ):
          <fpage>124</fpage>
          -
          <lpage>129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Ahmed</given-names>
            <surname>Mohamed Abdel Maksoud Ali</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>MultiDialect Arabic Broadcast Speech Recognition</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of Edinburgh, Edinburgh, UK.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>5</volume>
          :
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Hervé</given-names>
            <surname>Bourlard</surname>
          </string-name>
          and
          <string-name>
            <given-names>Stéphane</given-names>
            <surname>Dupont</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>A new ASR approach based on independent processing and recombination of partial frequency bands</article-title>
          .
          <source>In Proc. Int. Conf. on Spoken Language Processing (ICSLP)</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>426</fpage>
          -
          <lpage>429</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Büchi</surname>
          </string-name>
          , Malgorzata Anna Ulasik, Manuela Hürlimann, Fernando Benites, Pius von Däniken, and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Cieliebak</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Zhaw-init at germeval 2020 task 4: Low-resource speech-to-text.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Kai</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Qiang</given-names>
            <surname>Huo</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Training deep bidirectional lstm acoustic model for lvcsr by a contextsensitive-chunk bptt approach</article-title>
          .
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          ,
          <volume>24</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1185</fpage>
          -
          <lpage>1193</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Anoop</given-names>
            <surname>Deoras</surname>
          </string-name>
          , Tomáš Mikolov, and
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Church</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A Fast Re-scoring Strategy to Capture LongDistance Dependencies</article-title>
          .
          <source>In Proceedings of the 2011 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>1116</fpage>
          -
          <lpage>1127</lpage>
          , Edinburgh, Scotland, UK. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Philip N.</given-names>
            <surname>Garner</surname>
          </string-name>
          , David Imseng, and Thomas Meyer.
          <year>2014</year>
          .
          <article-title>Automatic Speech Recognition and Translation of a Swiss German Dialect: Walliserdeutsch</article-title>
          .
          <source>In Proc. of the Ann. Conf. of the Int. Speech Commun. Assoc</source>
          . (Interspeech),
          <source>Singapore.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Awni</given-names>
            <surname>Hannun</surname>
          </string-name>
          ,
          <source>Carl Case, Jared Casper</source>
          , Bryan Catanzaro, Greg Diamos, Erich Elsen, Ryan Prenger, Sanjeev Satheesh, Shubho Sengupta,
          <string-name>
            <given-names>Adam</given-names>
            <surname>Coates</surname>
          </string-name>
          , et al.
          <year>2014</year>
          .
          <article-title>Deep speech: Scaling up end-to-end speech recognition</article-title>
          .
          <source>arXiv preprint arXiv:1412</source>
          .
          <fpage>5567</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Nora</given-names>
            <surname>Hollenstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Noëmi</given-names>
            <surname>Aepli</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Compilation of a Swiss German dialect corpus and its application to PoS tagging</article-title>
          .
          <source>In Proceedings of the First Workshop on Applying NLP Tools to Similar Languages, Varieties and Dialects</source>
          , pages
          <fpage>85</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Amir</given-names>
            <surname>Hussein</surname>
          </string-name>
          , Shinji Watanabe, and
          <string-name>
            <given-names>Ahmed</given-names>
            <surname>Ali</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Arabic Speech Recognition by End-to-</article-title>
          <string-name>
            <surname>End</surname>
          </string-name>
          ,
          <source>Modular Systems and Human. arXiv:2101</source>
          .08454 [cs, eess].
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Tannon</given-names>
            <surname>Kew</surname>
          </string-name>
          , Iuliia Nigmatulina, Lorenz Nagele, and
          <string-name>
            <given-names>Tanja</given-names>
            <surname>Samardzic</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Uzh tilt: A kaldi recipe for swiss german speech to standard german text</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Reinhard</given-names>
            <surname>Kneser</surname>
          </string-name>
          and Hermann Ney.
          <year>1995</year>
          .
          <article-title>Improved backing-off for m-gram language modeling</article-title>
          .
          <source>In Proc. IEEE Int. Conf. on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>181</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Jason</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Vitaly</given-names>
            <surname>Lavrukhin</surname>
          </string-name>
          , Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen,
          <string-name>
            <given-names>Huyen</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , and Ravi Teja Gadde.
          <year>2019</year>
          .
          <article-title>Jasper: An end-to-end convolutional neural acoustic model</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .03288.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Iuliia</given-names>
            <surname>Nigmatulina</surname>
          </string-name>
          , Tannon Kew, and
          <string-name>
            <given-names>Tanja</given-names>
            <surname>Samardzic</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>ASR for non-standardised languages with dialectal variation: the case of Swiss German</article-title>
          .
          <source>In Proceedings of the 7th Workshop on NLP for Similar Languages, Varieties and Dialects</source>
          , pages
          <fpage>15</fpage>
          -
          <lpage>24</lpage>
          , Barcelona, Spain (Online).
          <source>International Committee on Computational Linguistics (ICCL).</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Kishore</given-names>
            <surname>Papineni</surname>
          </string-name>
          , Salim Roukos, Todd Ward, and
          <string-name>
            <given-names>WeiJing</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Michel</given-names>
            <surname>Plüss</surname>
          </string-name>
          , Lukas Neukom, and
          <string-name>
            <given-names>Manfred</given-names>
            <surname>Vogel</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Swiss parliaments corpus, an automatically aligned swiss german speech to standard german text corpus</article-title>
          . arXiv preprint arXiv:
          <year>2010</year>
          .02810.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Larissa</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , Lucy Linder, Sandra Djambazovska, Alexandros Lazaridis, Tanja Samardžic´, and
          <string-name>
            <given-names>Claudiu</given-names>
            <surname>Musat</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>A Swiss German Dictionary: Variation in Speech and Writing</article-title>
          . arXiv:
          <year>2004</year>
          .00139 [cs].
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Michael</given-names>
            <surname>Stadtschnitzer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Schmidt</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Adaptation and training of a swiss german speech recognition system using data-driven pronunciation modelling</article-title>
          .
          <source>Proceedings of DAGA-44. Jahrestagung für Akustik</source>
          , München, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Sundermeyer</surname>
          </string-name>
          , Ralf Schlüter, and Hermann Ney.
          <year>2012</year>
          .
          <article-title>LSTM neural networks for language modeling</article-title>
          .
          <source>In Proc. of the Ann. Conf. of the Int. Speech Commun. Assoc. (Interspeech)</source>
          , pages
          <fpage>194</fpage>
          -
          <lpage>197</lpage>
          , Portland,
          <string-name>
            <surname>OR</surname>
          </string-name>
          , USA.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Wayne</given-names>
            <surname>Xiong</surname>
          </string-name>
          , Jasha Droppo, Xuedong Huang, Frank Seide, Mike Seltzer, Andreas Stolcke,
          <string-name>
            <given-names>Dong</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Geoffrey</given-names>
            <surname>Zweig</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Achieving human parity in conversational speech recognition</article-title>
          .
          <source>arXiv preprint arXiv:1610</source>
          .
          <fpage>05256</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Xiaohui</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Vimal Manohar, Daniel Povey, and
          <string-name>
            <given-names>Sanjeev</given-names>
            <surname>Khudanpur</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Acoustic Data-Driven Lexicon Learning Based on a Greedy Pronunciation Selection Framework</article-title>
          .
          <source>In Proc. of the Ann. Conf. of the Int. Speech Commun. Assoc. (Interspeech)</source>
          , pages
          <fpage>2541</fpage>
          -
          <lpage>2545</lpage>
          . ISCA.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>