<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>JANUS: a
speech-to-speech translation system using connec-
tionist and symbolic processing strategies. In [Pro-
ceedings] ICASSP</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>ZHAW-CAI: Ensemble Method for Swiss German Speech to Standard German Text</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Malgorzata Anna Ulasik</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuela H u¨rlimann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bogumila Dubel</string-name>
          <email>bodubel@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yves Kaufmann</string-name>
          <email>y.kaufmann@yagan.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Silas Rudolf</string-name>
          <email>silasrudolf@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan Deriu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katsiaryna Mlynchyk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hans-Peter Hutter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Cieliebak</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Artificial Intelligence Zurich University of Applied Sciences</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1991</year>
      </pub-date>
      <volume>91</volume>
      <issue>1991</issue>
      <fpage>793</fpage>
      <lpage>796</lpage>
      <abstract>
        <p>This paper presents the contribution of ZHAW-CAI to the Shared Task ”Swiss German Speech to Standard German Text” at the SwissText 2021 conference. Our approach combines three models based on the Fairseq, Jasper and Wav2vec architectures trained on multilingual, German and Swiss German data. We applied an ensembling algorithm on the predictions of the three models in order to retrieve the most reliable candidate out of the provided translations for each spoken utterance. With the ensembling output, we achieved a BLEU score of 39.39 on the private test set, which gave us the third place out of four contributors in the competition.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Speech-to-Text (STT) enables transcribing spoken
utterances into text. For successfully performing
a transformation from speech to a text, the
existence of a standardised writing system of the target
language is of prime importance. This is where
Swiss German 1 poses a substantial challenge: it
does not have a standardised orthography since it
functions as the default spoken language in both
formal and informal situations, while for writing,
the Standard German language is used. This
phenomenon, called “medial diglossia”
        <xref ref-type="bibr" rid="ref24">(Siebenhaar
and Wyler, 1997)</xref>
        , occurs in the entire
Germanspeaking part of Switzerland, which is additionally
characterised by a high dialect diversity. Swiss
German is increasingly used for writing in informal
Copyright © 2021 for this paper by its authors. Use permitted
under Creative Commons License Attribution 4.0
International (CC BY 4.0)
      </p>
      <p>
        1To be precise, there is no single ”Swiss German”
language, but rather a collection of many different regional
dialects that are subsumed with this term.
contexts, but since there is no single standard
writing system, Swiss German speakers usually write
phonetically in their local dialect in informal
situations
        <xref ref-type="bibr" rid="ref23">(Siebenhaar, 2003)</xref>
        . On formal occasions
such as work meetings and political debate, speech
is typically transcribed into Standard German. As
there is a considerable linguistic distance between
Swiss German dialects and Standard German,
developing a model for transcribing Swiss German
speech into Standard German text actually involves
Speech Translation, which combines STT with
Machine Translation (MT)
        <xref ref-type="bibr" rid="ref2">(Be´rard et al., 2016)</xref>
        .
      </p>
      <p>
        As a response to the Shared Task “Swiss
German Speech to Standard German Text” organised
at Swisstext 2021, we provided a solution
consisting of three models based on different architectures:
Fairseq
        <xref ref-type="bibr" rid="ref4">(Wang et al., 2020a)</xref>
        , Jasper
        <xref ref-type="bibr" rid="ref13">(Li et al., 2019)</xref>
        and Wav2vec XLSR-5
        <xref ref-type="bibr" rid="ref1">(Baevski et al., 2020)</xref>
        which
were trained with various data sets, both in
Standard German and Swiss German. Their predictions
were subsequently fed into a majority voting
algorithm with the aim to select the most reliable
translation.
      </p>
      <p>The remainder of this paper is structured as
follows: Section 2 provides the description of the
Shared Task and Section 3 discusses relevant
literature. In Section 4 we present the systems which
make up our final solution, their architecture and
the applied training data. In section 5 we provide an
overview of all experiments performed with these
models and their outputs. Section 6 lays out the
ensembling approach and section 7 presents the
post-processing experiments we performed on the
predictions of the models. The paper ends with a
conclusion presented in section 8.
taining 293 hours of audio recordings, mostly in
the Bernese dialect, transcribed in Standard
German. Since the alignment between the recordings
and the transcripts was done automatically, each
utterance has an Intersection over Union (IoU) score
reflecting its alignment quality. Additionally, there
was an unlabelled data set consisting of 1208 hours
of recordings, mostly in the Zurich dialect. The
solutions were evaluated based on a 13 hours test set,
which contains recordings of speakers coming from
all German-speaking parts of Switzerland. The
dialect distribution of the test set is close to the actual
Swiss German dialect distribution in Switzerland.</p>
      <p>
        The translation accuracy of the provided
solutions is measured using BLEU, a standard metric
for automatic evaluation of machine translation
        <xref ref-type="bibr" rid="ref17">(Papineni et al., 2002)</xref>
        . The approach consists in
counting n-grams in the candidate translation matching
n-grams in the reference translation without taking
the word order into account. The metric ranges
from 0 to 100. A perfect match results in a score
of 100. A score of 0 occurs if there are no matches.
The tool used by the organisers for evaluating
solutions is the NLTK implementation of the BLEU
score with default parameters2. Prior to evaluation,
both the references and the translations are
normalised: the utterances are lowercased, the
punctuation is removed, the numbers are spelled out and
all non-ASCII characters except for the letters ”a¨”,
”o¨”, ”u¨” are removed.
      </p>
      <p>The test set was split into a public and a private
subset of equal sizes. For all evaluations presented
in this paper, the public test set was used.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Speech Translation (ST) is the task of translating
spoken text in a source language to text or speech
in a target language. The approaches to solve this
problem can be put into two categories: cascading
approaches and end-to-end approaches
        <xref ref-type="bibr" rid="ref1 ref12 ref21 ref25">(Sperber
and Paulik, 2020)</xref>
        .
      </p>
      <p>
        Cascaded Approaches work by splitting the
task into two steps: first, an STT model transcribes
speech of the source language to text in the
target language, and then a machine translation (MT)
module translates the generated text into the target
language (Waibel et al., 1991). The main issue
with the cascaded approach is the fact that errors
2https://www.nltk.org/api/nltk.
translate.html#nltk.translate.bleu_score.
corpus_bleu
made by the STT module are propagated to the
MT module
        <xref ref-type="bibr" rid="ref15">(Ney, 1999)</xref>
        . Thus, efforts are put into
coupling the STT and MT modules to prevent error
propagation, for instance, by generating multiple
hypotheses of the STT system via n-best search
or the creation of lattices
        <xref ref-type="bibr" rid="ref22 ref26">(Woszczyna et al., 1993;
Schultz et al., 2004)</xref>
        .
      </p>
      <p>
        End-to-End Approaches model ST as a single
task, where input is speech in the source language,
and the output consists of text or speech in the
target language. The main issue with this
modelling approach is the lack of sufficient training
data. Whereas data for STT typically consists of
several hundreds of hours of transcribed data, most
ST datasets contain only a fraction of this amount.
For instance, the Europarl-ST corpus contains on
average only 42 hours of transcribed data per
language pair
        <xref ref-type="bibr" rid="ref11">(Iranzo-Sa´nchez et al., 2020)</xref>
        , whereas
the Librispeech STT corpus contains around 1000
hours of transcribed data
        <xref ref-type="bibr" rid="ref16">(Panayotov et al., 2015)</xref>
        .
For this reason, end-to-end approaches nowadays
rely on leveraging multi-task learning and single
language pre-training of the STT and MT
submodules and use the ST dataset for fine-tuning
        <xref ref-type="bibr" rid="ref4">(Wang
et al., 2020b)</xref>
        .
      </p>
      <p>Most cascading approaches rely on data where
access to both the source language transcript and
its target language translation is needed. However,
in our scenario, we do not have access to written
text of the source language since Swiss German
is a spoken language, and thus, often directly
transcribed into Standard German (see 1 for more
details). Thus, our models follow the End-to-End
approach.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Systems Description</title>
      <p>This section describes the architecture of the three
models which build the foundation for the
experiments presented in Section 5 and are components of
the final solution which combines the three models’
outputs in an ensembling algorithm. The section
also explains what data was used for training the
models.
4.1
4.1.1</p>
      <sec id="sec-3-1">
        <title>Fairseq</title>
      </sec>
      <sec id="sec-3-2">
        <title>Model</title>
        <p>
          Fairseq is based on the transformer architecture for
Speech-to-Text provided by Fairseq S2T Toolkit
          <xref ref-type="bibr" rid="ref4">(Wang et al., 2020a)</xref>
          , which combines the tasks
of STT and ST under the same encoder-decoder
architecture
          <xref ref-type="bibr" rid="ref4">(Changhan Wang, 2020)</xref>
          . The
experiments were trained with the small transformer
model with 256 dimensions, 12 Layers encoder, 6
Layers decoder, 27M parameters, Adam optimiser,
and inverse square root for the learning rate
scheduler. Decoding is executed with a character-based
SentencePiece model (Taku Kudo, 2018) using an
n-best decoding strategy with n=5. The acoustic
model (encoder) can be pre-trained with the same
transformer architecture as described above.
4.1.2 Data
The audios were extracted to 80-dimensional log
mel-scale filterbank features (windows with 25 ms
size and 10 ms shift) and saved in NumPy format
for the training. To alleviate overfitting, speech
data transforms SpecAugment
          <xref ref-type="bibr" rid="ref18">(Park et al., 2019)</xref>
          ,
adopted by Fairseq S2T, were applied. For text
normalisation we used the script provided by the
task organisers. Additional numbers were spelled
out using num2words3. We use three additional
datasets:
• SwissDial
          <xref ref-type="bibr" rid="ref19">(Pelin Dogan-Scho¨nberger, 2021)</xref>
          :
26 hours of Swiss German
• ArchiMob (Tanja Samardzic, 2016): 80 hours
of Swiss German
• Common Voice German v4: 483 hours of
German4
The SwissDial dataset consists of 26 hours of
audios in 8 different Swiss dialects with
corresponding transcriptions in Swiss dialect and Standard
German translations. The Swiss German
transcription rules differ between dialects. ArchiMob
contains 70 hours of audios in 14 different Swiss
dialects with transcription in Swiss German, where
each word is additionally provided with a Standard
German normalisation. The transcription rules are
normalised and are equal for all dialects (Dieth
transcription,
          <xref ref-type="bibr" rid="ref7">(Dieth and Schmid-Cadalbert, 1986)</xref>
          ).
Common Voice German v4 consists of 483 hours
of audios in Standard German with corresponding
transcriptions.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>4.2 Jasper 4.2.1 Model</title>
        <p>
          We used the Jasper
          <xref ref-type="bibr" rid="ref13">(Li et al., 2019)</xref>
          configuration
corresponding to our best submission in the
pre3https://pypi.org/project/num2words/
4https://commonvoice.mozilla.org/en/
datasets/
decessor of this Shared Task
          <xref ref-type="bibr" rid="ref3">(Bu¨chi et al., 2020)</xref>
          .
The Acoustic Model as per Bu¨chi et al. (2020)
consists of 10x5 blocks and was pre-trained on 537
hours of Standard German data (see Bu¨chi et al.
(2020), Table 2). In all reported experiments, we
fine-tuned five blocks on the Shared Task data as
described in Section 5.2 below. We used last year’s
extended language model, a 6-gram model trained
with KenLM, without further fine-tuning on this
year’s data. For the data sources, see Table 2 in
Bu¨chi et al. (2020). Decoding was done using beam
search with a beam size of 1024.
4.2.2
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Data</title>
        <p>We extracted the audios to 64-dimensional
melfilterbank features with 20ms window size and
10ms overlap as input to the Jasper acoustic model.
The reference texts were preprocessed as described
in Bu¨chi et al. (2020). No additional Swiss German
audio data was used for training Jasper.
4.3
4.3.1</p>
      </sec>
      <sec id="sec-3-5">
        <title>Wav2vec XLSR-53</title>
      </sec>
      <sec id="sec-3-6">
        <title>Model</title>
        <p>Wav2vec XLSR-53 is a cross-lingual extension of
wav2vec 2.0 as per Baevski et al. (2020).
Pretrained on 53 different languages, it attempts to
learn a quantisation of the latent representations
shared across languages by solving a contrastive
task over masked speech representations. In the
experiment below, we fine-tuned wav2vec
XLSR53 on the Shared Task data. No explicit language
model was used to conduct the experiment.
4.3.2</p>
      </sec>
      <sec id="sec-3-7">
        <title>Data</title>
        <p>The labelled data used for fine-tuning XLSR-53
was based on the task training data. However, it
was further pre-processed removing all utterances
which contained special characters or were detected
as not being in German using langdetect5. Numeric
values were replaced by strings using num2words6.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments on Individual Models</title>
      <p>Sections 5.1 through 5.3 present the experiments
we performed to improve the individual models and
provide the BLEU scores achieved in each
experiment. We also discuss approaches to improve the
model outputs with the use of ensembling (Section
6) and post-processing (Section 7).</p>
      <p>5https://github.com/Mimino666/
langdetect
6https://pypi.org/project/num2words/
5.1
Below we describe the different models and
experimental results obtained with Fairseq. All
Experiments are trained with the same configuration as
described in Section 4.1 and can be divided into
three groups: extension of training data, inclusion
of a pre-trained encoder and ensembling.</p>
      <sec id="sec-4-1">
        <title>5.1.1 Extending the training data</title>
      </sec>
      <sec id="sec-4-2">
        <title>Fairseq F-SP-0.9 For F-SP-0.9 we trained</title>
        <p>the model from scratch on the Shared Task
training data. We used 176 hours, corresponding to an
Intersection over Union (IoU) greater or equal to
0.9.</p>
        <p>Fairseq F-SP-All We noted that the model
F-SP-0.9 generalises very poorly, so for
F-SP-All we trained a new model with the entire
task training data, which corresponds to 293 hours.
Despite partially poorly aligned translations, the
model benefits from the new data: the BLEU score
is improved by about 4.32 points.</p>
        <p>Fairseq F-SP-SD We decided to extend the
training data with the SwissDial Corpus. For this, we
trained a new model F-SP-SD with the entire task
training data plus all data from SwissDial. This
data extension improves the score by an additional
4.81 BLEU points in comparison to F-SP-All.</p>
      </sec>
      <sec id="sec-4-3">
        <title>5.1.2 Including pre-trained encoder</title>
        <p>Fairseq F-SP-DE We also investigated how to
improve the encoder (acoustic model). We
pretrained a Standard German (DE) encoder on the
Common Voice German v4 dataset. For F-SP-DE,
we added the pre-trained encoder and trained the
model on the entire Shared Task training data.
Including the DE encoder improves the score by 3.36
BLEU points in comparison to F-SP-All.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Fairseq F-SP-SD-DE Since both models</title>
        <p>F-SP-SD and F-SP-DE improved the BLEU
score, we decided to bring the two approaches
together. We trained a new model F-SP-SD-DE
with the entire Shared Task training data,
SwissDial data and include the pre-trained DE encoder
in the training. This brings an improvement of 8.37
BLEU points in comparison to F-SP-All.
Fairseq F-SP-AM-DE In this model we used the
entire task training data plus the data from
ArchiMob. For the training we included the pre-trained
DE encoder. This setup improves the BLEU score
by 14.01 in comparison to F-SP-All.</p>
        <p>Fairseq F-SP-SD-CH In order to further
improve the acoustic model, we trained an encoder
in Swiss German (CH) on the SwissDial and
ArchiMob dataset. We trained a new model
F-SP-SD-CH with the entire Shared Task
training data and SwissDial and included the pre-trained
CH encoder in the training. The BLEU score in
comparison to F-SP-All is improved by 12.54
points.</p>
      </sec>
      <sec id="sec-4-5">
        <title>5.1.3 Ensembling</title>
      </sec>
      <sec id="sec-4-6">
        <title>Fairseq Ensemble F-SP-SD &amp; F-SP-DE (F-E1)</title>
        <p>In this experiment, we ensembled the models
F-SP-SD and F-SP-DE. F-E1 achieves a BLEU
score of 28.74 . Ensembling is done with the
implementation provided by the Fairseq S2T Toolkit7. In
comparison to F-SP-SD-DE, which combines in
the training setup the same training dataset
SwissDial as F-SP-SD and the same DE encoder as
F-SP-DE, the ensembling performs slightly
better. In comparison to F-SP-All the BLEU score
improves by 9.94 points.</p>
      </sec>
      <sec id="sec-4-7">
        <title>Fairseq Ensemble F-SP-AM-DE &amp; F-SP-SD</title>
        <p>CH (F-E2) After the good performance of
F-E1, we decided to ensemble F-SP-AM-DE
and F-SP-SD-CH. This ensembling improves
the BLEU score in comparison to F-SP-All by
17.00 points.</p>
      </sec>
      <sec id="sec-4-8">
        <title>Fairseq F-E2 extended (F-E3) Finally, we</title>
        <p>trained a model on the entire available data for
Swiss German (task, SwissDial and ArchiMob)
and used this model to perform ensembling on top
of F-E2. For time reasons, we were not able to
complete the training and the output of this model
could not been included in the final solution
presented in 6. We only evaluated an intermediate
status of the model and achieved a score of 36.83
BLEU points. In comparison to F-SP-All, it
improves the score by 18.03 points.</p>
        <p>Table 1 shows the public BLEU scores
obtained with the Fairseq models on the Shared
Task public part of the test set. The table contains
additional information about applied train sets and
encoders. F-E3 achieved the best performance
with a BLEU score of 36.83 on the public part of
the test set (37.4 on the private part). In addition
to ensembling, the inclusion of a CH encoder in
7https://github.com/pytorch/fairseq/
issues/223
the training process as well as the extension of the
training data with the ArchiMob corpus benefited
the model performance most.</p>
        <p>Jasper-FT For Jasper-FT we fine-tune the
pre-trained Standard German model on the Shared
Task training data. We used 169 hours, sampled
from the set with an IoU greater or equal to 0.9,
which were augmented to 507 hours using 90% and
110% speed perturbation as in Bu¨chi et al. (2020).
Jasper-PL We noted that the task test set
differs acoustically from the training data since
different dialects are present and the audio quality
tends to be lower. This motivated the creation of
Jasper-PL, where we used pseudo-labeling on
the test set. More precisely, we used the
hypotheses of Jasper-FT on the task test set to fine-tune
Jasper-FT for 20 additional epochs.</p>
        <p>Jasper-PL-E We decided to further work on the
(comparatively) low-quality audio of the task test
set and used the Dolby Media Enhance API v1.18
to create an ”enhanced” version of the task test
set. The Enhance API automatically improves the
quality of audio files, e.g. by correcting the volume
and reducing noise and hum. We then fine-tuned
Jasper-FT on this data, this time using the
hypotheses provided by Jasper-PL as labels since
these achieve a higher BLEU score.</p>
        <p>
          Table 2 shows the public BLEU scores obtained
with the Jasper models on the two different test
sets (Jasper-PL-E was only evaluated on the
enhanced test set). The best-performing Jasper model
is Jasper-PL with a BLEU score of 32.97 on
the public part of the test set. Using the enhanced
audio data does not confer any advantage on either
prediction or pseudo-label fine-tuning compared
to the as-is data. We can, however, see the
benefit of rather naive pseudo-labelling in this setting
where training and testing data are quite different.
Future work could expand on the use of
pseudolabelling by using more advanced setups, such as
confidence-based
          <xref ref-type="bibr" rid="ref12 ref27">(Kahn et al., 2020)</xref>
          or iterative
          <xref ref-type="bibr" rid="ref27">(Xu et al., 2020)</xref>
          pseudo-labelling.
wav2vec XLSR-53 FT For wav2vec
XLSR-53 FT we fine-tuned the pre-trained
baseline (as published on HuggingFace9) on the
Shared Task training data. We used 227 hours,
corresponding to an IoU greater or equal than 0.8.
The data was pre-processed as outlined in Section
4.3.2.
Having trained and evaluated the three models
described in Sections 4.1, 4.2 and 4.3, we performed
experiments with two ensembling methods:
majority voting and a hybrid technique combining
majority voting with perplexity calculation. We
used the outputs of the best-performing models
of each of the three systems, aiming to select the
8https://dolby.io/developers/
media-processing/api-reference/enhance
9https://huggingface.co/facebook/
wav2vec2-large-xlsr-53
most reliable translation for each utterance from
among them. The best-performing models were
F-E2 (BLEU score of 35.8010), Jasper-PL
(BLEU score of 32.97) and wav2vec XLSR-53
FT (BLEU score of 30.4).
        </p>
        <p>The models were first categorised based on their
BLEU scores into a primary, first auxiliary and
second auxiliary models. F-E2 with the highest score
was selected as the primary model, Jasper-PL
with the second best score was set as the first
auxiliary model and wav2vec XLSR-53 FT was
used as the second auxiliary model.</p>
        <p>In the first step, we aligned the hypotheses of the
three models and extracted text passages where all
three hypotheses agree, leaving only text excerpts
where the hypotheses disagree.</p>
        <p>Majority Voting (MV) The majority voting
consisted in collecting votes for each text excerpt
defined in the previous step: a particular hypothesis
receives a vote for each word it has in common
with any other hypothesis. The hypothesis with the
most votes is chosen as the best candidate
translation. If multiple hypotheses score the same, the
output of the model categorised higher in the
hierarchy (primary, first auxiliary, second auxiliary) is
selected.</p>
        <p>Hybrid Ensembling (HE) The hybrid
ensembling method combines majority voting with
perplexity calculation. If more than one hypothesis
scores maximum and the hypotheses with the
maximum score are not equal, the perplexity of the
hypotheses is calculated. To this end, we extended the
particular text excerpt with 3 context words
preceding and following the excerpt. For these text
segments, we calculated perplexity with a pre-trained
uncased German BERT model11. The hypothesis
with the lower perplexity was selected.</p>
        <p>The results of the experiments are presented in
Table 4. Out of the two algorithms we applied on
the data, better results could be achieved with the
majority voting. The BLEU score improved by 2.9
points from 35.80 to 38.70 when compared to the
result of the best model (F-E2).</p>
        <p>10F-E3 as a last-minute submission could not be used for
ensembling</p>
        <p>11https://github.com/dbmdz/berts#
german-bert
Next to the Language Models for Speech
Recognition, we evaluated an approach to using text-only
data by training a supervised ”spelling correction”
(SC) model to correct the errors made by the STT
model explicitly. Instead of predicting the
likelihood of emitting a word based on the surrounding
context, the SC model only needs to identify likely
errors in the STT model output and propose
alternatives. Intuitively, this task highly depends on the
baseline model’s quality: if the model transcribes
very well, this task can be reduced to simply
copying the input transcript directly to the output.</p>
        <p>
          Most recent approaches for transcript
postprocessing use a transformer-based method:
          <xref ref-type="bibr" rid="ref14">(Liao
et al., 2021)</xref>
          use a modified RoBERTa structure
and show an increase of 17.53 BLEU points on
the self-augmented English Conversational
Telephone Speech data set. On the LibriSpeech dataset,
          <xref ref-type="bibr" rid="ref10">(Hrinchuk et al., 2019)</xref>
          show promising results
using a pre-trained BERT as initialisation for their
spell correction model, while
          <xref ref-type="bibr" rid="ref9">(Guo et al., 2019)</xref>
          takes a different approach with a bidirectional
LSTM.
        </p>
        <p>We compared different Transformer
architectures with their corresponding open-sourced
pretrained models and other post-processing methods.</p>
        <p>The objective for all transformer models was set
to next-sentence prediction (sequence to sequence
generation) with a vocabulary size of 30’000, batch
size of 16, and beam size for beam search set to 5.
The models were initialised with pre-trained
German embeddings and fine-tuned for up to 120’000
steps on the Shared Task training set described in
2.</p>
        <p>
          • BERT
          <xref ref-type="bibr" rid="ref6">(Devlin et al., 2018)</xref>
          , having both
encoder and decoder initialised with pre-trained
weights.
• DistilBERT
          <xref ref-type="bibr" rid="ref21">(Sanh et al., 2020)</xref>
          , the
lightweight alternative to BERT, reducing the
training time up to 60%.
• ELECTRA
          <xref ref-type="bibr" rid="ref5">(Clark et al., 2020)</xref>
          , which uses a
more sample-efficient pre-training approach
for the encoder, called replaced token
detection.
• SymSpell
          <xref ref-type="bibr" rid="ref8">(Garbe, 2020)</xref>
          ,which is a spelling
correction algorithm for correcting spelling
errors based on Damerau-Levenshtein distances,
stored in a pre-trained dictionary.
        </p>
        <p>The following table shows the BLEU scores
on the public test set, when performing
postprocessing on the output of the majority voting
algorithm as described in 6. The Baseline refers
to the BLEU score of the non-processed majority
voting output.</p>
        <p>As the evaluations show, most post-processing
attempts decrease the overall BLEU score, with
SymSpell as the most straightforward approach
performing best. Compared with previous work in this
area, this could be explained by the limited amount
of data available for training the transformer
models. Due to lack of performance, we exclude the
post-processing step in our final solution.
8</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this paper, we presented our contribution to the
Shared Task ”Swiss German Speech to Standard
German Text” at SwissText 2021. Our solution
combines the outputs of three models based on
Fairseq, Jasper and Wav2vec XLSR-53
architectures. Because of time and resource constraints,
we used only the labeled data set. Out of the 21
experiments we performed with the models,
including transcript post-processing and ensembling, we
achieved the best result by applying an ensembling
method on the outputs of Fairseq model F-E2
(BLEU score of 35.80) as the primary model, and
Jasper-PL (32.97) and wav2vec XLSR-53
FT (30.39) as auxiliary models. We processed the
three models’ predictions with a majority voting
algorithm and this way retrieved the most reliable
candidate out of the provided translations for each
utterance in the public test set. With this solution,
we achieved a BLEU score of 39.39 on the private
test set, which resulted in the third place out of four
contributors in the competition.</p>
      <p>Swiss German is a low-resource language, which
makes training an STT or a Speech Translation
system a challenging task. However, our experiments
show that applying ensembling both on various
models of the same architecture (as in Fairseq
models F-E1, F-E2 and F-E3) and on models based
on various architectures (as implemented in our
final solution) trained with limited data can lead
to a score improvement of several BLEU points.
Pseudo-labeling is another approach which
contributes to model enhancement as we could observe
with the Jasper-PL model. We will be further
investigating these two methods aiming at
improving the results despite the limited data currently
available for Swiss German.
Chengyi Wang, Yu Wu, Shujie Liu, Zhenglu Yang, and
Ming Zhou. 2020b. Bridging the Gap between
PreTraining and Fine-Tuning for End-to-End Speech
Translation. In Proceedings of the AAAI Conference
on Artificial Intelligence, volume 34, pages 9161–
9168.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Alexei</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <surname>Henry Zhou</surname>
            , Abdelrahman Mohamed, and
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Auli</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations</article-title>
          .
          <source>Facebook AI</source>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          Alexandre Be´rard, Olivier Pietquin, Christophe Servan, and
          <string-name>
            <given-names>Laurent</given-names>
            <surname>Besacier</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Listen and Translate: A Proof of Concept for End-to-End Speech-to-Text Translation</article-title>
          .
          <source>arXiv preprint arXiv:1612</source>
          .
          <fpage>01744</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Bu</surname>
          </string-name>
          <article-title>¨chi, Malgorzata Anna Ulasik, Manuela Hu¨rlimann, Fernando Benites, Pius von Da¨niken, and</article-title>
          <string-name>
            <given-names>Mark</given-names>
            <surname>Cieliebak</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>ZHAW-InIT at GermEval 2020 Task 4</article-title>
          :
          <string-name>
            <surname>Low-Resource</surname>
          </string-name>
          Speech-to-Text.
          <source>In Proceedings of the 5th Swiss Text Analytics Conference (SwissText) &amp; 16th Conference on Natural Language Processing (KONVENS)</source>
          .
          <article-title>CEUR-WS.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Jiatao</given-names>
            <surname>Gu Changhan Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Juan</given-names>
            <surname>Pino</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Improving Cross-Lingual Transfer Learning for End-to-End Speech Recognition with Speech Translation</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <surname>Minh-Thang</surname>
            <given-names>Luong</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quoc</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
            , and
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>ELECTRA: Pretraining Text Encoders as Discriminators Rather Than Generators</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Eugen</given-names>
            <surname>Dieth</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Schmid-Cadalbert</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Schwyzertu¨tschi diala¨ktschrift</article-title>
          . Sauerla¨nder, Aarau,
          <volume>2</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Wolf</given-names>
            <surname>Garbe</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>SymSpell: Fast spell correction algorithm</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Jinxi</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Tara N.</given-names>
            <surname>Sainath</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ron J.</given-names>
            <surname>Weiss</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A Spelling Correction Model for End-to-End Speech Recognition</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Oleksii</given-names>
            <surname>Hrinchuk</surname>
          </string-name>
          , Mariya Popova, and
          <string-name>
            <given-names>Boris</given-names>
            <surname>Ginsburg</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Correction of Automatic Speech Recognition with Transformer Sequence-to-sequence Model</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Javier</given-names>
            <surname>Iranzo-Sa</surname>
          </string-name>
          ´nchez,
          <string-name>
            <given-names>Joan</given-names>
            <surname>Albert</surname>
          </string-name>
          Silvestre-Cerda`,
          <string-name>
            <surname>Javier</surname>
            <given-names>Jorge</given-names>
          </string-name>
          , Nahuel Rosello´, Adria` Gime´nez,
          <string-name>
            <surname>Albert</surname>
            <given-names>Sanchis</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Jorge</given-names>
            <surname>Civera</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Alfons</given-names>
            <surname>Juan</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Europarl-ST: A Multilingual Corpus for Speech Translation of Parliamentary Debates</article-title>
          .
          <source>In ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>8229</fpage>
          -
          <lpage>8233</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Jacob</surname>
            <given-names>Kahn</given-names>
          </string-name>
          , Ann Lee, and
          <string-name>
            <given-names>Awni</given-names>
            <surname>Hannun</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Self-training for End-to-End Speech Recognition</article-title>
          .
          <source>In ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>7084</fpage>
          -
          <lpage>7088</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Jason</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Vitaly</given-names>
            <surname>Lavrukhin</surname>
          </string-name>
          , Boris Ginsburg, Ryan Leary, Oleksii Kuchaiev, Jonathan M Cohen,
          <string-name>
            <given-names>Huyen</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , and Ravi Teja Gadde.
          <year>2019</year>
          .
          <article-title>Jasper: An End-to-End Convolutional Neural Acoustic Model</article-title>
          .
          <source>In Proceedings of Interspeech</source>
          <year>2019</year>
          , pages
          <fpage>71</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Junwei</given-names>
            <surname>Liao</surname>
          </string-name>
          , Yu Shi, Ming Gong, Linjun Shou, Sefik Eskimez, Liyang Lu, Hong Qu, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Zeng</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Generating Human Readable Transcript for Automatic Speech Recognition with Pretrained Language Model</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Ney</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Speech Translation: coupling of recognition and translation</article-title>
          .
          <source>In 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No.99CH36258)</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>517</fpage>
          -
          <lpage>520</lpage>
          vol.
          <volume>1</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Vassil</given-names>
            <surname>Panayotov</surname>
          </string-name>
          , Guoguo Chen, Daniel Povey, and
          <string-name>
            <given-names>Sanjeev</given-names>
            <surname>Khudanpur</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Librispeech: An ASR corpus based on public domain audio books</article-title>
          .
          <source>In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pages
          <fpage>5206</fpage>
          -
          <lpage>5210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Kishore</given-names>
            <surname>Papineni</surname>
          </string-name>
          , Salim Roukos, Todd Ward, and
          <string-name>
            <given-names>WeiJing</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>BLEU: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In Proceedings of the 40th annual meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Daniel S Park</surname>
            ,
            <given-names>William</given-names>
          </string-name>
          <string-name>
            <surname>Chan</surname>
          </string-name>
          , Yu Zhang, Chung-Cheng Chiu, Barret Zoph,
          <string-name>
            <surname>Ekin D Cubuk</surname>
          </string-name>
          , and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Hofmann Pelin</surname>
          </string-name>
          Dogan-Scho¨nberger, Julian Ma¨der.
          <year>2021</year>
          .
          <article-title>SwissDial: Parallel Multidialectal Corpus of Spoken Swiss German</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Michel</given-names>
            <surname>Plu</surname>
          </string-name>
          <article-title>¨ss, Lukas Neukom</article-title>
          , and
          <string-name>
            <given-names>Manfred</given-names>
            <surname>Vogel</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>SwissText 2021 Task 3: Swiss German Speech to Standard German Text</article-title>
          . In preparation.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Victor</given-names>
            <surname>Sanh</surname>
          </string-name>
          , Lysandre Debut, Julien Chaumond, and
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Wolf</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Tanja</given-names>
            <surname>Schultz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vogel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Saleem</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Using Word Lattice Information for a Tighter Coupling in Speech Translation Systems</article-title>
          . In INTERSPEECH.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Beat</given-names>
            <surname>Siebenhaar</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Sprachgeographische Aspekte der Morphologie und Verschriftung in schweizerdeutschen Chats</article-title>
          .
          <source>Linguistik online</source>
          ,
          <volume>15</volume>
          (
          <issue>3</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Beat</given-names>
            <surname>Siebenhaar</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alfred</given-names>
            <surname>Wyler</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Dialekt und Hochsprache in der deutschsprachigen Schweiz</article-title>
          . Pro Helvetia.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Sperber</surname>
          </string-name>
          and
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Paulik</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Speech Translation and the End-to-End Promise: Taking Stock of Where We Are</article-title>
          .
          <source>In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          , pages
          <fpage>7409</fpage>
          -
          <lpage>7421</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Woszczyna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Coccaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Eisele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lavie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McNair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Polzin</surname>
          </string-name>
          , I. Rogina,
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Rose</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Sloboda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tomita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tsutsumi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Aoki-Waibel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Waibel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Ward</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Recent Advances in Janus: A Speech Translation System</article-title>
          .
          <source>In Proceedings of the Workshop on Human Language Technology, HLT '93, page 211-216</source>
          , USA. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Qiantong</given-names>
            <surname>Xu</surname>
          </string-name>
          , Tatiana Likhomanenko, Jacob Kahn, Awni Hannun, Gabriel Synnaeve, and
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Collobert</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Iterative Pseudo-Labeling for Speech Recognition</article-title>
          .
          <source>In Proceedings of Interspeech</source>
          <year>2020</year>
          , pages
          <fpage>1006</fpage>
          -
          <lpage>1010</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>