<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Wav2vec 2.0 Fine-tuning For The SE&amp;R 2022 Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alef Iury Siqueira Ferreira</string-name>
          <email>alef_iury_c.c@discente.ufg.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gustavo dos Reis Oliveira</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Federal University of Goias</institution>
          ,
          <addr-line>Goiânia</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents our eforts to build a robust ASR model for the shared task and prepared speech &amp; Speech Emotion Recognition in Portuguese (SE&amp;R 2022). The goal of the challenge is to advance the ASR research for the Portuguese language, considering prepared and spontaneous speech in diferent dialects. Our method consist on fine-tuning an ASR model in a domain-specific approach, applying gain normalization and selective noise insertion. The proposed method improved over the strong baseline provided on the test set in 3 of the 4 tracks available.</p>
      </abstract>
      <kwd-group>
        <kwd>speech recognition</kwd>
        <kwd>Portuguese</kwd>
        <kwd>prepared speech</kwd>
        <kwd>spontaneous speech</kwd>
        <kwd>wild data</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The performance of Automatic Speech Recognition
systems (ASRs) has increased significantly with the
development of modern neural network topologies and the
use of massive amount of data to train the models [1].
Although the accuracy of recent models improved for
high-resource languages, such as English, the
development of ASR models in other languages is still a dificult
task using the same technologies [
        <xref ref-type="bibr" rid="ref17 ref25">2, 3</xref>
        ]. In this scenario,
Self-Supervised Learning (SSL), a method in which
representations with semantic information are learned by
using unlabelled data, emerged as an important advance,
allowing the training of deeper models using less labelled
data [4, 5]. In this line of work, this paper explores the use
of the Wav2vec 2.0 [6], a framework for self-supervised
learning of discrete representations from raw audio data.
      </p>
      <p>Wav2vec 2.0 (Figure 1) is inspired by previous works
in unsupervised pre-training for speech recognition, that
is, Wav2vec [7] and Vq-Wav2vec [4]. During pre-training
the model learns speech representations solving a
contrastive task which requires identifying the correct
quantized latent speech representations of a masked time step
among a set of distractors. After the self-supervised
pretraining, the model can be fine-tuned on labeled data in a
supervised task, like ASR, adding a randomly initialized
linear projection on top of the context network with 
classes and a loss function specific to the task at hand,
like CTC, for instance.</p>
      <sec id="sec-1-1">
        <title>The model shows important results among low re</title>
        <p>Proceedings of the First Workshop on Automatic Speech Recognition
for Spontaneous and Prepared Speech &amp; Speech Emotion Recognition
in Portuguese (SE&amp;R 2022), co-located with PROPOR 2022. March 21st,
nEvelop-O
2022 (Online).
[9] demonstrated that the fine-tuning of the Wav2vec
2.0 model achieves state-of-the-art (SOTA) results only
using publicly available datasets.</p>
      </sec>
      <sec id="sec-1-2">
        <title>An important aspect to consider when training an</title>
        <p>
          ASR model is the quality and the domain of the data [
          <xref ref-type="bibr" rid="ref50">10,
11, 12</xref>
          ]. While most of the available public datasets are
composed of prepared speech [9], mostly read sentences
[13, 14], the domain of real ASRs are far more complex,
mainly because it is formed by spontaneous speech and
diferent speech dialects. Quality is another issue: most
of ASR use cases involve high noise environments or low
recording equipment, which is not adressed in most of
the public datasets available [9].
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>To stimulate research that can advance the present</title>
        <p>SOTA in ASR in Portuguese, for both prepared and
spontaneous speech, the shared-task Automatic Speech
Recognition for spontaneous and prepared speech &amp; Speech
Emotion Recognition in Portuguese (SE&amp;R 2022) introduces a
new baseline for ASR and a new dataset in Portuguese
[9]. The Corpus of Annotated Audios (CORAA ASR),
a large corpus of spontaneous and prepared speech, is
composed by various subsets in Portuguese with
diferent characteristics. The baseline achieves a Word Error
Rate (WER) of 24.18% on CORAA ASR test set, a dificult
dataset containing samples with low quality, noise, and
a variety of domains and dialects.</p>
        <p>In this work, we investigate the fine-tuning of the
baseline model [9] proposed by the shared task, a fine-tuned
model based on the Wav2vec 2.0 XLSR-53 [15], using
only public available Portuguese datasets, including the
CORAA ASR dataset. We conducted several experiments
in diferent domains for the challenge and explored the
use of selective noise insertion and audio normalization
during training. This work is organized as follows:
Section 2 discuss the proposed methods and Section 3 shows
and discusses the obtained results. Finally, Section 4
presents the conclusions of this work.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <sec id="sec-2-1">
        <title>2.1. Datasets</title>
        <p>We used several publicly available datasets in Portuguese.
Besides CORAA ASR, most of them is composed by
prepared speech. In general, we opted to use all the data in
the gathered datasets for training, except the dev part of
the CORAA ASR, as presented in Table 1. The datasets
used in this work are:
• CETUC [13]: contains approximately 145 hours
of Brazilian Portuguese speech distributed among
50 male and 50 female speakers, each
pronouncing approximately 1,000 phonetically balanced
sentences selected from the CETEN-Folha1
corpus;
• Common Voice (CV) 7.0 [16]: is a project
proposed by Mozilla Foundation with the goal to
create a wide open dataset in diferent languages.
In this project, volunteers donate and validate
speech using the oficial site 2;
• Multilingual LibriSpeech (MLS) [14]: a massive
dataset available in many languages. The MLS is
based on audiobook recordings in public domain
like LibriVox3. The dataset contains a total of
6k hours of transcribed data in many languages.
The set in Portuguese used in this work4 (mostly
Brazilian variant) has approximately 284 hours
of speech, obtained from 55 audiobooks read by
62 speakers;
• Multilingual TEDx [17]: a collection of audio
recordings from TEDx talks in 8 source languages.
The Portuguese set (mostly Brazilian Portuguese
variant) contains 164 hours of transcribed speech.
1https://www.linguateca.pt/cetenfolha/
2https://commonvoice.mozilla.org/pt
3https://librivox.org/
4http://www.openslr.org/94/
• Corpus of Annotated Audios (CORAA ASR) v1
[9]: is a public available dataset that contains
290.77 hours of validated pairs (audio-transcription)
in Portuguese (mostly Brazilian Portuguese
variant) and is comprised by five other corpora: ALIP
[18], C-ORAL Brasil I [19], NURC-Recife [20],
SP2010 [21] and TEDx Portuguese talks.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Experiments</title>
        <sec id="sec-2-2-1">
          <title>Our experiments consists on the fine-tuning of the base</title>
          <p>line model of [9]. For each experiment, we trained the
model for 5 epochs, using a batch size of 192, and
using Adam [22] where the learning rate is 3−05 that is
warmed up for the first 400 updates, then linearly
decayed for the remained. For the experiments, we used a
NVIDIA TESLA V100 32GB, a NVIDIA TESLA Tesla P100
16GB and a NVIDIA A100 80GB, depending on the type
of audio pre-processing used. The code to replicate the
results is available at https://github.com/alefiury/SE-R_
2022_Challenge_Wav2vec2.</p>
          <p>In total, we conducted five main experiments to test
our methods:
• Experiment 1: Wav2vec 2.0 XLSR-53 - Base: For
this experiment, the model was fine-tuned
considering the whole train set, but did not receive
neither normalization nor noise addition;
• Experiment 2: Wav2vec 2.0 XLSR-53 - Norm: The
model was fine-tuned with the whole train set
with normalization. For the normalization, the
mean gain of all the audios in the train set was
considered;
• Experiment 3: Wav2vec 2.0 XLSR-53 - Norm and
SNA: The model was fine-tuned with gain
normalization and selective noise addition. The audios
were normalized considering the mean gain of all
the audios in the train set, and those audios
pertaining to datasets that were considered to have a
low presence of noise, namely MLS and CETUC,
received randomly one of the following 5
possible types of noises: additive noise, being music
or nonspeech noises from the MUSAN Corpus
[23], Room impulse responses [24], Addition or
reduction of gain, Pitch shift and Gaussian noise;
• Experiment 4: Wav2vec 2.0 XLSR-53 - Norm +
Prepared Speech: Model fine-tuned based on the
ifnal trained model of Experiment 2, but
considering just the prepared speech data from the
CORAA ASR dataset, trained for more 5 epochs;
• Experiment 5: Wav2vec 2.0 XLSR-53 - Norm +
Spontaneous Speech: Model fine-tuned based
on the final trained model of Experiment 2, but
considering just the spontaneous speech data</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>3.2. Final Results</title>
        <p>Table 4 compares the baseline with our selected models
in the test set. The model Wav2vec 2.0 XLSR-53 - Norm
3. Results and Discussion + Prepared Speech surpassed the strong baseline in the
Prepared Speech PT_BR, Prepared Speech PT_PT and the
The shared-task consists of 4 tracks. Each track has a Mixed tracks. As seen in the results based on the dev set,
domain specific scenario, that includes prepared speech the fact that most of the data of the datasets that were
and spontaneous speech. In this regard, we conducted added to the train set are comprised of prepared speech,
a prior analysis (Section 3.1) using the dev set to select might have contributed to the increase in performance
the best approaches based on 3 of the 4 tracks available: in this domain in both Portuguese variants. Lastly, even
Mixed, Prepared Speech PT_BR and Spontaneous Speech. though we were not able to surpass the baseline model in
The best models were selected and then submitted for the Spontaneous Speech track, we achieved competitive
evaluation. Our final results are presented in Section 3.2. results with both submitted models.</p>
      </sec>
      <sec id="sec-2-4">
        <title>3.1. dev set Analysis</title>
      </sec>
      <sec id="sec-2-5">
        <title>3.3. Additional Experiments</title>
        <p>Overall, our models did not show a huge improvement in After selecting and submitting the best results, we
perperformance when compared to the baseline model. Even formed some additional experiments to further explore
though we fine-tuned a model that is considered the state- our proposed methods using the dev set. We tried to use
of-the-art in Brazilian Portuguese, we suspect that the text correction in the outputs of the ASR models, and
number of training epochs might have been insuficient train the normalized models for a longer period of time
to obtain an increase in performance, or that the baseline using early stopping considering the prepared speech and
model might have already reached a local optima. spontaneous speech data from the CORAA ASR dataset.</p>
        <p>Furthermore, as presented in Table 2, the model that The text correction was done with an additional
postwas fine-tuned using prepared speech clearly improved processing step, using a KenLM [25] model. For the
difthe results in the Prepared Speech subset (and conse- ferent tasks, we used 2 diferent KenLM models: one for
quently the Mixed subset). The same phenomenon could spontaneous speech, which was built using subsets of
not be seen in the Spontaneous Speech subset. A possible the CORAA ASR dataset containing spontaneous speech
explanation to this fact is that most of the data of the phrases. And the other one was built considering wikipedia
datasets that were added to the train set are comprised in portuguese texts, as proposed by [3]. Both were
4of prepared speech, which might have contributed to the grams. We found that this post-processing of the
predicincrease in performance in this particular domain. An- tions from the ASR models did not improve the results
other possible explanation is the low number of training from the dev set on the Prepared Speech PT_BR and
epochs used to train the models. Spontaneous Speech tracks, as can be seen in Table 5, in</p>
        <p>Additionally, the noise insertion did not gave further fact they were worse. One possible explanation is that
improvement in performance. Nevertheless, the results some of the decoder hyper-parameters did not work well
of the SDA model in some more noisy subsets of the with our ASR models. Another possibility is that the
CORAA ASR dataset, like ALIP and NURC, for instance, 4-gram trained with spontaneous text was built with a
showed some interesting and promising results when small amount of text, which might have decreased the
compared to the baseline. These results are shown in performance of the model. However, the results on the
Table 3. Prepared Speech PT_PT track were much better
compared to the previous experiments. This result suggests
that the LM might improve results when there are few
domain data used to train the Wav2vec model, since most of
our training data was composed by Brazilian Portuguese
audios.</p>
        <p>Furthermore, as we had suspected earlier, the model
with gain normalization that were trained considering the
spontaneous speech data from the CORAA ASR corpus
and for a longer period of time performed better in its
respective subtrack, strengthening our hypothesis that
our results did not improve in the main experiments due
to a low number of training epochs.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusions</title>
      <sec id="sec-3-1">
        <title>In this work we presented our eforts to build a robust ASR model using multiple approaches, such as selective noise insertion and domain specific fine-tuning. In our experiments we found that fine-tuning a strong baseline</title>
        <p>with additional public available data in multiple domains
and using normalization, even for a few epochs, can
improve performance. With our results we were able to
improve on the test set in 3 of the 4 tracks available over
the strong baseline provided.</p>
        <p>As future works, we plan to train a ASR model using
a dynamic noise insertion approach that do not depend
on choosing specific datasets previously.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This research was funded by CEIA with support by the
Goiás State Foundation (FAPEG grant #201910267000527)5.
We also would like to thank Cyberlabs Group6 for the
support for this work.</p>
      <sec id="sec-4-1">
        <title>5http://centrodeia.org/ 6https://cyberlabs.ai/</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>speech manually validated for speech recognition [1</article-title>
          ]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Recent advances in end-to-end automatic in brazilian portuguese</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2110</volume>
          .
          <fpage>15731</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>speech recognition</source>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2111</volume>
          .
          <fpage>01690</fpage>
          . [10]
          <string-name>
            <surname>M. L. Seltzer</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            , An investigation [2]
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ananthanarayanan</surname>
          </string-name>
          , R. Anubhai,
          <article-title>of deep neural networks for noise robust speech</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Battenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Case</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Casper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Catanzaro</surname>
          </string-name>
          , recognition, in: 2013 IEEE international conference
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Q.</given-names>
            <surname>Cheng</surname>
          </string-name>
          , G. Chen, et al.,
          <article-title>Deep speech 2: End-to- on acoustics, speech and signal processing</article-title>
          , IEEE,
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>end speech recognition in english and mandarin,</article-title>
          <year>2013</year>
          , pp.
          <fpage>7398</fpage>
          -
          <lpage>7402</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>in: International conference on machine learning</source>
          , [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Likhomanenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Pratap</surname>
          </string-name>
          , P. Tomasello,
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>PMLR</surname>
          </string-name>
          ,
          <year>2016</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          . J.
          <string-name>
            <surname>Kahn</surname>
            , G. Avidov,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Collobert</surname>
            , G. Synnaeve, Re[3]
            <given-names>I. M.</given-names>
          </string-name>
          <string-name>
            <surname>Quintanilha</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          <string-name>
            <surname>Netto</surname>
            ,
            <given-names>L. W. P.</given-names>
          </string-name>
          <string-name>
            <surname>Biscainho</surname>
          </string-name>
          ,
          <article-title>thinking evaluation in asr: Are our models robust</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>An open-source end-to-end asr system for brazilian enough?</article-title>
          , arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>11745</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>portuguese using dnns built from newly assembled</article-title>
          [12]
          <string-name>
            <given-names>W.-N.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sriram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          , T. Likhomanenko,
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>tion Systems</source>
          <volume>35</volume>
          (
          <year>2020</year>
          )
          <fpage>230</fpage>
          -
          <lpage>242</lpage>
          . G. Synnaeve, et al.,
          <source>Robust wav2vec 2</source>
          .0:
          <issue>Analyzing</issue>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Auli, vq-wav2vec: domain shift in self-supervised pre-training, arXiv</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>Self-supervised learning of discrete speech rep-</article-title>
          preprint
          <source>arXiv:2104.01027</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          resentations, in: International Conference on [13]
          <string-name>
            <given-names>V.</given-names>
            <surname>Alencar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alcaim</surname>
          </string-name>
          ,
          <article-title>Lsf and lpc-derived features</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>Learning Representations (ICLR)</source>
          ,
          <year>2020</year>
          .
          <article-title>URL: https: for large vocabulary distributed continuous speech</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          //openreview.net/pdf?id=rylwJxrYDS.
          <article-title>recognition in brazilian portuguese</article-title>
          , in:
          <year>2008</year>
          42nd [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaiswal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Babu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Z.</given-names>
            <surname>Zadeh</surname>
          </string-name>
          , D. Baner- Asilomar conference on signals,
          <source>systems and com-</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>jee</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Makedon</surname>
          </string-name>
          ,
          <article-title>A survey on contrastive self- puters</article-title>
          , IEEE,
          <year>2008</year>
          , pp.
          <fpage>1237</fpage>
          -
          <lpage>1241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>supervised learning, Technologies</source>
          <volume>9</volume>
          (
          <year>2021</year>
          )
          <article-title>2</article-title>
          . [14]
          <string-name>
            <given-names>V.</given-names>
            <surname>Pratap</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sriram</surname>
          </string-name>
          , G. Synnaeve, R. Col[6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Auli, wav2vec lobert, Mls: A large-scale multilingual dataset for</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>2.0: A framework for self-supervised learning speech research</article-title>
          ,
          <year>Interspeech 2020</year>
          (
          <year>2020</year>
          ). URL:
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>of speech representations</article-title>
          , in: H. Larochelle, http://dx.doi.org/10.21437/Interspeech.2020-
          <volume>2826</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Ranzato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hadsell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Balcan</surname>
          </string-name>
          , H. Lin doi:
          <volume>10</volume>
          .21437/interspeech.2020-
          <volume>2826</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          (Eds.),
          <source>Advances in Neural Information Pro</source>
          <volume>-</volume>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandelwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          , V. Chaud-
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>cessing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            As- hary, G. Wenzek,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Guzmán</surname>
          </string-name>
          , É. Grave, M. Ott,
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>sociates</surname>
          </string-name>
          , Inc.,
          <year>2020</year>
          , pp.
          <fpage>12449</fpage>
          -
          <lpage>12460</lpage>
          . URL: L.
          <string-name>
            <surname>Zettlemoyer</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Stoyanov</surname>
          </string-name>
          , Unsupervised cross-
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>https://proceedings.neurips.cc/paper/2020/file/ lingual representation learning at scale, in: Proceed-</mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          92d1e1eb1cd6f9fba3227870bb6d7f07-
          <fpage>Paper</fpage>
          .pdf.
          <source>ings of the 58th Annual Meeting of the Association</source>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Collobert</surname>
          </string-name>
          , M. Auli, for Computational Linguistics,
          <year>2020</year>
          , pp.
          <fpage>8440</fpage>
          -
          <lpage>8451</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <article-title>wav2vec: Unsupervised pre-training for speech</article-title>
          [16]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ardila</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Branson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kohler</surname>
          </string-name>
          , J. Meyer,
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          recognition.,
          <source>in: INTERSPEECH</source>
          ,
          <year>2019</year>
          . M.
          <string-name>
            <surname>Henretty</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Morais</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Saunders</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Tyers</surname>
            , G. We[8]
            <given-names>L. R. S.</given-names>
          </string-name>
          <string-name>
            <surname>Gris</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Casanova</surname>
          </string-name>
          , F. S. de Oliveira,
          <article-title>ber, Common voice: A massively-multilingual</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>A. da Silva</given-names>
            <surname>Soares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Junior</surname>
          </string-name>
          ,
          <article-title>Brazilian por- speech corpus</article-title>
          ,
          <source>in: Proceedings of the 12th Lan-</source>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <source>tuguese speech recognition using wav2vec 2.0</source>
          ,
          <year>2021</year>
          . guage Resources and Evaluation Conference,
          <year>2020</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>arXiv:2107.11414</source>
          . pp.
          <fpage>4218</fpage>
          -
          <lpage>4222</lpage>
          . [9]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Junior</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Casanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Soares</surname>
          </string-name>
          , F. S. de Oliveira, [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Salesky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiesner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bremerman</surname>
          </string-name>
          , R. Cattoni,
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>translation</surname>
          </string-name>
          ,
          <source>arXiv preprint arXiv:2102.01757</source>
          (
          <year>2021</year>
          ). [18]
          <string-name>
            <given-names>S. C. L.</given-names>
            <surname>Gonçalves</surname>
          </string-name>
          ,
          <article-title>Projeto alip (amostra linguística</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>do interior paulista)</source>
          e banco de dados iboruna:
          <volume>10</volume>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>brasileiro</surname>
          </string-name>
          , Estudos
          <string-name>
            <surname>Linguísticos</surname>
          </string-name>
          (São Paulo.
          <year>1978</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <volume>48</volume>
          (
          <year>2019</year>
          )
          <fpage>276</fpage>
          -
          <lpage>297</lpage>
          . URL: https://revistas.gel.org.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>br/estudos-linguisticos/article/view/2430. doi:10.</mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          21165/el.v48i1.
          <fpage>2430</fpage>
          . [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Raso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mello</surname>
          </string-name>
          ,
          <string-name>
            <surname>C-</surname>
          </string-name>
          oral - brasil i: Corpus de refer-
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <article-title>ência do portugues brasileiro falado informal</article-title>
          ,
          <year>2012</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          pp.
          <fpage>362</fpage>
          -
          <lpage>367</lpage>
          . doi:
          <volume>10</volume>
          .1007/978- 3-
          <fpage>642</fpage>
          - 28885- 2_
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          40. [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Oliveira</surname>
          </string-name>
          , Jr,
          <article-title>Nurc digital um protocolo para a dig-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <source>Linguistic Studies</source>
          <volume>3</volume>
          (
          <year>2016</year>
          )
          <fpage>149</fpage>
          -
          <lpage>174</lpage>
          . [21]
          <string-name>
            <surname>M. R B</surname>
            ,
            <given-names>O. L</given-names>
          </string-name>
          ,
          <article-title>Mapping paulistano portuguese: the</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>sp2010 project, Firenze, Italy: Fizenze University</mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          Press,
          <year>2012</year>
          , pp.
          <fpage>459</fpage>
          -
          <lpage>463</lpage>
          . [22]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          optimization, in: Y. Bengio, Y. LeCun (Eds.),
          <fpage>3rd</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <surname>tions</surname>
          </string-name>
          ,
          <source>ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-9,
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          2015, Conference Track Proceedings,
          <year>2015</year>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          http://arxiv.org/abs/1412.6980. [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Snyder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          , Musan: A
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          music, speech, and noise corpus,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <source>arXiv:1510</source>
          .
          <fpage>08484</fpage>
          . [24]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Peddinti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Seltzer</surname>
          </string-name>
          , S. Khudan-
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <article-title>speech for robust speech recognition</article-title>
          , in: 2017 IEEE
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          <source>Signal Processing (ICASSP)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>5220</fpage>
          -
          <lpage>5224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          <source>doi:10</source>
          .1109/ICASSP.
          <year>2017</year>
          .
          <volume>7953152</volume>
          . [25]
          <string-name>
            <given-names>K.</given-names>
            <surname>Heafield</surname>
          </string-name>
          , KenLM: Faster and smaller lan-
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          <string-name>
            <surname>inburgh</surname>
          </string-name>
          , Scotland,
          <year>2011</year>
          , pp.
          <fpage>187</fpage>
          -
          <lpage>197</lpage>
          . URL: https:
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>//aclanthology.org/W11-2123.</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>