<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Machine Translation of Covid-19 Information Resources via Multilingual Transfer</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ivana Kvapilíková</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ondrˇej Bojar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University, Faculty of Mathematics and Physics Institute of Formal and Applied Linguistics Malostranské námeˇstí 25</institution>
          ,
          <addr-line>118 00 Prague</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The Covid-19 pandemic has created a global demand for accurate and up-to-date information which often originates in English and needs to be translated. To train a machine translation system for such a narrow topic, we leverage in-domain training data in other languages both from related and unrelated language families. We experiment with different transfer learning schedules and observe that transferring via more than one auxiliary language brings the most improvement. We compare the performance with joint multilingual training and report superior results of the transfer learning approach.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>A global crisis such as the current Covid-19 pandemic
requires information to be spread as efficiently as possible.
Working with data from different international resources
in multiple languages can resolve possible inconsistencies
and prevent misinformation. In an emergency situation,
new data is released constantly and is communicated to the
public not only via national news and authorities, but also
foreign media, scientific journals or statements of
international agencies. There are extensive data resources written
in English which are not accessible for non-English
speakers.</p>
      <p>In order to quickly access the information in a foreign
language, machine translation (MT) can be of great help.
However, Covid-related texts use a specific terminology
and MT models are known to struggle outside of the
general domain.</p>
      <p>More than a year after the Covid outbreak, there
already is a significant amount of domain-specific
multilingual text resources. Furthermore, Covid-related texts
are a part of a broader medical domain which can
provide additional authentic data for training. Thanks to the
MLIA @ Eval1 initiative who gathered training data for
MT and information retrieval related to the pandemic, we
can successfully adapt an MT system to the Covid domain
or even train it from scratch.</p>
      <p>This paper gives an overview of possible methods to
automatically translate Covid-related texts. Section 2
outlines different approaches to train a domain-specific MT
system using multilingual corpora. Section 3 describes
our training data and section 4 gives more details about
our MT systems and presents the results. Section 5
concludes the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>In this section we outline several strategies applicable in
the situation where we need to translate from English to
multiple languages, we are confined within a specific
domain and we have mid-size parallel corpora for every
language pair of interest.</p>
      <p>Firstly, we can train a standard MT system from scratch
for each language separately, possibly resorting to some
data augmentation method, e.g. back-translation.
Secondly, we can use transfer learning to transfer from a
pretrained MT system in one language to another. Finally, we
can train a multilingual MT system which learns jointly
from all available data.</p>
      <p>In our experiments, we do not consider any additional
monolingual resources. Although monolingual data are
generally easier to obtain, we remain constrained by the
datasets provided by the organizers of MLIA @ Eval
which are described in Section 3. We also do not evaluate
a transfer from a large MT model pretrained on texts from
the general domain which would be a promising strategy
as well.
2.1</p>
      <sec id="sec-2-1">
        <title>Low-resource Neural Machine Translation</title>
        <p>
          When neural machine translation (NMT) became the
dominant paradigm in MT [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ], it was believed that extremely
large parallel resources are required for training. However,
Sennrich and Zhang [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] showed that with careful tuning
of the hyperparameters, an NMT model can be
successfully trained already on 100k sentence pairs, which is less
than we have available in the Covid/medical domain for
the language pairs of our interest. Furthermore, Conneau
and Lample [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] show that translation quality can be
further boosted by pretraining a language model and using
it to initialize the parameters of both the encoder and the
decoder.
        </p>
        <p>
          An NMT system directly trained to translate in the
Covid domain serves as our baseline.
Back-translation is a crucial method in NMT used to
augment training data by translating an existing monolingual
corpus [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. The synthetic text can be either on the source
[
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] or the target [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] side of the training corpus, or both
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>
          When using a bidirectional model (sharing the
encoder and decoder for both translation directions),
backtranslation can be performed on-the-fly. During training,
the model switches between the training and the inference
mode to produce batches of synthetic sentence pairs and
learn from both authentic and synthetic samples in each
training step. As the system improves, the quality of the
generated samples improves as well. This approach was
originally proposed for training an unsupervised MT
system [
          <xref ref-type="bibr" rid="ref14 ref2">2, 14</xref>
          ].
        </p>
        <p>In our systems, we translate only the target sentences
and generate a synthetic source side on the fly. We do not
use any additional monolingual data for back-translation.
2.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Transfer Learning</title>
        <p>
          The first strategy to utilize multilingual in-domain training
corpora is transfer learning. It can be used to transfer from
a different domain [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] or a different language [
          <xref ref-type="bibr" rid="ref12 ref22">12, 22</xref>
          ]. In
this work we focus on the latter.
        </p>
        <p>
          A trivial transfer learning approach was proposed by
Kocmi and Bojar [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] who fine-tune a low-resource child
model from a high-resource parent model pretrained for a
different language pair. The training procedure consists of
first training an NMT model on the parent parallel corpus
until it converges and then replacing the training data with
the child corpus.
        </p>
        <p>
          Before training the parent model, it is necessary to
designate some vocabulary entries for the new language.
Otherwise the model would be forced to completely re-learn
its subword embeddings and their connections and would
lose its ability to transfer. Kocmi [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] shows that the best
strategy is to generate the vocabulary in advance from the
concatenation of corpora of both the child and the
parent language pair. However, if the child language is not
known prior to training the parent, it is enough to leave
some "free" slots in the vocabulary and later fill them in
with the vocabulary of the child language.
        </p>
        <p>In this work, we experiment with several transfer
learning schedules. We repeat the transfer procedure several
times with the child becoming the parent for either a
completely new language (e.g. German ! English ! Spanish
! . . . ) or for the original parent (e.g. German ! English
! German ! . . . ), as illustrated in Figure 1. We always
generate the vocabulary from the concatenation of the
parent and its "first child". When adding a third (or fourth)
language, the joint BPE vocabulary has to be modified by
replacing the original parent vocabulary entries with the
new child ones. The schedules and their results are
described in Section 4.
2.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Multilingual Training</title>
        <p>The second strategy to utilize multilingual in-domain
training corpora is joint multililngual training.</p>
        <p>
          Multilingual translation systems are either trained with
full parameter sharing [
          <xref ref-type="bibr" rid="ref1 ref7 ref9">1, 7, 9</xref>
          ], with language-specific
encoders and decoders relying on shared attention [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] or an
attention bridge [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. The results show that multilingual
models yield comparable or even superior results to the
standard bilingual setup.
        </p>
        <p>In this work, we rely on full parameter sharing and
use the same architecture as our bilingual systems, while
training it to translate from English into three languages
(French, Italian and Spanish) at once. During inference,
the target language is determined from indicated language
embeddings of the target sentence. We selected these three
languages for their similarity which could help the model
re-use and share some knowledge. The BPE vocabulary
was extracted from the concatenation of all four corpora,
using only unique English sentences to reach a comparable
corpus size.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>Covid-19 MLIA @ Eval organized a community
evaluation effort aimed at accelerating the creation of
resources and tools for improved Multilingual Information
Access (MLIA). A part of this initiative is a competition
to develop the best MT system translating from English
to several European languages: German, Modern Greek,
French, Italian, Spanish and Swedish. The competition is
incremental and so far only the first round was concluded.</p>
      <p>The parallel training data provided by the organizers for
the first round and used in this paper is summarized in
Table 1.2 It was created based on existing corpora from the
medical domain, enriched with sentences directly about
2The development test set used for the final model selection was
obtained by cutting 500 sentences off of either the train set or the
development set, depending on the original development set size.</p>
      <p>English → German
dev
BLEU</p>
      <p>EN → DE</p>
      <p>20.76</p>
      <p>English → Italian
dev
BLEU</p>
      <p>EN → IT
30.97</p>
      <p>EN → ES
EN → DE</p>
      <p>21.26
EN → ES
EN → IT
31.68</p>
      <p>EN → DE
EN → ES
EN → DE</p>
      <p>22.60
EN → DE
EN → ES
EN → IT
32.10</p>
      <p>EN → DE
EN → SV
EN → DE</p>
      <p>22.55
EN → DE
EN → ES
EN → IT
EN → ES
EN → IT
33.07</p>
      <p>EN → DE
EN → SV
EN → DE
EN → SV
EN → DE
22.50</p>
      <p>
        Train
Fine-tune
Fine-tune
Fine-tune
Fine-tune
Train
Fine-tune
Fine-tune
Fine-tune
Fine-tune
Covid, mostly harvested through web crawling and
parallel sentence mining [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The sentences in different
languages might be similar, but the entire corpus collection is
not multi-parallel.
      </p>
      <p>
        All data was segmented into BPE units [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] with a
vocabulary of 30k items for the training.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experiments &amp; Results</title>
      <p>
        We participated in the MT shared task of the Covid-19
MLIA @ Eval initiative and trained a model for translation
into each of the six languages listed in Section 3. The
results of our submitted systems are summarized in the
preliminary report [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the overall results are discussed in the
shared task findings [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our English ! German and
English ! Swedish systems ranked first (tied with one other
system), our other models ranked second.
      </p>
      <p>We experimented with three training strategies
compared against one baseline BASE:
1. unidirectional training without back-translation
(BASE);
2. bidirectional training with online back-translation
(BT);
3. transfer learning (TRANSFER);
4. multilingual training (MULTILING).</p>
      <p>
        For all our MT models we use a 6-layer Transformer
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] architecture with 8 heads, embedding dimension of
1024 and GELU [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] activations. The training is performed
using the XLM3 toolkit. The translation models were
trained on 4 GPUs4 with 2-step gradient accumulation to
reach an effective batch size of 8 3400 tokens.
Effective batch size has a significant impact on the training
and we observe that the models converge on lower BLEU
scores for smaller batch sizes. We used Adam [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
optimizer with inverse square root decay (b1 = 0:9, b2 = 0:98,
lr = 0:0001). Beam search with the beam size of 4 was
used during final decoding; greedy decoding was used for
back-translation. The vocabulary size was set to 30k.
Using larger vocabulary leads to a performance drop. All our
model parameters are initialized with a pretrained masked
language model as described in Conneau and Lample [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4.1
      </p>
      <sec id="sec-4-1">
        <title>Online Back-Translation</title>
        <p>For each language pair we first trained a bidirectional
back-translation model described in Section 2 and
compared it to a standard unidirectional model without
backtranslation. Online back-translation improved the score by
0.5–2.9 BLEU points, depending on the language, but
surprisingly caused a decrease of 0.4 BLEU in the case of the
English–Modern Greek model. The reason for this drop is
likely in the bidirectionality of the model rather than the
data augmentation itself. The results are summarized in
Table 2.</p>
        <p>
          We experimented with a dropout of 0.1 and 0.2 and
concluded that higher dropout helps in most settings. This
observation is in line with Sennrich and Zhang [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] who
emphasize the role of higher dropout when working with
low- to medium- sized resources.
3https://github.com/facebookresearch/XLM
4Quadro P5000, 16GB of RAM
Transfer Combination
en-es ! en-de
en-de ! en-es
en-de ! en-es ! en-de
en-de ! en-es ! en-fr
en-de ! en-es ! en-it
en-de ! en-es ! en-it ! en-es
en-de ! en-es ! en-it ! en-es ! en-it
en-es ! en-fr
en-es ! en-it
en-de ! en-el
en-es ! en-el
en-de ! en-sv
en-de ! en-sv ! en-de
en-de ! en-sv ! en-de ! en-sv
en-de ! en-sv ! en-de ! en-sv ! en-de
We used the best-performing BASE / BT models as the
parent models and continued with unidirectional training
(English ! foreign language) for our transfer learning
experiments. Since the fine-tuning is unidirectional, we can
no longer perform online back-translation.
        </p>
        <p>We observed that it often helped to use the transfer
incrementally, having the model converge on one
parallel corpus, switch the target language, wait for
convergence and switch again. We hypothesize that the model
benefits from seeing a larger variety of sentences. For
example transferring from German to Spanish to Italian
(32.10 BLEU) performs better than transferring directly
from Spanish to Italian (31.68 BLEU). The best
combination is to even repeat the Spanish-Italian transfer twice
(33.07 BLEU).</p>
        <p>When translating from English to German,
finetuning the en-de BT model on English!Spanish (or
English!Swedish) and switching back to English !
German adds around 1 BLEU on top of the original BT model.
All language combinations used in our transfer learning
multilingual
best transfer
best base
multilingual
best transfer
best base
experiments are described in Table 3 and selected
schedules are illustrated in Figure 1.</p>
        <p>
          We observe that transfer learning improves the
performance in all cases but French, where the BASE model with
BT reaches 38.5 BLEU, which is 3 BLEU points more
than transfer learning. There is a significant overlap
between the training sets in different languages and it is
possible that French does not benefit from the transfer because
it does not provide enough new sentences. On the other
hand, the largest improvement is seen by the language pair
with the least amount of training data, English-Swedish,
where BLEU increases by 1.1 points on the dev set and
1.6 points on the test set.
We train a multilingual model for translation from English
to French, Italian and Spanish. The model has the same
architecture as our bilingual models, all parameters are
shared for all languages. Its encoder and decoder were
first pretrained on monolingual data in all three languages
and English using the MLM criterion [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>Table 4 shows the comparison of the TRANSFER
models with a multilingual model trained jointly. We observe
that transfer learning yields superior results and is thus a
more effective way to leverage multilingual data than joint
multilingual training. However, there is an advantage of a
joint model in terms of the training and storage cost. After
three days of training, the multilingual model can be used
for translation into all three languages. The initial BASE
models can take between one (without BT) and five (with
BT) days to train and fine-tuning on a child language pair
adds around 6 hours.</p>
        <p>Table 5 lists our task submissions and compares all
approaches on the official Covid-19 MLIA @ Eval blind test
set.5
5The BLEU scores in Table 4 and Table 5 cannot be directly
com</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We trained several MT systems specialized in translation
of texts related to the topic of Covid-19 and the pandemic
from English to six European languages.</p>
      <p>We experimented with three training approaches and we
conclude that there is not a universal winner that would
work the best for all language pairs. However, transfer
learning brings promising results across the board,
especially when training data is limited. We observed an
interesting phenomenon where incremental fine-tuning on
multiple languages brings additional gains, as we expose the
model to a larger variety of training sentences.</p>
      <p>In our setting, transferring knowledge is a more efficient
way to leverage multilingual data than joint training. For
English!German, we observe that a transfer learning
detour via Spanish or Swedish improves the parent model
itself. For English!Modern Greek, transfer learning via
German works well, despite the unrelatedness of the two
languages. For English!French, on the other hand, a
bidirectional model with back-translation beats both
multilingual and transfer-based models.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This study was supported in parts by the grants
CZ.07.1.02/0.0/0.0/16_023/0000108 (Operational
Programme – Growth Pole of the Czech Republic),
19-26934X of the Czech Science Foundation, and by the
SVV project number 260 575.
pared as the dev scores were calculated by authors and test scores by the
organizers.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Aharoni</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Johnson,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Firat</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          :
          <article-title>Massively multilingual neural machine translation</article-title>
          .
          <source>In: Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pp.
          <fpage>3874</fpage>
          -
          <lpage>3884</lpage>
          , Association for Computational Linguistics, Minneapolis,
          <source>Minnesota (Jun</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Artetxe</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Labaka</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agirre</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Unsupervised neural machine translation</article-title>
          .
          <source>In: Proceedings of the Sixth International Conference on Learning Representations (April</source>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Casacuberta</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceausu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deligiannis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Domingo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Martinez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herranz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Papavassiliou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piperidis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prokopidis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , Roussis.,
          <string-name>
            <surname>D.</surname>
          </string-name>
          :
          <article-title>The Covid19 MLIA @ Eval Initiative: Overview of the machine translation task (</article-title>
          <year>2021</year>
          ), URL http: //eval.covid19-mlia.eu/meetings/round1/ report/20210112-task3-overview.pdf
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Conneau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lample</surname>
          </string-name>
          , G.:
          <article-title>Cross-lingual language model pretraining</article-title>
          . In: Wallach,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Beygelzimer</surname>
          </string-name>
          , A.,
          <string-name>
            <surname>d'</surname>
            Alché-Buc,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garnett</surname>
            ,
            <given-names>R</given-names>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          <volume>32</volume>
          , pp.
          <fpage>7059</fpage>
          -
          <lpage>7069</lpage>
          , Curran Associates, Inc. (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Firat</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
          </string-name>
          , Y.:
          <article-title>Multi-way, multilingual neural machine translation with a shared attention mechanism</article-title>
          .
          <source>In: Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pp.
          <fpage>866</fpage>
          -
          <lpage>875</lpage>
          , Association for Computational Linguistics, San Diego, California (Jun
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Freitag</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Onaizan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Fast domain adaptation for neural machine translation</article-title>
          .
          <source>CoRR abs/1612</source>
          .06897 (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ha</surname>
            ,
            <given-names>T.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niehues</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waibel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Toward multilingual neural machine translation with universal encoder and decoder (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Hendrycks</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gimpel</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bridging nonlinearities and stochastic regularizers with gaussian error linear units</article-title>
          .
          <source>CoRR abs/1606</source>
          .08415 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krikun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thorat</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viégas</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wattenberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hughes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Google's multilingual neural machine translation system: Enabling zero-shot translation</article-title>
          .
          <source>Transactions of the Association for Computational Linguistic</source>
          <volume>5</volume>
          ,
          <fpage>339</fpage>
          -
          <lpage>351</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ba</surname>
          </string-name>
          , J.:
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>In: Proceedings of the 3rd International Conference for Learning Representations</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Kocmi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Exploring Benefits of Transfer Learning in Neural Machine Translation</article-title>
          .
          <source>Ph.D. thesis</source>
          , Charles University (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Kocmi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bojar</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Trivial transfer learning for low-resource neural machine translation</article-title>
          .
          <source>In: Proceedings of the Third Conference on Machine Translation: Research Papers</source>
          , pp.
          <fpage>244</fpage>
          -
          <lpage>252</lpage>
          , Association for Computational Linguistics, Brussels (Oct
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Kvapilíková</surname>
            ,
            <given-names>I.:</given-names>
          </string-name>
          <article-title>CUNI machine translation systems for the Covid-</article-title>
          19
          <source>MLIA initiative</source>
          (
          <year>2021</year>
          ), URL http://eval.covid19-mlia.eu/meetings/ round1/report/20210114-cunimt.pdf
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Lample</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Denoyer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ranzato</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Unsupervised machine translation using monolingual corpora only</article-title>
          .
          <source>In: 6th International Conference on Learning Representations (ICLR</source>
          <year>2018</year>
          )
          <article-title>(</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Niu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Denkowski</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carpuat</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Bi-directional neural machine translation with synthetic parallel data</article-title>
          .
          <source>In: Proceedings of the 2nd Workshop on Neural Machine Translation and Generation</source>
          , pp.
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          , Association for Computational Linguistics, Melbourne,
          <source>Australia (Jul</source>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Sennrich</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haddow</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Improving neural machine translation models with monolingual data</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the ACL (Volume 1: Long Papers)</source>
          , pp.
          <fpage>86</fpage>
          -
          <lpage>96</lpage>
          , Association for Computational Linguistics, Berlin, Germany (Aug
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Sennrich</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haddow</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birch</surname>
            ,
            <given-names>A.:</given-names>
          </string-name>
          <article-title>Neural machine translation of rare words with subword units</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the ACL</source>
          , pp.
          <fpage>1715</fpage>
          -
          <lpage>1725</lpage>
          , Association for Computational Linguistics, Berlin (Aug
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Sennrich</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Revisiting low-resource neural machine translation: A case study</article-title>
          .
          <source>In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</source>
          , pp.
          <fpage>211</fpage>
          -
          <lpage>221</lpage>
          , Association for Computational Linguistics, Florence,
          <source>Italy (Jul</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
          </string-name>
          , Ł.,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          . In: Guyon,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.V.</given-names>
            ,
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Vishwanathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Garnett</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pp.
          <fpage>6000</fpage>
          -
          <lpage>6010</lpage>
          , Curran Associates, Inc. (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Vázquez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raganato</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tiedemann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Creutz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Multilingual NMT with a language-independent attention bridge</article-title>
          .
          <source>In: Proceedings of the 4th Workshop on Representation Learning for NLP (RepL4NLP-2019)</source>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>39</lpage>
          , Association for Computational Linguistics, Florence,
          <source>Italy (Aug</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xia</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Liu, T.Y.:
          <article-title>Exploiting monolingual data at scale for neural machine translation</article-title>
          .
          <source>In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP)</source>
          , pp.
          <fpage>4207</fpage>
          -
          <lpage>4216</lpage>
          , Association for Computational Linguistics, Hong Kong,
          <source>China (Nov</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Zoph</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuret</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>May</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Transfer learning for low-resource neural machine translation</article-title>
          .
          <source>In: Proceedings of the 2016 Conference on EMNLP</source>
          , pp.
          <fpage>1568</fpage>
          -
          <lpage>1575</lpage>
          , Association for Computational Linguistics, Austin, Texas (Nov
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>