<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Findings of a Machine Translation Shared Task Focused on Covid-19 Related Documents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francisco Casacuberta</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexandru Ceausu</string-name>
          <email>Alexandru.Ceausu@curia.europa.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khalid Choukri</string-name>
          <email>choukri@elda.org</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miltos Deligiannis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel Domingo</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mercedes García-Martínez</string-name>
          <email>Ma@</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Herranz</string-name>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guillaume Jacquet</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vassilis Papavassiliou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stelios Piperidis</string-name>
          <email>spip@athenarc.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Prokopis Prokopidis</string-name>
          <email>prokopis@athenarc.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitris Roussis</string-name>
          <email>dimitris.roussis@athenarc.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marwa Hadj Salah</string-name>
          <email>marwa@elda.org</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Athena Research Center</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Court of Justice of the EU</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>European Commission</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Evaluations and Language resources Distribution Agency</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>PRHLT Research Center, Universitat Politècnica de València</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Pangeanic S. L. - PangeaMT</institution>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>ValgrAI - Valencian Graduate School and Research Network for Artificial Intelligence</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This work presents the results of a machine translation shared task focused on Covid-19 related documents. Nine teams took part in this event, which was divided in two rounds and involved seven different language pairs. Two different scenarios were considered: one in which only the provided data was allowed, and a second one in which the use of external resources was allowed. Overall, the best approaches were based on multilingual models and transfer learning, with an emphasis on the importance of applying a cleaning process to the training data.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Machine Translation</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Shared Task</kwd>
        <kwd>Covid-19</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>consisted of three natural language processing tasks: (1)
information extraction, (2) multilingual semantic search
In emergency situations, the public as well as many other and (3) machine translation.
stakeholders need to aggregate and summarize different In this paper, we focus on the third task about machine
sources of information into a single coherent synopsis translation (MT), which was focused on texts from the
or narrative, complementing different pieces of informa- Covid-19 crisis that shocked the world and for which
tion, resolving possible inconsistencies and preventing there were not many processed text or corpora. The task
misinformation. This should happen across multiple lan- was divided in two rounds. At the end of each round,
guages, sources and levels of linguistic knowledge that participants wrote or updated their report describing their
varies depending on social, cultural or educational factors. system and highlighting which methods and data had been</p>
      <p>As a response to the Covid-19 crisis, the Covid-19 used.</p>
      <p>MLIA initiative1 organized a community evaluation
effort aimed at accelerating the creation of resources and
tools for improving the deployment of automatic systems 2. Task description
focused on Covid-19 related documents. This initiative
30% of children and adults infected with measles can</p>
      <p>develop complications.</p>
      <p>The first dose is given between 10 and 18 months of age
in European countries.</p>
      <p>thorities and public health agencies4, EU agencies and
specific broadcast websites (e.g., voxeurop 5, GlobalVoices6
or Voltairnet7).</p>
      <p>For acquiring domain-specific bilingual corpora, we
Figure 1: Examples of English sentences from Covid-19 used a recent version of ILSP-FC [1], a modular toolkit
related documents. that integrates modules for text normalization, language
identification, document clean-up, text classification,
bilingual document alignment (i.e., identification of pairs of
sively with data provided by the organizers (in- documents that are translations of each other) and
sencluding data from a different language pair, mono- tence alignment. As mentioned above, taking into account
lingual data, etc). The use of basic linguistic tools the emergency situation, a “rapid” approach based on
keysuch as taggers, parsers or morphological analyz- words was adopted for text classification (i.e., keeping
ers or multilingual systems was allowed for this only documents that are strongly related to the current
scenario. worldwide health crisis). Specifically for sentence
align• Unconstrained: systems could be trained using ment, the LASER8 toolkit was used instead of the
intedata not provided by the organizers, from any ex- grated aligner. Then, a battery of criteria was applied on
ternal resource not allowed in the constrained sce- aligned sentences to automatically filter out sentence pairs
nario. with potential alignment or translation issues (e.g., with
score less than a predefined threshold) or of limited use</p>
      <p>Systems were evaluated and compared according to the for training MT systems (e.g., duplicate pairs, identical
scenario to which they belonged. It was mandatory that segments in a pair, etc.) and, thus, generate precision-high
one of the submitted systems per language pair belonged language resources.
to the constrained scenario. Participants were able to take For the second round, we repeated the previous
part in any or all of the language pairs. They used their process—re-crawling several websites of national
authorsystems to translate a test set of unseen sentences in the ities and public health agencies—in order to enrich the
source language. Evaluation consisted on assessing the data that had already been collected. Additionally, we
extranslation quality of the submissions. Different metrics ploited the outcomes of an available infrastructure, namely
were used on each round. Medical Information System (MediSys9), with the
purpose of constructing parallel corpora beneficial for MT
2.1. Data collection [2]. Similarly to the first round’s approach, it could be
seen as an application of implementing a quick response
In the context of the first round of this initiative, we de- of the MT community to the pandemic crisis.
cided to collect an initial collection of parallel corpora MediSys is one of the publicly accessible systems of the
in health and medicine domains from well-known web Europe Media Monitor (EMM) which processes media to
sources and enrich them with identified Covid-19 parallel identify potential public health threats in a fully automated
data. The purpose of following this approach was to im- fashion [3]. Focusing on the current pandemic, a dataset
plement a very quick response of the MT community in of metadata which concerns Covid-19 related news was
an emergency situation, like the current pandemic. made publicly available in RSS/XML format, which
cor</p>
      <p>To this end, we first collected an updated version of the responded to millions of news articles [4]. The dataset
European Medicines Agency (EMEA) corpus2, and ap- was divided into subsets according to the articles’ month
plied new (more robust and efficient) methods for text ex- of publication. First, the metadata were parsed and the
traction from pdf files, sentence splitting, sentence align- URL and language of each article were extracted. Then,
ment and parallel corpus filtering. Moreover, medical- each web page of the targeted languages was fetched and
related multilingual collections which were offered by the its main content was stored in a text file. The generated
Publications Ofcfie of EU 3 were processed in a similar text files were merged to create a single document for
manner, increasing the volume of the “general” subset of each language and each period. Thus, these documents
the training data. constituted the Covid-19 related monolingual corpora and</p>
      <p>The first step of acquiring Covid-19-related data was
the identification of several bilingual websites with such
content. With the aim of constructing datasets that could
be publicly available, we targeted websites of national
au</p>
      <sec id="sec-1-1">
        <title>2https://www.prhlt.upv.es/~mt/</title>
        <p>prokopidis-and-papavassiliou-emea.html.
3https://op.europa.eu/en/home.</p>
      </sec>
      <sec id="sec-1-2">
        <title>4Such a list is available at https://www.ecdc.europa.eu/en/</title>
        <p>COVID-19/national-sources.</p>
        <p>5https://voxeurop.eu/.
6https://globalvoices.org/.
7https://www.voltairenet.org/.
8https://github.com/facebookresearch/LASER.</p>
        <p>9https://jeodpp.jrc.ec.europa.eu/ftp/jrc-opendata/
LANGUAGE-TECHNOLOGY/EMM_collection/2020_MediSys_
Covid19_dataset/.
2.3. Quality assessment
were considered comparable (in pairs), due to their narrow
topic and the fact that they were published in the same
time period. To this end, the LASER toolkit was applied
on each document pair to mine sentence alignments for
each EN-X language pair. Finally, several filtering
methods were adopted (i.e., thresholding the alignment score
by 1.04, removing near de-duplicates, etc) to compile the
ifnal dataset.</p>
      </sec>
      <sec id="sec-1-3">
        <title>Taking into account that the corpora were obtained from</title>
        <p>crawling (see Section 2.1), it is important to assess the
quality of the reference sets. To do so, we selected a
2.2. Corpora subset of the Spanish round 1 corpora and post-edited it
For the first round, we selected the data described in the with the help of a team of professional translators. This
previous section (Section 2.1) and split them into train, subset consisted of the worst 500 segments according to
validation and test. Then, to ensure that the tests were a the alignment probability between source and reference.
good representation of the task and were appropriate for Overall, translators thought that the translations in general
being used for evaluation, we sorted all segments from the are good, but some are very free, adding things that are
initial test according to the alignment probability between not in the source, or they are too literal.
source and target. After that, we filtered them according to To assess the quality of the reference sets, we compared
their number of words: removing those segments whose the reference and its post-edited version using human
source had either less than 0.7 or more than 1.3 times the TER (hTER) [5]. This metric computes the number of
average number of words per sentence from the training errors between a translation hypothesis and its post-edited
set. Finally, we selected the first two thousand segments version (in this case, between the automatic reference and
to construct the final version of the test set for round 1. its post-edited version). Thus, the smallest the value the</p>
        <p>This process was improved for selecting round 2’s cor- highest the quality. We obtained a fairly low hTER value
pora. Given the data used for this round, we computed (18.8), which indicates that the translation quality of the
some statistics and removed the outliers (segments that reference is generally good and, thus, is coherent with the
contained more than 100 words in either its source or translators’ opinion.
target). Then, we split the data into train, validation and
test sets. Since the data came from different sources, we 2.4. Evaluation
wanted to ensure that both the validation and tests sets
were representative enough of the training sets. For this In order to evaluate the participant’s systems, we selected
reason, for each language pair, we computed the represen- the bilingual evaluation understudy (BLEU) [6]—which
tation of each source in the total data (i.e., the number of computes the geometric average of the modified n-gram
segments from this source divided by the total number of precision, multiplied by a brevity factor—as our main
segments). Then, out of the total segments we wanted to metric, using sacreBLEU [7] to compute it ensuring
select for validation and test (4000 for each), we select consistent scores.
that same percentage from each source. Additionally, we selected different alternative
well</p>
        <p>Additionally, to ensure that validation and test did not known MT metrics for each round:
contain low-quality segments (given that the data had been
crawled from the web), we sorted the segments according • Round 1:
to its alignment quality. Finally, we shuffled the selected Character n-gram F-score (ChrF) [8]:
segments and split them equally into validation and test. character n-gram precision and recall</p>
        <p>Therefore, the procedure we followed for each language arithmetically averaged over all character
pair was: n-grams.</p>
        <p>1. We computed the ratio of data from each different</p>
        <p>source over the total data.
2. We computed the average number of words per</p>
        <p>segment over this set.
3. We constituted a subset [0.7 * average words per</p>
        <p>segment, 1.3 * average words per segment].
4. We sorted this subset (from best to worst)
accord</p>
        <p>ing to its alignment score.
5. We selected the best 8000 * the percentage
obtained at step 1 segments.</p>
        <p>• Round 2:</p>
        <sec id="sec-1-3-1">
          <title>Translation Edit Rate (TER) [5]: this metric</title>
          <p>computes the number of word edit
operations (insertion, substitution, deletion and
swapping), normalized by the number of
words in the final translation.</p>
        </sec>
        <sec id="sec-1-3-2">
          <title>BEtter Evaluation as Ranking (BEER) [9]:</title>
          <p>a sentence level metric that incorporates
a large number of features combined in a
linear model.</p>
          <p>We applied approximate randomization testing RNN
(ART) [10]—with 10, 000 repetitions and using a -value
of 0.05—to determine whether two systems presented These systems were trained using the standard parameters
statistically significance. The scripts used for conducting for RNN MT systems: long short-term memory units [16],
the automatic evaluation are publicly available together with all model dimensions set to 512; Adam [17], with a
with some utilities which were useful for the shared ifxed learning rate of 0.0002 and a batch size of 60; label
task10. smoothing of 0.1 [18]; beam search with a beam size of</p>
          <p>Following the WMT criteria [11], we grouped systems 6; and joint byte pair encoding (BPE) [19] applied to all
together into clusters according to the statistical signifi- corpora, using 32, 000 merge operations. In light of the
cance of their performance (as determined by ART). With results, this architecture was only used for the first round.
that purpose, we sorted the submissions according to each
metric and computed the significance of the performance Transformer
between one system and the following. If it was not
significant, we added the second system into the cluster of
the first system 11. Otherwise, we added it into a new
cluster. This way, systems from one cluster significantly
outperformed all others in lower ranking clusters.</p>
          <p>These systems were trained using the standard parameters:
6 layers; Transformer [14], with all dimensions set to 512
except for the hidden transformer feed-forward (which
was set to 2048); 8 heads of Transformer self-attention;
2 batches of words in a sequence to run the generator on
in parallel; a dropout of 0.1; Adam [17], using an Adam
2.5. Baselines beta2 of 0.998, a learning rate of 2 and Noam learning rate
At each round, we trained two different constrained sys- decay with 8000 warm up steps; label smoothing of 0.1
tems to use as baselines in order to have an estimation of [18]; beam search with a beam size of 6; and joint BPE
the expected translation quality of each scenario. The first applied to all corpora, using 32, 000 merge operations.
system was based on recurrent neural network (RNN)
[12, 13] while the other one was based on the Trans- 2.6. Participants’ approaches
former architecture [14]. All systems were built using
OpenNMT-py [15].</p>
        </sec>
      </sec>
      <sec id="sec-1-4">
        <title>In this subsection, we present the different approaches</title>
        <p>submitted by each team.
10Hidden GitHub repository.</p>
        <p>11Considering that, at the start of this process, there is an initial
cluster containing the first system.</p>
      </sec>
      <sec id="sec-1-5">
        <title>This team only participated in round 1, with an approach based on multilingual BART [20].</title>
      </sec>
      <sec id="sec-1-6">
        <title>EMEA Corpus and a health related subset of the Euramis dataset.</title>
        <sec id="sec-1-6-1">
          <title>Lingua Custodia</title>
          <p>CdT-ASL Lingua Custodia’s submissions for round 1 consisted of
a multilingual model able to translate from English to
CdT-ASL team developed NICE which integrates neural French, German, Spanish, Italian and Swedish; and
inmachine translation (NMT) custom engines for confiden- dividual translation models for English–German and
Ential adapted translations. They submitted constrained and glish–French. They applied unigram SentencePiece for
unconstrained systems, they added generic and public subword segmentation using a source and target shared
health domains to internal data for unconstrained systems. vocabulary of 50K for individual models and 70K for
mulThey applied cleaning processes to prepare the data for tilingual models. Additionally, authors split the numbers
training with big transformer using OpenNMT-tf. They character by character. For multilingual models, a
lanonly took part in round 2. guage token is added to the source in order to indicate the
target language. The English–German multilingual model
CUNI-MT achieved much higher score than the English–German
This team took part in both rounds. For the first round, single model. This improvement is not shown in the
Enthey submitted approaches based on standard NMT with glish–French model.
online back-translation [21]; a transfer learning approach For round 2, they participated in the constrained
scebased on Kocmi and Bojar [22]; and multilingual models nario. The pre-processing used was based on Moses’
in which, during inference, the corresponding embedding tokenizer and cleaning techniques such as removing much
of the target language was selected. For the second round, longer sentences comparing source and target lengths,
rethey trained a multilingual model jointly on all languages. placement of consecutive spaces by one space. They used
inline casing consisting of adding a tag with the casing
information. They finally append the language token to
CUNI-MTIR each source sentence in the pre-processing in order to
indicate the target language for multilingual models.
Standard transformer architecture in Sockeye toolkit was used
for training in multiple GPUs instead of Seq2SeqPy used
previously because the data loading is more efficient and
has better support for multiple GPUs.</p>
        </sec>
      </sec>
      <sec id="sec-1-7">
        <title>CUNI-MTIR only took part in round 1, training constrain</title>
        <p>models using the Transformer architecture from
MarianNMT [23], and using the UFAL Medical Corpus12 for
training unconstrained data and then fine-tuning the
models with the constrained data. All the data was tokenized
using Khresmoi13’s tokenizer and, then, encoded using
BPE with 32K merges.</p>
        <sec id="sec-1-7-1">
          <title>LIMSI</title>
        </sec>
      </sec>
      <sec id="sec-1-8">
        <title>LIMSI only took part in round 1, submitting a Transformer</title>
        <p>E-Translation model using BPE with 32K vocabulary units was applied
For round 1, this team used transfer learning and a 12K to the constrained system. They submitted four
unconsize vocabulary created using SentencePiece over Trans- strained systems: 1) one system build using an external
former models trained with MarianNMT [23]. Addition- in-domain biomedical corpora; 2) a system first trained
ally, they submitted to the unconstrained category their on WMT1414 general data and fine-tuned on the shared
WMT system and a new version of that system, fine-tuned task’s corpus; 3) same as 2) but adding BERT [24]; and 4)
with the constrained data. a system only trained with constrained data but computing</p>
        <p>For round 2, they focused on performing a general the BPE codes using all the external in-domain corpus.
clean-up including a language identifier and checking the
match of the number of tokens in source and target to filter PROMT
noisy segments. For Greek and Spanish they did not do
pre- or post-processing, only sanity checking. They
experimented with standard Transformer and big Transformer
in MarianNMT [23]. For the unconstrained scenario, they
made use of the TAUS Corona Crisis Corpora, the OPUS
For round 1, PROMT’s approaches consisted in a
multilingual model trained using MarianNMT’s [23] Transformer
architecture. For the constrained scenario, all data was
concatenated using de-duplication to one single
multilingual corpus to build a 8k SentencePiece [25] model for
subword segmentation. In addition, a language-specific
tag was added to the source side of the parallel sentence
12http://ufal.mff.cuni.cz/ufal_medical_corpus.</p>
        <p>13http://www.khresmoi.eu/assets/Deliverables/WP4/
KhresmoiD412.pdf.
14http://www.statmt.org/wmt14/translation-task.html.
pairs (e.g., &lt;  &gt; token was added to the beginning of cases. This approach also used a smaller vocabulary and
the English sentence of the English–Italian sentence pair). SentencePiece instead of BPE.</p>
        <p>They also removed all tokens that appeared less than ten In general, the differences from one position to the next
times in the combined de-duplicated monolingual corpus one were of a few points (according to both metrics), with
from their vocabulary. a case (English–French) in which there are two points of</p>
        <p>For the unconstrained scenario, all available data difference (according to BLEU) between the first and last
mainly from the OPUS [26] and statmt15 with the addition approaches of the same ranking. Our baselines worked
of private data harvested from the Internet were added to well as delimiters: more sophisticated approaches
generthe training data. A special BPE implementation [27] de- ally ranked above our Transformer baselines, while the
veloped by the team was applied instead of SentencePiece, rest ranked either between them or below the RNN
basebut the authors used SentencePiece in the constrained sce- lines. Moreover, the RNN baselines established the limit
nario as it seemed to work better in low-resource settings. before a significant drop in translation quality between
The size of the BPE models and vocabularies varied from approaches of one position in the ranking and the next
8k to 16k and shared vocabulary was not used (separate position (sometimes it is the exact limit, while other times
BPE models were trained) for the English–Greek pair as there is a cluster above it of a similar quality).
the two languages have different alphabets. Regarding the unconstrained scenario, it had less
par</p>
        <p>For round 2, they trained a transformer multilingual ticipation than the constrained one. With an exemption
model with a single encoder and a single decoder with (E-Translation’s approaches based on their WMT
sysMarian toolkit and performing fine-tuning for each lan- tem [28] yielded the best results for English–German),
guage pair. For the unconstrained scenario, they used the PROMT’s multilingual approach achieved the best results
same approach as in round 1. for all language pairs. In general, approaches were
similar to the constrained ones but using additional external
TARJAMA-AI data. Additionally, due to the use of external data, the best
unconstrained systems yielded around 10 BLEU points
and 7 ChrF points of improvement compared to the best
constrained systems for each language pair.
3.1. Round 1 This work presents a community evaluation effort to
improve the generation of MT systems as a response to a
Table 2 presents the results of the first round. Overall, global problem. The initiative consisted of generating
spemultilingual and transfer learning approaches yielded cialized corpus for a new and important topic: Covid-19.
the best results for all language pairs in the constrained This initiative was divided into two rounds.
scenario. In fact, except for English–German (in which This first round addressed 6 different language pairs
they shared the same ranking), PROMT’s multilingual ap- and was divided into two scenarios: one in which
particproach—which was the only multilingual system trained ipants were limited to using only the provided corpora
for all language pairs—achieved the best results in all (constrained) and another one in which the use of external
tools and data was allowed (unconstrained). 8 different
This team submitted a single system consisting in a model
trained with all the language pairs data adding a special
token for the non-target languages. Additionally, they
over-sampled the corpus of the desired target language
(i.e., the English–Spanish corpus for training the
constrained English–Spanish, etc). They only took part in
round 1.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Results</title>
      <sec id="sec-2-1">
        <title>In this section, we present the results from each round.</title>
        <p>Following the WMT criteria [11], we grouped systems
together into clusters according to which systems
significantly outperformed all others in lower ranking clusters,
according to ART (see Section 2.4). For clarity purposes
and space constrains, we use BLEU as the main metric
for performing the ranking. Nonetheless, we tried using
each metric from Section 2.4 as the main one for ranking,
observing that all of them resulted in similar clusters.
3.2. Round 2
Table 3 presents the results of the second round. With
the exception of English–French, in which monolingual
approaches achieved the best results, multilingual
approaches yielded the best performances. In the case of
English—German, system ensembling also ranked at first
position.</p>
        <p>In general, the differences from one position to the
next one were of a few points (according to all metrics).
Our baselines worked well as delimiters: more
sophisticated approaches generally ranked above our baselines,
while the following cluster after them obtained the highest
quality drop between consecutive ranks.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Conclusions</title>
      <p>English–Arabic
Description
7lang-ov
multilingual-model-round2-tuned-ar
7lang
multilingual-model-round2
transfer
1lang
Transformer
multiling
only-round2-data</p>
      <p>Transformer
BLEU [↑]
25.1
22.9
22.0
21.7
19.1
19.1
18.8
17.0
15.9</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>The Covid-19 MLIA @ Eval initiative has received
support from the European Commission, the European
Language Resources Coordination (ELRC), European
Language Resources Association (ELRA), the European
Research Infrastructure for Language Resources and
Technology (CLARIN), the CLEF Initiative and the Joint
Research Centre (JRC). We gratefully acknowledge the
translation team from Pangeanic for their help with the quality
assessment.</p>
    </sec>
    <sec id="sec-5">
      <title>Author contribution</title>
      <sec id="sec-5-1">
        <title>Khalid Choukri was one of the main organizers of the</title>
        <p>Covid-19 MLIA @ Eval initiative. Francisco
Casacuberta, Miguel Domingo, Mercedes García-Martínez and
Manuel Herranz organized the machine translation shared
task. Alexandru Ceausu, Miltos Deligiannis, Guillaume
Jacquet, Vassilis Papavassiliou, Stelios Piperidis, Prokopis
Prokopidis, Dimitris Roussis and Marwa Hadj Salah
were responsible of the data acquisition and
engineering. Miguel Domingo and Mercedes García-Martínez
wrote the main manuscript text. All authors reviewed the
manuscript and contributed in generating the final version.
teams took part in this round. Among their approaches,
the most successful ones were based on multilingual MT
and transfer learning, using the Transformer architecture.</p>
        <p>The second round addressed 7 different language pairs
and was also divided into constrained and unconstrained
scenarios, and engaged 5 different teams. Overall, a focus
on the cleaning process of the data yielded great
improvements. Most approaches were based on the Transformer
and big Transformer architectures. Once more,
multilingual models achieved great results, showing to be
specially beneficial for languages with less resources. These
results make sense due to the fact that Covid-19 corpora
are needed to specialize models in the domain, even if
they are from another language.
[14] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, OPUS, in: Proceedings of the Eighth International
L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, Conference on Language Resources and Evaluation
Attention is all you need, in: Advances in Neural (LREC’12), 2012, pp. 2214–2218.
Information Processing Systems, 2017, pp. 5998– [27] A. Molchanov, PROMT systems for WMT 2019
6008. shared translation task, in: Proceedings of the
[15] G. Klein, Y. Kim, Y. Deng, J. Senellart, A. M. Rush, Fourth Conference on Machine Translation (Volume
OpenNMT: Open-Source Toolkit for Neural Ma- 2: Shared Task Papers, Day 1), 2019, pp. 302–307.
chine Translation, in: Proceedings of the Associa- [28] C. Oravecz, K. Bontcheva, L. Tihanyi, D.
Kolovrattion for Computational Linguistics: System Demon- nik, B. Bhaskar, A. Lardilleux, S. Klocek, A. Eisele,
stration, 2017, pp. 67–72. etranslation’s submissions to the wmt 2020 news
[16] F. A. Gers, J. Schmidhuber, F. Cummins, Learning translation task, in: Proceedings of the Fifth
Conferto forget: Continual prediction with LSTM, Neural ence on Machine Translation, 2020, pp. 254–261.
computation 12 (2000) 2451–2471.
[17] D. P. Kingma, J. Ba, Adam: A method for
stochastic optimization, arXiv preprint arXiv:1412.6980
(2014).
[18] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed,</p>
        <p>D. Anguelov, D. Erhan, V. Vanhoucke, A.
Rabinovich, Going deeper with convolutions, in:
Proceedings of the IEEE Conference on Computer
Vision and Pattern Recognition, 2015, pp. 1–9.
[19] R. Sennrich, B. Haddow, A. Birch, Neural
machine translation of rare words with subword units,
in: Proceedings of the Annual Meeting of the
Association for Computational Linguistics, 2016, pp.</p>
        <p>1715–1725.
[20] Y. Liu, J. Gu, N. Goyal, X. Li, S. Edunov,</p>
        <p>M. Ghazvininejad, M. Lewis, L. Zettlemoyer,
Multilingual denoising pre-training for neural
machine translation, arXiv preprint arXiv:2001.08210
(2020).
[21] G. Lample, L. Denoyer, M. Ranzato, Unsupervised
machine translation using monolingual corpora only,
CoRR abs/1711.00043 (2017). URL: http://arxiv.</p>
        <p>org/abs/1711.00043. arXiv:1711.00043.
[22] T. Kocmi, O. Bojar, Trivial transfer learning for
lowresource neural machine translation, in: Proceedings
of the Third Conference on Machine Translation,
2018, pp. 244–252.
[23] M. Junczys-Dowmunt, R. Grundkiewicz, T. Dwojak,</p>
        <p>H. Hoang, K. Heafield, T. Neckermann, F. Seide,
U. Germann, A. Fikri Aji, N. Bogoychev, A. F. T.</p>
        <p>Martins, A. Birch, Marian: Fast neural machine
translation in C++, in: Proceedings of ACL 2018,</p>
        <p>System Demonstrations, 2018, pp. 116–121.
[24] J. Zhu, Y. Xia, L. Wu, D. He, T. Qin, W. Zhou,</p>
        <p>H. Li, T.-Y. Liu, Incorporating bert into neural
machine translation, arXiv preprint arXiv:2002.06823
(2020).
[25] T. Kudo, J. Richardson, Sentencepiece: A
simple and language independent subword tokenizer
and detokenizer for neural text processing, CoRR
abs/1808.06226 (2018). URL: http://arxiv.org/abs/
1808.06226. arXiv:1808.06226.
[26] J. Tiedemann, Parallel data, tools and interfaces in</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>