<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>XL-WA: a Gold Evaluation Benchmark for Word Alignment in 14 Language Pairs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>FedericoMartel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrei StefaBnejgu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CesareCampagnan o</string-name>
          <email>campagnano@di.uniroma1</email>
          <xref ref-type="aff" rid="aff6">6</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>JakaČibej</string-name>
          <xref ref-type="aff" rid="aff8">8</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>RuteCosta</string-name>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ApolonijaGanta</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>JelenaKalla</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SvetlaKoeva</string-name>
          <email>vetla@dcl.bas.b</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>KristinaKoppel</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon Krek</string-name>
          <email>simon.krek@ijs.si</email>
          <email>simon.laszlo@nytud.hu</email>
          <xref ref-type="aff" rid="aff5">5</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>MargitLangemet</string-name>
          <email>margit@eki.ee</email>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>VeronikaLipp</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>SanniNimb</string-name>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sussi Olsen</string-name>
          <email>saolsen@hum.ku.dk</email>
          <xref ref-type="aff" rid="aff7">7</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bolette SandfoPrdedersen</string-name>
          <xref ref-type="aff" rid="aff7">7</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ValeriaQuochi</string-name>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Salgad o</string-name>
          <email>anasalgado@fcsh.unl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>LászlóSimon</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CaroleTiberius</string-name>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael-JUreña-Ruiz</string-name>
          <email>rafa@rae.es</email>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>RobertoNavigl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>NOVA CLUNL</string-name>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Portugal</string-name>
          <xref ref-type="aff" rid="aff9">9</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Academia das Ciências de Lisboa</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Babelscape</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Bulgarian Academy of Sciences</institution>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>HUN-REN Hungarian Research Centre for Linguistics</institution>
          ,
          <country country="HU">Hungary</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Institute of the Estonian Language</institution>
          ,
          <country country="EE">Estonia</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Jožef Stefan Institute</institution>
          ,
          <country country="SI">Slovenia</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>Sapienza University of Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff7">
          <label>7</label>
          <institution>University of Copenhagen</institution>
          ,
          <country country="DK">Denmark</country>
        </aff>
        <aff id="aff8">
          <label>8</label>
          <institution>University of Ljubljana</institution>
          ,
          <country country="SI">Slovenia</country>
        </aff>
        <aff id="aff9">
          <label>9</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Word alignment plays a crucial role in several Natural Language Processing tasks, such as lexicon injection and cross-lingual label projection. The evaluation of word alignment systems relies heavily on manually-curated datasets, which are not always available, especially in mid- and low-resource languages. In order to address this limitation, we propose XL-WA, a novel entirely manually-curated evaluation benchmark for word alignment covering 14 language pairs. We illustrate the creation process of our benchmark and compare statistical and neural approaches to word alignment in both language-specific and zero-shot settings, thus investigating the ability of state-of-the-art models to generalize on unseen language pairs. We release our new benchmark ath:ttps://github.com/SapienzaNLP/XL-W. A CLiC-it 2023: 9th Italian Conference on Computational Linguistics, level between parallel senten1c,e2s][. Historically, word apolonija.gantar@guest.arne(As.s.iGantar)j;elena.kallas@eki.ee - also thanks to novel neural approaches - still plays</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR</p>
      <p>ceur-ws.org</p>
      <p>martelli@diag.uniroma1(F.i.tMartellib);ejgu@babelscape.com
CEUR
htp:/ceur-ws.org
ISN1613-073</p>
      <p>Attribution 4.0 International (CC BY 4.0).</p>
      <p>CEUR</p>
      <p>Workshop ProceedingsC(EUR-WS.org)</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <sec id="sec-2-1">
        <title>Word alignment is the computational task of identify</title>
        <p>ing translation correspondences at word and multi-word
alignment played a crucial role in Statistical Machine
Translation3,[4, SMT]. However, while SMT has been
replaced by end-to-end neural architectures which
attain considerably higher performances, word alignment
a crucial role in many other Natural Language
Processing (NLP) tasks, such as lexicon injection and, most
importantly, cross-lingual annotation proje5c]t. iFoonr[
instance, Procopio et a6l]. [recently proposed a
state-ofthe-art approach to cross-lingual label projection based
on word alignment which allows high-quality
sensemore, word alignment has also been leveraged efectively
to create silver datasets, not only for Word Sense Disam- # of Sentences # of Alignments
biguation7[, 8, WSD] but also for other semantic tasks, Lang. Dev Test Dev Test
such as Semantic Role Labelin9g, [10, SRL], thereby
addressing the knowledge acquisition bottlen1e1c]k,e[s- en-ar 90 210 1591 3597
pecially when dealing with mid- and low-resource lan- en-bg 105 245 1719 4179
guages. eenn--edsa 110055 224455 11894611 44712326</p>
        <p>While, on the one hand, current architectures for word en-et 105 245 1614 3722
alignment are achieving increasingly better performance, en-hu 105 245 1580 3781
on the other hand, the lack of high-quality manual data in en-it 103 243 1980 4765
multiple languages significantly limits their potential and en-ko 90 210 1277 3007
scalability. With a view to addressing the aforementioned en-nl 105 245 1886 4490
drawbacks, our contributions are as follows: en-pt 105 245 1849 4578
en-ru 90 210 1114 2582
1. We propose a fully manually-annotated evalua- en-sl 105 245 1942 4537
tion benchmark for word alignment with a to- en-sv 90 210 1522 3530
tal of 14 language pairs, each composed of En- en-zh 90 210 1724 4135
glish and one of the following languages: Ara- ∑ 1393 3253 23600 55761
bic, Bulgarian, Chinese, Danish, Dutch, Estonian,
Hungarian, Italian, Korean, Portuguese, RussiaTnab,le 1
Slovenian, Spanish and Swedish. Composition ofXL-WA. We report from left to right: the
available language combinations, the number of sentences
2. We experiment with statistical and neuralaandp-alignments divided by data split. In our experiments, we
proaches to word alignment and evaluate thuesmeapproximately 30% of our data for development so as to
against our newly created benchmark. obtain a more representative set.
3. We demonstrate that the concatenation of our
novel datasets can be exploited efectively to train
a neural approach that generalizes on unseen lEanng-lish-Swedish4 [26], Chinese-English5 [27].
Interestingly, Graca et al.28[] proposed a collection of small
guages in a zero-shot setting, thereby
address</p>
        <p>datasets for word alignment in 6 language combinations;
ing the lack of training data in low-resource elaacnh- dataset being composed of 100 sentences derived
guages. from the Europarl corp6u[s29]. Among the currently
available resources, we highlight the following
contri2. Related Work butions which we use in our experiments: the
English</p>
        <p>French and Romanian-English corpora released during
Approaches Initial approaches to word alignmetnhte HLT-NAACL-2003 workshop on Building and Using
leveraged statistical and heuristic mo1d2e]l.sA[long Parallel Tex7ts[30], and the German-English data8set
these lines, several systems were proposed such as HMpMroposed by Vilar et al3.1][. Finally, Neubig3[2]
pre[1], GIZA++1 [13], PGIZA++, MGIZA++ [14] and FastAl- sented a Japanese-English data9 soebttained by
transign2 [15]. Subsequently, statistical approaches were grlaadt-ing Wikipedia pages. However, despite the
precedually substituted by neural counterparts and the adinvegnetforts undertaken in this direction, to the best of
of Transformer architectur16e]s s[et a new standard inour knowledge, no entirely manually-curated evaluation
this task1[7, 18, 19, 20, 21]. More recently, Procopio etbenchmark, which matches XL-WA in both size and
lanal. [6] proposed a novel neural discriminative model gfuoarge pairs covered, is currently available.
word alignment based on multilingual BE22R]T, c[apable
of significantly reducing the processing time.</p>
        <p>3. XL-WA
Data Over the course of the last few decades,Toa tackle the aforementioned gap, we introXdLu-cWeA,
number of datasets for word alignment, both maanno-vel entirely manually-curated evaluation benchmark
ual and automatic, have been created, e.g.
CzechEnglish3 [23], Dutch-English 2[4], English-Turkish 2[5], 4https://www.ida.liu.se/divisions/hcs/nlplab/resources/ges/
5https://nlp.csai.tsinghua.edu.cn/~ly/systems/TsinghuaAligner/
TsinghuaAligner.html
6https://www.statmt.org/europarl/
1https://github.com/moses-smt/giza-pp 7https://web.eecs.umich.edu/~mihalcea/wpt/
2https://github.com/clab/fast_align 8https://www-i6.informatik.rwth-aachen.de/goldAlignment/
3https://ufal.mff.cuni.cz/czech-english-manual-word-alignment9http://www.phontron.com/kftt
the
holds its first plenary session in</p>
        <p>El</p>
        <sec id="sec-2-1-1">
          <title>CDR celebra su</title>
          <p>primer pleno en</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Bruselas en marzo de 1994 .</title>
          <p>for word alignmenXtL.-WA is currently composed of We compute such tha t  contains: i1) if, for the
14 datasets, out of which 9 are parallel. The langu a-gtehsEnglish sentence, a translation int o-tthhtearget
included in XL-WA cover 7 diferent language families, i.lea. nguage is available, 0iio)therwise. We first extract the
Afro-Asiatic, Indo-European (Germanic), Indo-Europeasenntences shared in the highest number of languages.
Sub(Romance), Indo-European (Slavic), Sino-Tibetan, Uralsiecquently, we manually discard sentences which are not
(Finno-Ugric) and Koreanic. well-formatted or contain significant grammatical errors.</p>
          <p>Importantly, all datasets include English as sourceWleanth-en ask annotators to provide the missing
translaguage. This choice is motivated by the fact that enabltiniogns in order to fill the gaps in our parallel d1a2t.aset
word alignment from English to multiple target langFuiangaelsly, we ask our annotators to perform word
alignis crucial for tasks such as label projection, wheremtehnet from scratch.
majority of high-quality annotated data whose labels can
be propagated is typically available in English. Guidelines All annotators are required to follow
spe</p>
          <p>We show the composition of our dataset in T1a.blceific annotation guidelines for word alignment inspired
ImportantlXy,L-WA is annotated exclusively by profebsy- Lambert et al3.5[], who provide detailed instructions
sional mother tongue annotators with a solid acadaenmdicsuggestions regarding the annotation of datasets for
background and proven experience in linguistic annowtoar-d alignment, including specific cases and exceptions.
tion tasks. A detailed description of the format whicIhmwpeortantly, annotators are asked to align source and
adopt is provided in Sectio3.n2. target words also when these do not share the same part
of speech. Furthermore, annotators are required to align
3.1. Creation process complex lexical units such as compounds and multi-word
expressions. For instance, given an open compound word
In this section, we detail the creation process and i l lu,es-.g. bus driver in the English source sentence,
transtrate the guidelines adopted during the annotation plhaatseed. into Dutch with the compound w
orbduschauf</p>
          <p>The creation oXfL-WA can be divided into three stepsf:eur, each component o f should be aligned t o.
i) automatic extraction of candidate sentences from a
corpus, ii) manual selection of sentences satisfying specific 3.2. Alignment Format
linguistic criteria, and miiia)nual alignment.</p>
          <p>In order to obtain a balanced corpus in terms ofWdeon-ow describe our alignment format and provide an
mains and genres, similarly to the procedure adopetxeadmple for the language combination English-Spanish
by Martelli et al3.3][, we extract our data from Wik(ie-n-es).</p>
          <p>Matri1x0 [34], a wide-coverage collection of parallel senW-e adopt the Pharaoh alignment form3a6t]. [
tences derived from the Wikipe1d1iacorpus using an Specifically, we use a Tab-Separated Values (TSV)
automatic approach based on multilingual sentencefeomrm-at, where each row is formatted as folsloouwrcs:e
beddings, covering 1620 diferent language combinationss.entence&lt;tab&gt;target sentence&lt;tab&gt;alignments.
First, we consider the WikiMatrix datasets containToinkgens and alignments are separated by spaces; each
English as the source language, and extract the highaelsitgnment is composed of a pair of integers which
number of overlapping source sentences across datasiedtesn. tify the corresponding positions of source and
To this end, we compute a Boolean ma t∈ri{x0, 1} × target tokens, starting from zero. In order to deal with
where is the number of English sentences in WikiMam-ulti-word expressions in which 1:1 alignments are
trix an d the number of the target languages other tnhaont possible, e.g., due to collocations or idiomatic
English covered inXL-WA. expressions, we align all components of a given
10https://ai.facebook.com/blog/wikimatrix/
11https://www.wikipedia.org/
12Due to time constraints and the limited availability of professional
annotators for specific language combinations, we carry out this
step in 9 language combinations only.
3.3. Inter-annotator agreement
multi-word expression in English with all component4s.o1f.1. Language-specific setting
the corresponding multi-word expression in the target</p>
          <p>Systems In this setting, we experiment with two
stalanguage. tistical approaches, namely GIZA++ and FastAlign, and</p>
          <p>Below we report an example extracted fro ment-ehse two state-of-the-art neural models, i.e. the SQuAD-style
dataset: formulation for word align m14e, nwthich relies on
mul• Source: In March 1994 the CoR holds its tilingual BERT, proposed by Nagata et2a0l]. a[nd the
first plenary session in Brussels . MultiMirror neural word aligner by Procopio e6t]. al. [</p>
          <p>For each language pair, the aforementioned statistical
• Target: El CDR celebra su primer pleno systems are trained on a randomly selected sample of
en Bruselas en marzo de 1994 . 0.5M parallel sentences concatenated with our test data.</p>
          <p>Instead, for neural approaches requiring aligned data,
• Alignments: 3-0 4-1 5-2 6-3 7-4 8-5 9-5 which is not available in all our language combinations,
10-6 11-7 0-8 1-9 2-10 2-11 12-12 we follow Garg et a1l.7][. Specifically, we use sentences
A visual representation of the above example is providdeerdived from the aforementioned silver training data,
in Figure1. tagged both with GIZA++ and FastAlign, and randomly
choose 1,000 sentences with the highest number of
overlapping alignments.</p>
          <p>Finally, in order to assess the reliability of our manuDalaatan- For this setting, we derive training data from
notations, we compute the inter-annotator agr1e3e.metnhtree well-established parallel corpora, namely, Europarl,
To this end, we randomly select a sample of approWxiik-iMatrix and UNP1C5 [37]. Importantly, this choice
mately 50 sentence pairs in two language combinatioanllso,ws us to cover all language combinations considered.
namelyen-da and en-it, and ask new annotators tIonstead, for validation and evaluation purposes we use
align these manually. We compute the Cohen’s kaptphaeXL-WA datasets whose composition is reported in
and obtain 0.94 and 0.89 ienn-da anden-it, respectively. Table1. In this case, our goal is to show and analyze the
Importantly, these results indicate a remarkable lepveelrfoofrmance achieved by state-of-the-art models on each
agreement, which suggests a high degree of annotatiloannguage pair.
consistency across datasets.</p>
          <p>4.1.2. Zero-shot setting</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experimental Setup</title>
      <sec id="sec-3-1">
        <title>In the zero-shot setting, we experiment with MultiMirror</title>
        <p>only, since this model shows a reasonable balance
beIn this section, we illustrate our experimental setuptwaneden results and processing speed. Specifically, we train
carry out a performance analysis. To this end, we MpuutltiMirror on the concatenation of our datasets and
forward two diferent experimental settings. Specificalelvya,luate it against unseen language pairs, thus
demonwe propose a comparison between statistical and nsterua-ting the efectiveness oXfL-WA when no aligned data
ral approaches tested against our novel benchmariks ianvailable in a given language combination. In this case,
a language-specific setting, i.e. we train and test on tohuer goal is to determine the extent to which the model
same language pairs (Sectio4n.1.1). Subsequently, we is able to generalize on language pairs unseen during
investigate the behavior of models in a zero-shot setttrinagin,ing, i.ee.n-de, en-fr, en-ja anden-ro. The data is
thus exploring the ability of state-of-the-art modsepllsittoas in Nagata et a2l0.].[
deal with languages unseen during the training phase
(Section4.1.2). Finally, we describe the evaluation metrics
adopted. 4.2. Evaluation metrics</p>
      </sec>
      <sec id="sec-3-2">
        <title>As customary in the word alignment task, we adopt the</title>
        <p>4.1. Settings following evaluation metrics: precision, recall and F1. In
this work, we do not use the Alignment Error Rate (AER)
We now describe our two experimental settings. Techmnei-tric, since previous works argue that AER is unlikely
cal details regarding hyperparameters and hardwarteoarbee a useful metric for word alignment, due to its bias
reported in AppendiAx. towards precision4][.
13Due to time constraints we compute the inter-annotator agree14mhetnttps://github.com/nttcslab-nlp/word_align
in two language combinations. 15https://opus.nlpl.eu/UNPC.php</p>
        <p>Lang. P R F1 Instead, as far as the zero-shot scenario is concerned,
we observe a good generalization capability of
Multien-de 89.4 78.5 83.6 Mirror when trained on the concatenation of our novel
en-fr 94.7 55.7 70.1 datasets and tested against unseen language pairs, as
en-ja 79.4 42.5 55.3 reported in Tabl3e. In particular, the language
combien-ro 86.7 80.7 83.6 nationsen-de and en-ro attain a remarkable 83.6 F1
score. Importantly, this seems to suggest that the
zero</p>
        <p>Avg 87.5 64.4 73.2 shot paradigm can be employed as a viable approach to
Table 3 compensate efectively for the lack of annotated data in
Results of MultiMirror trained onXaLl-lWA datasets and eval- many low-resource languages.
uated on unseen data, i.e., in a zero-shot setting. To facilitate Finally, we investigate the impact of the size of the
analysis and comparison, we keep English as source language.training data generated with GIZA++ and FastAlign,
as described in Sectio4n.1, on the overall performance
achieved by MultiMirr1o6r.To this end, we increase the
5. Results size of the silver training data to 10,000 sentence pairs
and compare the results obtained with those achieved in
In this section, we discuss the results obtained. As ctahne previous setting where we use 1,000 sentence pairs.
be seen in Table2, in the language-specific setting, we As can be seen in Table4, the greater quantity of data
observe a remarkable diference between statistical aanldlows us to achieve better results in terms of both
precineural approaches, with the latter outperforming thseiofnora-nd F1 score. However, interestingly, when training
mer by up to 17.5 points in terms of F1 score on averaogne. 10,000 sentence pairs, MultiMirror reports a slightly
In this setting, the best results are attained by Naingafetraior performance in terms of recall, with a decrease
et al. 2[0] in the English-Dutchen(-nl) combination. of 0.3 on average.</p>
        <p>Interestingly, we note that even neural models struggle
to achieve good results in topic-prominent languages
such as Chinese, Hungarian and Korean. In fact, in these
languages, both statistical and neural approaches o b16tAasimnentioned in Sectio4n.1.2, we use MultiMirror in this
expersignificantly below-average results. iment due to a satisfactory trade-of between performance and
processing speed.</p>
        <p>R</p>
        <p>F1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <sec id="sec-4-1">
        <title>The authors gratefully acknowledge the support of the</title>
        <p>ELEXIS project No. 731015 under the European Union’s
Horizon 2020 research and innovation programme.
Furthermore, the authors are sincerely thankful for the
support of the Estonian Research Council grant (PRG 1978).</p>
        <p>Finally, the authors gratefully acknowledge the support
of the PNRR MUR project PE0000013-FAIR. This work has
been carried out while Andrei Stefan Bejgu was enrolled
in the Italian National Doctorate on Artificial Intelligence
run by Sapienza University of Rome.
Lang.
[1] S. Vogel, H. Ney, C. Tillmann, Hmm-based word
alignment in statistical translation, in: COLING
1996 Volume 2: The 16th International Conference</p>
        <p>Avg 84.8 83.3 83.9 85.7 83.0 84.2 on Computational Linguistics, 1996. URhtLt:ps://
Table 4 aclanthology.org/C96-21.41
Comparison between MultiMirror trained on diferent size [2] J. Tiedemann, Word to word alignment strategies,
silver datasets. P, R and F1 stand for Precision, Recall and in: COLING 2004: Proceedings of the 20th
InterF1-score, respectively; all the scores are calculated using the national Conference on Computational Linguistics,
Micro average. 2004, pp. 212–218. URL: https://aclanthology.org/
C04-1031.</p>
        <p>[3] F. J. Och, C. Tillmann, H. Ney, Improved
align6. Conclusion ment models for statistical machine translation,
in: 1999 Joint SIGDAT Conference on Empirical
In this work, we introducXeL-WA, a novel evaluation Methods in Natural Language Processing and Very
benchmark for word alignment in 14 language pairs. Large Corpora, 1999. URLh:ttps://aclanthology.org/
We detail the creation process for our novel evalua-W99-0604.
tion suite, as well as our experimental setup in whi[c4h] A. Fraser, D. Marcu, Measuring word alignment
we compare statistical and neural approaches to word quality for statistical machine translation,
Compualignment. We investigate the behavior of models in tational Linguistics 33 (2007) 293–303. URhLt:tps:
zero-shot scenarios and show that the concatenation of//aclanthology.org/J07-30.02
our datasets can be used efectively to align language[s5] D. Yarowsky, G. Ngai, Inducing multilingual pos
unseen during training, thus tackling the paucity or taggers and np bracketers via robust projection
limited availability of data for word alignment in lowa-cross aligned corpora, in: Second Meeting of
resource languages. We release our new benchmark at: the North American Chapter of the Association
https://github.com/SapienzaNLP/XL-W. A for Computational Linguistics, 2001. UhRtLt:ps:</p>
        <p>As future work, we intend to investigate the impact //aclanthology.org/N01-10.26
of language-specific peculiarities on the overall perfo[r6]- L. Procopio, E. Barba, F. Martelli, R. Navigli,
Mulmance of neural models for word alignment. Furthermore,timirror: Neural cross-lingual word alignment
we plan to increase the language coveragXeLo-WfA and, for multilingual word sense disambiguation, in:
importantly, investigate the role played by additional lowP-roceedings of the Thirtieth International Joint
resource languages in zero-shot settings. Finally, we aim Conference on Artificial Intelligence, IJCAI-21,
to explore novel neural approaches to word alignment 2021, pp. 3915–3921. URL: https://www.ijcai.org/
which can be employed in the field of cross-lingual label proceedings/2021/0539.pd.f
projection in order to create multilingual silver t[r7a]inE-. Barba, L. Procopio, N. Campolungo, T. Pasini,
ing datasets for several Natural Language UnderstandingR. Navigli, Mulan: Multilingual label
propagatasks, such as WSD, SRL and Semantic Parsing. tion for word sense disambiguation, in:
Proceedings of the Twenty-Ninth International
Conference on International Joint Conferences on
Artiifcial Intelligence, 2021, pp. 3837–3844. URLh:ttps: URL: https://aclanthology.org/D19-1.453
//www.ijcai.org/Proceedings/2020/0531.p.df [18] E. Stengel-Eskin, T.-r. Su, M. Post, B. Van Durme, A
[8] M. Bevilacqua, T. Pasini, A. Raganato, R. Navigli, discriminative neural model for cross-lingual word
Recent trends in word sense disambiguation: A sur- alignment, in: Proceedings of the 2019
Confervey, in: Z. Zhou (Ed.), Proceedings of the Thirtieth ence on Empirical Methods in Natural Language
International Joint Conference on Artificial Intelli-Processing and the 9th International Joint
Congence, IJCAI 2021, Virtual Event / Montreal, Canada, ference on Natural Language Processing
(EMNLP19-27 August 2021, ijcai.org, 2021, pp. 4330–4338. IJCNLP), Association for Computational
LinguisURL: https://doi.org/10.24963/ijcai.2021/5.9d3oi:10. tics, Hong Kong, China, 2019, pp. 910–920. URL:
24963/ijcai.2021/593. https://aclanthology.org/D19-1.0d8o4i:10.18653/
[9] S. Padó, M. Lapata, Cross-lingual annotation v1/D19-1084.</p>
        <p>projection for semantic roles, Journal of [A19r]- T. Zenkel, J. Wuebker, J. DeNero, Adding
intificial Intelligence Research 36 (2009) 307–340. terpretable attention to neural translation
modURL: https://www.jair.org/index.php/jair/article/ els improves word alignment, arXiv preprint
download/10629/2541.6 arXiv:1901.11359 (2019). URL:http://arxiv.org/abs/
[10] A. Daza, A. Frank, X-SRL: A parallel cross- 1901.11359.</p>
        <p>lingual semantic role labeling dataset, in: Proc[e2e0d]- M. Nagata, K. Chousa, M. Nishino, A supervised
ings of the 2020 Conference on Empirical Meth- word alignment method based on cross-language
ods in Natural Language Processing (EMNLP), As- span prediction using multilingual BERT (2020).
sociation for Computational Linguistics, Online, URL: https://aclanthology.org/2020.emnlp-mai.n.41
2020, pp. 3904–3914. URL: https://aclanthology[.21] M. J. Sabet, P. Dufter, F. Yvon, H. Schütze,
Simaorg/2020.emnlp-main.32.1doi:10.18653/v1/2020. lign: High quality word alignments without
paremnlp-main.321. allel training data using static and contextualized
[11] W. A. Gale, K. W. Church, D. Yarowsky, A method embeddings, in: Findings of the Association for
for disambiguating word senses in a large corpus, Computational Linguistics: EMNLP 2020, 2020,
Computers and the Humanities 26 (1992) 415–439. pp. 1627–1643. URL: https://aclanthology.org/2020.</p>
        <p>URL: https://www.jstor.org/stable/30204.634 findings-emnlp.14.7
[12] V. J. Della Pietra, The mathematics of stati[s2t2i]- J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT:
cal machine translation: Parameter estimation,Pre-training of deep bidirectional transformers for
Using Large Corpora (1994) 223. URLh:ttps:// language understanding, in: Proceedings of the
aclanthology.org/J93-2003.p.df 2019 Conference of the North American
Chap[13] F. J. Och, H. Ney, A systematic comparison of vari- ter of the Association for Computational
Linguisous statistical alignment models, Computational lin-tics: Human Language Technologies, Volume 1
guistics 29 (2003) 19–51. URL:https://aclanthology. (Long and Short Papers), Association for
Comorg/J03-1002. putational Linguistics, Minneapolis, Minnesota,
[14] Q. Gao, S. Vogel, Parallel implementations of word 2019, pp. 4171–4186. URL: https://aclanthology.org/
alignment tool, in: Software engineering, testing, N19-1423. doi:10.18653/v1/N19-1423.
and quality assurance for natural language pro[c2e3s]s- D. Mareček, Automatic Alignment of
Tectograming, 2008, pp. 49–57. URL: https://aclanthology.org/ matical Trees from Czech-English Parallel
CorW08-0509. pus, Master’s thesis, Charles University, MFF UK,
[15] C. Dyer, V. Chahuneau, N. A. Smith, A sim- 2008. URL: https://ufal.mff.cuni.cz/pcedt3.0/pubs/
ple, fast, and efective reparameterization of ibm Marecek2008_diplomka.pd.f
model 2, in: Proceedings of the 2013 Confer[2-4] L. Macken, An annotation scheme and gold
ence of the North American Chapter of the As- standard for dutch-english word alignment, in:
sociation for Computational Linguistics: Human 7th conference on International Language
ReLanguage Technologies, 2013, pp. 644–648. URL: sources and Evaluation (LREC 2010), European
https://aclanthology.org/N13-1073.pdf Language Resources Association (ELRA), 2010,
[16] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, pp. 3369–3374. URL: http://www.lrec-conf.org/
L. Jones, A. N. Gomez, Ł. Kaiser, I. Polosukhin, proceedings/lrec2010/pdf/100_Paper.p.df
Attention is all you need, Advances in neur[2a5l] M. T. Cakmak, S. Acar, G. Eryigit, Word
aligninformation processing systems 30 (2017). URL: ment for english-turkish language pair., in: LREC,
http://arxiv.org/abs/1706.037.62 2012, pp. 2177–2180. URL: http://www.lrec-conf.
[17] S. Garg, S. Peitz, U. Nallasamy, M. Paulik, Jointly org/proceedings/lrec2012/pdf/380_Paper.pdf
learning to align and translate with transfo[2r6m] eMr. Holmqvist, L. Ahrenberg, A gold standard for
models, arXiv preprint arXiv:1909.02074 (2019). english-swedish word alignment, in: Proceedings
of the 18th Nordic conference of computational 2004.amta-papers.13./
linguistics (NODALIDA 2011), 2011, pp. 106–113. [37] M. Ziemski, M. Junczys-Dowmunt, B. Pouliquen,
URL: https://aclanthology.org/W11-4615.pdf The united nations parallel corpus v1. 0, in:
Pro[27] Y. Liu, M. Sun, Contrastive unsupervised word ceedings of the Tenth International Conference
alignment with non-local features, in: Proceed- on Language Resources and Evaluation (LREC’16),
ings of the AAAI Conference on Artificial Intelli- 2016, pp. 3530–3534. URL: https://aclanthology.org/
gence, volume 29, 2015. URL:http://arxiv.org/abs/ L16-1561/.</p>
        <p>1410.2082. [38] L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, J. Han,
[28] J. Graca, J. P. Pardal, L. Coheur, D. Caseiro, On the variance of the adaptive learning rate and
Building a golden collection of parallel multi-beyond, in: International Conference on Learning
language word alignment., in: LREC, 2008. URL: Representations, 2019. URLh:ttps://iclr.cc/virtual_
http://www.lrec-conf.org/proceedings/lrec2008/ 2020/poster_rkgz2aEKDr.htm.l
pdf/250_paper.pd.f
[29] P. Koehn, Europarl: A parallel corpus for
statistical machine translation, in: ProcAee.d-Hyperparameters and
ings of machine translation summit x: papers,
2005, pp. 79–86. URL: https://aclanthology.org/2005. Hardware
mtsummit-papers.1.1</p>
        <p>In this appendix, we report the hyperparameters and
[30] R. Mihalcea, T. Pedersen, An evaluation exercise</p>
        <p>hardware setup for the experiments described in the
pafor word alignment, in: Proc. of HLT-NAACL, 2003,
pp. 1–10. URL: https://aclanthology.org/W03-0.301per.
[31] D. Vilar, M. Popović, H. Ney, Aer: Do we need to We adopt four approaches with the following
hyperparameters:
“improve” our alignments?, in: Proc. of Workshop
on Spoken Language Translation, 2006. URhLt:tps: • We use two statistical approaches, namely
//aclanthology.org/2006.iwslt-papers.7..pdf GIZA++ [13] and FastAlign1[5]. We compile the
[32] G. Neubig, The Kyoto free translation task, 2011. code downloaded from the original repositories</p>
        <p>URL: http://www.phontron.com/k.ftt and we run all the experiments on CPU. Neither
[33] F. Martelli, R. Navigli, S. Krek, J. Kallas, P. Gan- of the approaches requires any parameter tuning.
tar, S. Koeva, S. Nimb, B. Sandford Pedersen,
S. Olsen, M. Langemets, K. Koppel, T. Üksik, • SQuAD mBERT-based model2[0], whose code is
J. Dobrovoljc, R.-J. Ureña-Ruiz, J.-L. Sancho- downloaded from the oficial repository. We run
Sánchez, V. Lipp, T. Váradi, A. Győrfy, S. László, all the experiments using the default
hyperparamV. Quochi, M. Monachini, F. Frontini, C. Tiberius, eters. For the sake of consistency and fairness, we
R. Tempelaars, R. Costa, A. Salgado, J. Čibej, do not tune any hyperparameters and use the
opT. Munda, Designing the ELEXIS parallel timal ones according to the authors, as specified
sense-annotated dataset in 10 european lan- in their paper. All the experiments run for 2
trainguages, in: Proceedings of the eLex Conference, ing epochs with a learning rate3o×f10−5 and
2021. URL: https://elex.link/elex2021/wp-content/ a batch size of 6. Language-specific experiments
uploads/2021/08/eLex_2021_22_pp377-395.pd.f run for approximately 20 minutes each. We also
[34] H. Schwenk, V. Chaudhary, S. Sun, H. Gong, experiment with the whole multilingual dataset,
F. Guzmán, Wikimatrix: Mining 135m parallel which requires 4 hours to complete the training.
sentences in 1620 language pairs from wikipedia, Inference for the language-specific experiments
arXiv preprint arXiv:1907.05791 (2019). URhLt: tps: takes around one minute per language on GPU.
//arxiv.org/pdf/1907.05791.p d.f • MultiMirror6][ is an mBERT-based model whose
[35] P. Lambert, A. De Gispert, R. Banchs, J. B. Mariño, code is obtained from the authors for research
Guidelines for word alignment evaluation and man- purposes. All the experiments run with a patience
ual alignment, Language Resources and Evaluation of 50, using the RAdam3[8] optimizer with a
39 (2005) 267–285. URL: https://link.springer.com/ learning rate o1f− 05 and a token batch size
article/10.1007/s10579-005-4822-.5 of 512. Language-specific experiments run for
[36] P. Koehn, Pharaoh: A beam search decoder for approximately 10 to 15 minutes each, while the
phrase-based statistical machine translation mod- multilingual experiment on the whXoLl-eWA
els, in: R. E. Frederking, K. B. Taylor (Eds.), Ma- dataset runs for approximately 1 hour. Inferences
chine Translation: From Real Users to Research, time is negligible: a few seconds on CPU for the
Springer Berlin Heidelberg, Berlin, Heidelberg, language-specific data and around one minute for
2004, pp. 115–124. URL: https://aclanthology.org/ the whole dataset.</p>
        <p>All the experiments are conducted on the same
hardware, i.e. an Intel Core i7 7800x CPU and NVidia RTX
2080ti GPU with 11GB of VRAM.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>