<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Language
Proceedings of the Eighth Evaluation Campaign of acquisition and conceptual development</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>LG at WiC-ITA: Exploring the relation between semantic distance and equivalence in translation.</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lorenzo Gregori</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <volume>3</volume>
      <issue>2001</issue>
      <fpage>406</fpage>
      <lpage>411</lpage>
      <abstract>
        <p>The Word in Context task has been addressed here on the basis of a simple intuition: if the same lemma has diferent senses in two diferent contexts, it tends to be translated (in other languages) with two diferent lemmas; conversely, in the case of the same sense, translation lemmas tend to be the same. The proposed methodology is based on the translation of sentence pairs in 21 languages, and the use of a SVM classifier/regressor. Obtained results are excellent in binary classification and average in regression tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;EVALITA</kwd>
        <kwd>Word In Context</kwd>
        <kwd>Machine Translation</kwd>
        <kwd>Semantics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>languages, by studying the semantic variation of general
verbs “cut” and “break”. Their work showed how
lanThe Word-in-Context (WiC) task is a new task [1, 2], that guages use their own verbs to partition diferently the
aims to identify if the same word used in two contexts semantic space related to these two events. IMAGACT1
has the same sense in both contexts or if it’s used with highlights the relation between diferent semantic types
diferent senses. Interestingly, in the current EVALITA of an action verb, and diferent groups of verbs usable
task [3, 4] sentences are manually judged by several an- to predicate them. From this resource, we can clearly
notators. observe that, in general, the verb set allowed for one</p>
      <p>The method chosen to solve this task is completely dif- type is not the same set allowed for another type, even if
ferent from the previously used methods, mostly based the two types belong to the same verb. This is a shared
on the use of pre-trained language models: see recent property across languages: it occurs in Italian, English,
works by Alan Ansell and colleagues, and Qianchu Liu and in many other languages.
and colleagues, among others [5, 6]. Actually, in the The proposed approach is based on high-quality
transproposed approach language models are used not to di- lation of the provided sentences in several languages;
rectly accomplish the given task, but to provide accurate translated sentences are then word-aligned to the
origsentence translations in several languages. inal ones, and lemmatized. With this data, it’s easy to</p>
      <p>The intuition behind the proposed methodology is that verify if the lemma used to translate the target word in
a word with the same sense in two contexts is probably the two original sentences is the same or not. This binary
translated with the same lemma in a target language; oth- information repeated for each language is used to
comerwise, a word that is used in two diferent senses tends pile a feature vector. Then, a Support Vector Machines
to be translated with two diferent lemmas. This idea has (SVM) classifier and an SVM regressor are trained on
been first explored by Gale, Church, and Yarowsky in the these vectors to decide if (or how much) the two word
early 1990s [7, 8], and is a basic concept behind the work senses are diferent.
on semantic spaces produced by Melissa Bowerman and
colleagues in the early 2000s [9, 10], and the IMAGACT
Ontology of Action [11, 12]. 2. Word semantics in translation</p>
      <p>Gale and colleagues showed that having a text
translated into another language can be useful for word sense
disambiguation, given that a word with two diferent
senses is frequently translated with two diferent words
in the target language. Bowerman and colleagues
analyzed the variation of event categorization in diferent
Translating a word to another language is not a trivial
task, because it’s pretty rare to find a word in the target
language with exactly the same meaning as the original
word. More likely, there are several possible translators,
each one suitable for some contexts and not for others.</p>
      <p>Moreover, even at the lower level of word senses, it is
hard to find perfect matches between two languages,
given that languages partition semantic spaces in their
own way [13, 14].
• English uses complex in (1) and set in (2).
(1) = (2)
(1) ̸= (2)
() = ()
 =?
 =?
() ̸= ()
 =?
 =?</p>
      <sec id="sec-1-1">
        <title>These premises, plus the fact that the same word can have diferent meanings, and diferent words can be synonyms, make the picture of semantics in translation very complex.</title>
        <p>Consider two occurrences of the same word (in
diferent contexts),  and , translated to another language
with two words, 1 and 2. Then,  and  can have
the same sense or diferent senses; 1 and 2 can be the
same word or diferent words, and in both cases, they can
have the same sense or two diferent senses (they can be
synonyms if they are diferent word with the same sense,
or polysemous if they are the same word with diferent
senses).</p>
        <p>So, the 4 cases represented in Table 1 (() means the
sense of the word ) are all possible, both if 1 and 2
are a single word, or if they are diferent words. It is the
probability of these cases that is undetermined ( =?).</p>
        <p>The assumption behind this experiment is that the
following two cases are more frequent than others:</p>
      </sec>
      <sec id="sec-1-2">
        <title>The adoption of two diferent lemmas to translate these</title>
        <p>Table 1 occurrences of “complesso” is spread over several
lanThe set of 21 translation languages used to solve the WiC task. guages, but, of course, there are some exceptions. In
Bulgarian, for example, in both of the sentences,
“complesso” is translated with a unique word “комплекс”,
similar to Italian.</p>
        <p>Dealing with translations, it’s not possible to rely on a
perfect word-to-word alignment: it can occur that
multiple words are translated with only one word in the target
language, or, vice-versa, a unique word translated with
more words, or some words are just omitted in
translation. Moreover, some alignment errors must be expected
by the word aligner. In this example, the word in (1) is
translated into Lithuanian with “kompleksas”, while in
(2) there is not any word aligned to “complesso”. It could
be due to the linguistic properties of Lithuanian or to an
error in the word alignment.</p>
        <sec id="sec-1-2-1">
          <title>Example 2. The target word has the same sense.</title>
          <p>The following two sentences have been labeled as 1 (i.e.</p>
          <p>target word belonging to the same sense):
1. Chi ha intenzioni meno serie, troverà godibile il The system proposed to solve the WiC-ITA task is divided
tour fotografico del complesso , completato dalla into two parts: (a) the creation of a feature vector related
visita virtuale del museo... to each sentence pair, and (b) the training of a machine
2. ...la Camera dei deputati aveva approvato un com- learning algorithm on the feature vectors.
plesso di disposizioni leggermente diverse da quelle
recepite dalla n. 180.</p>
          <p>3.1. Feature vectors</p>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>In fact, in most languages two diferent lemmas have been used in translation:</title>
        <p>• French uses complexe in (1) and ensemble in
(2);
• Finnish uses kompleksi in (1) and joukko in (2);
Given two sentences 1 and 2, containing both an
occurrence of the same lemma; these occurrences are the
target words 1 and 2. For each of the 21 languages
() considered (see Table 2 for the full list), the algorithm
performs the following steps:
• (() = ()) ∧ ((1) = (2)) if 1 and 2
are the same word;
• (() ̸= ()) ∧ ((1) ̸= (2)) if 1 and 2
are two diferent words.</p>
      </sec>
      <sec id="sec-1-4">
        <title>If this is true for most sentences in most languages,</title>
        <p>then the use of many translation languages provides
reliable information to identify semantic distance between
 and .</p>
        <p>The two examples reported below are derived from the
test set and aim to clarify the idea behind the proposed
methods, and the practical issues.</p>
        <sec id="sec-1-4-1">
          <title>Example 1. The target word has two diferent senses.</title>
          <p>The following two sentences have been labeled as 0 (i.e.
target word belonging to diferent senses):
1. inserendo in questa maschera la parola greca
“mache” (battaglia) si otterranno tutti i termini
collegati...
2. Questa maschera consente di visualizzare alcune
informazioni in forma sintetica: il titolo del
documento,...</p>
        </sec>
      </sec>
      <sec id="sec-1-5">
        <title>As in the previous example, most languages used the same lemma to translate “maschera” in these two contexts:</title>
        <p>• Lithuanian used kauk ė for both (1) and (2);
• Malay used topeng for both (1) and (2);
• Spanish used máscara for both (1) and (2).</p>
      </sec>
      <sec id="sec-1-6">
        <title>Icelandic used two diferent words: sjónarhóll for (1), and gríma for (2).</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Description of the system</title>
      <p>1. Translate 1 in language  → 1;
2. Translate 2 in language  → 2;
3. Align at word-level 1 to 1 → 1;
4. Align at word-level 2 to 2 → 2;
5. Lemmatize 1 and find the lemma of 1 aligned
to 1 → 1;
6. Lemmatize 2 and find the lemma of 2 aligned
to 2 → 2;
7. Assign 0.5 if 1 or 2 are non-words;
Assign 0 if 1 = 2;</p>
      <p>Assign 1 if 1 ̸= 2;</p>
      <sec id="sec-2-1">
        <title>The result of this algorithm is a feature vector  with</title>
        <p>a number {0, 0.5, 1} for each language. The size of 
is the number of languages used (21 in the proposed
experiment).</p>
        <p>As an example, consider the following sentences pair,
belonging to the test set:
translation system must be available in several languages.
To this aim, Opus-MT system [15, 16] has been selected:
it is a state-of-the-art system of neural machine
translation, for which more than 1,400 pretrained models are
freely available2: each model is specifically trained on a
language pair.</p>
        <p>To align sentences and translations at the word level,
the multilingual word aligner created by Dou &amp; Neubig
has been selected [17]: the pretrained model aligned on
multilingual BERT is freely available online3.</p>
        <p>The lemmatizer used for this work is Simplemma4, a
tool available as a Python library that performs sentence
lemmatization in over 50 languages.</p>
        <p>Both the aligner and the lemmatizer have been chosen
to be multilingual ready-to-use tools, that can be easily
included in the implemented pipeline.
3.3. Training the machine
Support Vector Machines algorithms have been used on
the current dataset both for regression and classification.</p>
        <p>Algorithms have been trained on the Italian dataset
and tested on the Italian dataset only.
• ... che vi sia alla base un accordo tra i coniugi,
soprattutto in relazione all’educazione del minore.
• Infatti ogni religione istituzionalizzata annette
importanza maggiore o minore alla propagazione ...
dei suoi riti.</p>
        <sec id="sec-2-1-1">
          <title>Binary classification task After the conversion of</title>
          <p>each pair of input sentences in a vector, a classifier was
trained on the training set and tested on the development</p>
          <p>After the algorithm execution, the  vector appears as set. An SVM classifier with a linear kernel was chosen for
below (only the first 10 elements are reported). its good performance on this task. The complexity
hyperparameter () has been tuned on the dev set, obtaining
     ℎ   ℎ  the best results with  = 0.1.
( 1 1 1 0.5 1 1 0 1 1 1 . . . ) The dataset has been balanced with a random
undersampling technique, to train the algorithm with an equal
Most of the values are 1, meaning that many languages number of examples of positive and negative classes (i.e.
translated the word minore with diferent words; there the sentences where the target word has the same sense,
is a 0.5 in 4th position, meaning that there is a missing and the ones where the target word has diferent senses).
alignment of the word minore with a corresponding In- As reported in Table 3, the training set for the
classificadonesian (id) word (in at least one of the two sentences); tion task is highly unbalanced, with a proportion between
the value of 0 (7th position) means that in Spanish (es) the two classes of 29% - 71%.
minore has been translated with the same word in both
the contexts. Ranking task
3.2. External tools
The key point of this algorithm is the sentence transla- 2https://huggingface.co/Helsinki-NLP
tion engine, which needs to have a high accuracy, and 3https://github.com/neulab/awesome-align
ensure a really contextualized translation; moreover, the 4https://pypi.org/project/simplemma/
The ranking task has been solved with an SVM regressor,
using a Gaussian kernel, that led to better performance</p>
          <p>Results of the Italian task are resumed in Table 5, where
the highest accuracy for each team is reported. The
results obtained in the classification task (on the test set) is
an accuracy of 0.73, which is very high for a WiC
competition [2]. Conversely, this algorithm didn’t emerge
in the ranking task, obtaining an average score of 0.49,
which is below the baseline.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Discussion</title>
      <sec id="sec-3-1">
        <title>The high accuracy obtained in the classification task with</title>
        <p>only features related to lemma equivalences in translation
is, first of all, a piece of strong evidence to support some
than the linear kernel5. linguistic theory about semantic spaces.</p>
        <p>A preliminary analysis of the training set for the rank- The implemented model is easy to interpret, and quick
ing task highlighted that data are highly unbalanced; to train. This makes the proposed system, a very
flexmoreover, the score values assigned to each pair of sen- ible tool to perform other experiments, like changing
tences are exactly 6: {1, 1.5, 2, 3, 3.5, 4 } (see Table 4). the number and the type of languages used, finding the</p>
        <p>It is possible to identify two sources of biases: (a) 71% of minimal set of languages with the maximum
discriminainstances have a high score (3 to 4), while only 29% has a tive power, and so on. It would be also interesting to try
low score (1 to 2); there is an increasing number of scores, to apply on another language the model trained in one
moving form central scores (2 and 3) to the extreme scores language, as suggested by the task organizers.
(1 and 4). This suggested performing a random under- The proposed algorithm can be improved:
sampling to balance the sentences assigned to each score
and train the algorithm in a reduced unbiased space.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>The first issue that emerged in the training was about the</title>
        <p>number of languages that should be used to obtain higher
accuracy. So, at first, the algorithm has been tested using
an increasing number of languages (1 to 21).</p>
        <p>The graph reported in Fig. 1 shows the change in
accuracy of the binary classification algorithm with respect
to the number of languages used for training. This graph
is based on the development set.</p>
        <p>In general, we can see that with just one language the
algorithm accuracy is about 0.60, moving up towards
0.70 as the number of languages increases. The choice
of 21 languages seems reasonable, considering the trend
of the curve, which starts flattening with more than 10
languages.
5Complexity hyper-parameter  is tuned to 0.1, as in the
classification problem
• The number of languages is probably enough to
reach the maximum accuracy with feature
vectors; otherwise the set of used languages could be
changed, by introducing, for example, some Asian
languages, like Chinese, Japanese, or Korean, that
probably would bring a big contribution to this
task;
• Alignment and lemmatization could be improved,
by using, for each language, a state-of-the-art tool
that is specifically tuned for that language; this
would probably lead to results that are better than
the ones obtained with multi-language tools.</p>
        <p>About the computational cost of the proposed
approach, it is completely moved from the training stage
(which has almost no cost) to the feature extraction: in
fact, the computationally intense stage is the neural
machine translation in 21 languages. This task is also very
slow, but easy to parallelize, using simultaneous
translation engines. Interestingly enough, once the vectors have
been compiled for the dataset in use, all the experiments</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>