<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UNIMIB @ DIACR-Ita: Aligning Distributional Embeddings with a Compass for Semantic Change Detection in the Italian Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Federico Belotti</string-name>
          <email>f.belotti8@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federico Bianchi</string-name>
          <email>f.bianchi@unibocconi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Palmonari</string-name>
          <email>matteo.palmonari@unimib.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bocconi University</institution>
          ,
          <addr-line>Via Sarfatti 25, 20136, Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <addr-line>Viale Sarca 336, 20126, Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we present our results related to the EVALITA 2020 challenge, DIACR-Ita, for semantic change detection for the Italian language. Our approach is based on measuring the semantic distance across time-specific word vectors generated with Compass-aligned Distributional Embeddings (CADE). We first generate temporal embeddings with CADE, a strategy to align word embeddings that are specific for each time period; the quality of this alignment is the main asset of our proposal. We then measure the semantic shift of each word, combining two different semantic shift measures. Eventually, we classify a word meaning as changed or not changed by defining a threshold over the semantic distance across time.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Semantic change detection is the task of detecting
if a word has shifted in meaning between different
periods of time
        <xref ref-type="bibr" rid="ref10 ref7">(Tahmasebi et al., 2018; Kutuzov
et al., 2018)</xref>
        . The DIACR-Ita
        <xref ref-type="bibr" rid="ref1 ref2">(Basile et al., 2020a)</xref>
        challenge (at EVALITA
        <xref ref-type="bibr" rid="ref1 ref2">(Basile et al., 2020b)</xref>
        ) is
meant to evaluate approaches for semantic change
detection for the Italian Language.
      </p>
      <p>The task is described as follows: for training,
two corpora t1 and t2, consisting of text coming
from different periods are given, for testing, a set
of unlabeled target words is given, where for each
of them a binary scores has to be predicted: 1
identifies lexical change between t1 and t2 while 0
does not.</p>
      <p>In this paper, we present our approach to
semantic change detection that is based on two
components: 1) an alignment procedure to generate
distributional vector spaces that are comparable for
t1 and t2 and 2) the use of distance metrics to
compute the degree of semantic change for a given
word. Our alignment procedure is based on
Compass Aligned Distributional Embeddings (CADE)
proposed by Bianchi et al. (2020) (note the
approach was introduced as Temporal Word
Embeddings with a Compass by Di Carlo et al. (2019),
but the name was changed to enforce the idea that
the embeddings can be used to align more general
corpora and not just diachronic ones). Given the
aligned embeddings, we use two measures to
compute the degree of change based on the similarities
of the vectors in the embedded space. Our results
show that our methodology for aligning spaces can
be useful in detecting lexical semantic change.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of the System: Semantic</title>
    </sec>
    <sec id="sec-3">
      <title>Change Detection with Compass</title>
    </sec>
    <sec id="sec-4">
      <title>Aligned Embeddings</title>
      <p>
        Our approach is based on measuring the
semantic distance across time of time-specific word
vectors generated with CADE and on the use of two
measures for detecting semantic shifts i.e., the
semantic distance between word vectors across time.
This distance can be interpreted as a function of
the words’ self-similarity across time, where the
similarity is measured by a linear combination of
cosine and second-order similarity
        <xref ref-type="bibr" rid="ref5 ref6">(Hamilton et
al., 2016a)</xref>
        .
      </p>
      <p>Finally, a threshold over this self-similarity is
used to classify a word as changed or not changed.</p>
      <p>
        This methodology was applied also in the
semantic shift detection challenge presented at
SemEval2020
        <xref ref-type="bibr" rid="ref9">(Schlechtweg et al., 2020)</xref>
        (to which
we participated after the end of the challenge).
The challenge allowed us to explore and
understand how the alignment and our self-similarity
behaved. In the classification task of the
SemEval2020 challenge (the one similar to this task),
we eventually achieved 0.703, 0.771, 0.725, 0.742,
in accuracy for respectively the English, German,
Latin and Swedish languages; these results have
been obtained with extensive parameter search
given the gold standard available in the
postevaluation.1 In DIACR-Ita, the threshold and few
other hyper parameters can be heuristically set to
account for the limited number of possible
submissions. In the next subsections we provide more
details about the alignment methodology and the
similarity function; more details about how we set
the hyper parameters are provided in Section 3.
2.1
      </p>
      <sec id="sec-4-1">
        <title>Aligning Embeddings</title>
        <p>
          Word2vec
          <xref ref-type="bibr" rid="ref8">(Mikolov et al., 2013)</xref>
          is a useful
methodology to generate vectors of words
allowing us to study word similarity through vector
similarity. However, due to the stochasticity of the
training procedure, running word2vec on
different corpora creates word vectors that are not
comparable. Thus, an alignment procedure that puts
the temporal word vectors in the same space is
needed.
        </p>
        <p>
          There are different approaches to generate these
aligned embeddings (see for example the work by
          <xref ref-type="bibr" rid="ref5 ref6">(Hamilton et al., 2016b)</xref>
          and
          <xref ref-type="bibr" rid="ref11">(Yao et al., 2018)</xref>
          ).
In this paper, we generate aligned embeddings
with Compass Aligned Distributional Embeddings
(CADE)
          <xref ref-type="bibr" rid="ref3">(Bianchi et al., 2020)</xref>
          (See Figure 1 for a
schematic description of the model). CADE is a
strategy to align word embeddings that are specific
for each time period that extends the word2vec
Continuous Bag Of Word (CBOW) model
proposed by Mikolov et al. (2013). CADE can be
used to generate aligned temporal word
embeddings (i.e., time-specific vectors of words, like
“amazon1974”) from the different slices.
        </p>
        <p>Given in input a set of slices of text, where each
slice corresponds to text coming from a specific
period of time, the alignment procedure is as
follows:</p>
        <p>First, the text from all the slices is concatenated
and CBOW is run on this corpus in order to
obtain a “compass” model, i.e., a model defining the
embedding space. The CBOW model uses two
matrices to generate the embeddings (U and C in
Figure 1), one for the context words and one for
the target words. The target word matrix of the
compass is then used to initialize the target
matri1Check the belerico entry in the challenge
leaderboard at https://competitions.codalab.org/
competitions/20948#results
ces for each new CBOW model fitted on each of
the slices. During training, these new target
matrices are frozen, i.e., they are not updated during the
training on the slice. This ensures that at the end of
the training process, the various temporal
embeddings are all aligned in the same embedding space,
making them comparable without losing their
individual temporal distinctions. We use the
publicly available online implementation of CADE.2
2.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Computing Semantic Change</title>
        <p>Once the embeddings are aligned, we need
measures to evaluate the degree of semantic change.
We compute the semantic shift of each word,
i.e. the semantic distance between word vectors
across time using the combination of two
different measures: Local Neighbors (ln), introduced by
Hamilton et al. (2016a) and cosine similarity (cos),
merging them with a weighted linear combination
into a new measure called Move.</p>
        <p>Local Neighbors ln is based on the similarity
between a word and its neighbor words in the two
different time periods. Essentially we compute
the degree of semantic change of the word w in
two slices by first collecting the nearest neighbors
(NNs) of wt and wt+1 in the two respective slices,
then given the embeddings at time t the
similarities between the vector of wt and the vectors of all
the neighbors are computed.3 The same process is
run for time t + 1 with wt+1, eventually giving us
two vectors of similarity scores. These two vectors
are again compared using cosine similarity. The
higher the value of this measure the less the vector
has changed with respect to its neighbors and thus
the less the word should have shifted in meaning.
Cosine Similarity The second measure we use
is simply the cosine similarity of the vectors of a
word in two different time periods. Similarly as
before , the higher the value the less the vector has
changed and thus the less the word should have
shifted in meaning.</p>
        <p>The Move Measure We merge these measures
together using a weighted linear combination, that
is:
s(wt; wt+1) = (1</p>
        <p>) ln(wt; wt+1)
+</p>
        <p>cos-sim(wt; wt+1)
2http://github.com/vinid/cade
3When a neighbor is missing in one time slice, we replace
it with the average vector of the space.</p>
        <p>Di
Ci</p>
        <p>Dn</p>
        <p>Cn</p>
        <p>U
1) Train the
compass from the
concatenation</p>
        <p>C</p>
        <p>U
2) Initialize and freeze
each CBOW target matrix
with the same U matrix
3) training
3) training
3) training
3) training
3) training
U</p>
        <p>U</p>
        <p>U</p>
        <p>U
with 2 [0; 1]. In particular express the usage
strength of the two measures: a high will shift
Move towards the cosine similarity, while a low
one towards the ln measure. As introduced before
we classify if the meaning has changed by
defining a threshold over s (more details about this are
presented in the next Section).
3</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Evaluation</title>
      <p>
        The dataset provided by the challenge’s organizers
        <xref ref-type="bibr" rid="ref1 ref2">(Basile et al., 2020a)</xref>
        is a collection of documents
extracted by newspapers written in the Italian
language labeled with temporal information.
Participants must train their models only on the data
provided, so a pre-processed corpus is given: tab
separated, with one token per line, where for each
token there are its corresponding part-of-speech
(POS) tag and lemma, with sentences separated by
empty lines. The corpus is split into two slices,
each belonging to a specific period of time, t1 and
t2, where t1 &lt; t2.
3.1
      </p>
      <sec id="sec-5-1">
        <title>Dataset</title>
        <p>
          For the training data we used the flat version
with only the lemmas, obtained by the
organizers’ script
          <xref ref-type="bibr" rid="ref1 ref2">(Basile et al., 2020a)</xref>
          ; in addition we
applied a pre-processing step, in which we removed
punctuation and non alpha-numeric symbols and
we kept only those sentences with at least two
tokens.
3.2
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Models Considered</title>
        <p>We use the embeddings aligned with CADE and
the move measure. The parameters of the moving
average we need to consider are: the number of
nearest neighbors (NNs) to be collected by ln,
for the moving average and the threshold for the
similarity. We set the threshold to decide if a word
is stable or not is set to 0.7, with the decision given
by:
(0 if s(wt; wt+1)</p>
        <p>0:7</p>
        <sec id="sec-5-2-1">
          <title>1 otherwise</title>
          <p>Essentially, the less changed are the two
vectors of the words (for cos) and the neighbors (for
ln) the more the word has been stable between
the two time periods. As heuristics we chose
2 f0:3; 0:5; 0:7g to evaluate the relationship
between the two measures used to build move, and
we set to 22 the number of nearest neighbors to
be considered by the ln; this is the general setup
that gave the results that have been submitted to
the challenge.</p>
          <p>We trained CADE for 10 epochs to learn
100dimensional vectors, with the window size set to 5,
10 negative examples for every positive one, with
the initial learning rate set to 0.025 and decreased
linearly during training.</p>
          <p>As other models, in the post evaluation we also
considered one that only uses the cos (CADE
(cos)) similarity measure and one that uses only
the ln metric CADE (ln)) (again with 0.7 as
threshold and with the number of NNs for ln set to 22).</p>
          <p>
            As baselines, the authors propose to use
baseline-freq, that is the absolute value of the
difference between the words’ frequencies and
baseline-colloc, where the Bag-of-Collocations of
the two words in the two different periods is built
and then cosine similarity is applied. A
threshold is used on both metrics to define semantic
change
            <xref ref-type="bibr" rid="ref1 ref2">(Basile et al., 2020a)</xref>
            . We report also the
results of the other participants.
team1
team2
team4
team5
team6
team7
team8
team3
CADE (move)y
team9
baseline-colloc
baseline-freq
          </p>
        </sec>
        <sec id="sec-5-2-2">
          <title>CADE (move)y</title>
          <p>CADE (move)y</p>
        </sec>
        <sec id="sec-5-2-3">
          <title>CADE (cos)</title>
          <p>CADE (ln)
/
/
/
0.3
/
/
/
/
/
/
/
/
The evaluation metric used in this challenge is the
accuracy, that is, the number of correct predictions
over the target data. Table 1 shows the results. Our
model was the third most accurate. However, in
the post-evaluation we discovered that just using
the ln metric and ignoring the use of cos (this is
equivalent to using = 0 in our move measure)
improves the performance leading to the second
best accuracy score in the leaderboard.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>
        Our results show that CADE
        <xref ref-type="bibr" rid="ref3">(Bianchi et al., 2020)</xref>
        is an effective method to generate aligned
embeddings for the Italian language. This result,
together with those obtained on the SemEval2020
data, suggest that CADE can support models of
semantic shift detection in several languages.
Indeed, we show that in combination with some
simple semantic change measures it is possible to
provide a good model for semantic change detection
that can be subsequently extended with more
features. Appendix A contains some more detailed
examples of the words that CADE (ln) and CADE
(move), with lambda set to 0.3, could not
classify correctly. Also, we show the neighborhood
for some of those words to give more context on
why we get those errors. A more precise use of
pre-processing techniques with the combination of
other metrics to compute semantic change might
help in reducing these errors.
      </p>
      <p>A</p>
    </sec>
    <sec id="sec-7">
      <title>CADE Misclassifications</title>
      <p>We report in Tables 2 and 3 CADE’s
misclassifications with the two best metrics, namely CADE
(move) with = 0.3 and CADE (ln). Eventually,
we also show in Tables 4 and 5 some examples of
neighborhood for the target words.</p>
      <p>Wrong predictions done by CADE</p>
      <p>= 0.3.</p>
      <p>Table 4 shows the top 10 nearest neighbors of
the target word “pacchetto” and we think CADE
classifies its meaning as changed because during
time t1 the meaning is more focused in the
economic area, as one can see from neighbors like
“azionario”, “obbligazione” or “contante”
(translated to “stock” as referred to the market, “bond”
and “cash” resp.); while at time t2 shifts to a more
political sense, as shown by words such as
“decreto” or “emendamento” (“decree” and
“amendment” resp.).
azionario
obbligazione
azionista
azionano
edison
casseforte
contante
siap
shell
prestire
maxiemendamento
finanziaria
decretone
decreto
ddl
emendamento
liberalizzazioni
decretere
maxidecreto
ecobonus</p>
      <p>The same it seems to happen for the target word
“piovra”, as one can see from Table 5, where at
time t1 CADE gathers senses from both
considering it as the animal, for example from the word
“tentacle”, or as someone tied to crime in
general, given words such as “profittatore” or
“ruberia” (“profiteer” and “robbery” resp.); while at
time t2 captures a shift towards the Italian crime
TV series “La piovra”, as emerge from words such
as “fiction”, “camorra” or “retequattro”, which is
an Italian television channel.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          , Annalina Caputo, Tommaso Caselli, Pierluigi Cassotti, and
          <string-name>
            <given-names>Rossella</given-names>
            <surname>Varvara</surname>
          </string-name>
          .
          <year>2020a</year>
          .
          <article-title>DIACR-Ita @ EVALITA2020: Overview of the EVALITA2020 Diachronic Lexical Semantics (DIACR-Ita) Task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          . 2020b.
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Federico</given-names>
            <surname>Bianchi</surname>
          </string-name>
          , Valerio Di Carlo, Paolo Nicoli, and
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Palmonari</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Compass-aligned distributional embeddings for studying semantic differences across corpora</article-title>
          . arXiv preprint arXiv:
          <year>2004</year>
          .06519.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Carlo</surname>
          </string-name>
          , Federico Bianchi, and
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Palmonari</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Training temporal word embeddings with a compass</article-title>
          .
          <source>In Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          , pages
          <fpage>6326</fpage>
          -
          <lpage>6334</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>William L. Hamilton</surname>
            , Jure Leskovec, and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          . 2016a.
          <article-title>Cultural shift or linguistic drift? comparing two computational measures of semantic change</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>2116</fpage>
          -
          <lpage>2121</lpage>
          , Austin, Texas, November. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>William L. Hamilton</surname>
            , Jure Leskovec, and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          . 2016b.
          <article-title>Diachronic word embeddings reveal statistical laws of semantic change</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>1489</fpage>
          -
          <lpage>1501</lpage>
          , Berlin, Germany, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Andrey</given-names>
            <surname>Kutuzov</surname>
          </string-name>
          , Lilja Øvrelid, Terrence Szymanski, and
          <string-name>
            <given-names>Erik</given-names>
            <surname>Velldal</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Diachronic word embeddings and semantic shifts: a survey</article-title>
          .
          <source>In Proceedings of the 27th International Conference on Computational Linguistics</source>
          , pages
          <fpage>1384</fpage>
          -
          <lpage>1397</lpage>
          ,
          <string-name>
            <given-names>Santa</given-names>
            <surname>Fe</surname>
          </string-name>
          , New Mexico, USA,
          <year>August</year>
          . Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barbara</surname>
            <given-names>McGillivray</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Simon</given-names>
            <surname>Hengchen</surname>
          </string-name>
          , Haim Dubossarsky, and
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Semeval-2020 task 1: Unsupervised lexical semantic change detection</article-title>
          . arXiv preprint arXiv:
          <year>2007</year>
          .11464.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          , Lars Borin, and
          <string-name>
            <given-names>Adam</given-names>
            <surname>Jatowt</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Survey of computational approaches to lexical semantic change</article-title>
          . arXiv preprint arXiv:
          <year>1811</year>
          .06278.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Zijun</given-names>
            <surname>Yao</surname>
          </string-name>
          , Yifan Sun, Weicong Ding,
          <string-name>
            <given-names>Nikhil</given-names>
            <surname>Rao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Hui</given-names>
            <surname>Xiong</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Dynamic word embeddings for evolving semantic discovery</article-title>
          .
          <source>In Proceedings of the eleventh acm international conference on web search and data mining</source>
          , pages
          <fpage>673</fpage>
          -
          <lpage>681</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>