<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benyou Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Di Buccio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Massimo Melucci</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Engineering University of Padova</institution>
          ,
          <addr-line>Padova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Semantic change detection task in a relatively low-resource language like Italian is challenging. By using contextualized word embeddings, we formalize the task as a distance metric for two flexible-size sets of vectors. Various distance metrics like average Euclidean Distance, average Canberra distance, Hausdorff distance, as well as Jensen-Shannon divergence between cluster distributions based on K-means clustering and Gaussian mixture model are used. The final prediction is given by an ensemble of top-ranked words based on each distance metric. The proposed method achieved better performance than a frequency and collocation based baselines.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Lexical Semantic Change detection aims at
identifying words that change meaning over time; this
problem is of great interest for NLP,
lexicography, and linguistics. A semantic change detection
task in English, German, Latin, and Swedish was
proposed by Schlechtweg et al. (2020). Recently,
Basile et al. (2020a) organized a lexical
semantic change detection task in Italian called
DIACRIta at EVALITA 2020
        <xref ref-type="bibr" rid="ref1 ref2">(Basile et al., 2020b)</xref>
        . This
technical report describes the methodology
designed and developed by the University of Padova
for the participation to DIACR-Ita.
      </p>
      <p>
        Some previous approaches for semantic change
modelling were based on static word embedding,
where word vectors were trained for each
timestamped corpus and then were aligned, e.g. by
orthogonal projections
        <xref ref-type="bibr" rid="ref8">(Hamilton et al., 2016)</xref>
        ,
vec“Copyright c 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).”
tor initialization
        <xref ref-type="bibr" rid="ref10">(Kim et al., 2014)</xref>
        , and
temporal teferencing
        <xref ref-type="bibr" rid="ref4">(Dubossarsky et al., 2019)</xref>
        . This
work relies on contextualized word embeddings
as the basic word representation component
        <xref ref-type="bibr" rid="ref9">(Hu
et al., 2019)</xref>
        , since they have been shown to be
effective in many NLP tasks including document
classification and question answering. The
methods relying on contextualized word embeddings
performed worse than those based on static word
embedding in Semantic Change detection tasks in
many languages
        <xref ref-type="bibr" rid="ref11 ref11 ref11 ref16 ref16 ref18 ref19 ref19 ref19 ref5 ref6 ref6 ref6">(Kutuzov and Giulianelli, 2020;
Po¨msl and Lyapin, 2020; Schlechtweg et al., 2020;
Vani et al., 2020; Giulianelli et al., 2020;
Giulianelli, 2019)</xref>
        . However, it is our opinion that
the use of contextualized word embeddings for
this task is worth investigating because (1) they
have highly expressive power as demonstrated in
many downstream tasks e.g., document
classification and question answering, and (2) they could
handle fine-grained representations of individual
context at the level of tokens.
      </p>
      <p>By using contextualized word embedding, each
word in a specific sentence is represented as a
vector depending on the neighboring words which
form the context of the word; a word
appearing many times in a corpus is therefore
represented as a set of vectors since one vector
corresponds to each occurrence). In this paper,
semantic change detection is addressed by computing the
distance between two flexible-size sets consisting
of vectors with respect to two time-stamped
corpora. We investigated several distance metrics:
average Euclidean Distance, average Canberra
distance, and Hausdorff distance. Our
methodology also relies on a clustering algorithm (e.g.
Kmeans clustering and Gaussian Mixture Model)
on the joint set and calculates a Jensen–Shannon
divergence between cluster distributions in the
two sub-corpora. We aggregate top-ranked words
based on each distance metric as the final
prediction. The proposed method achieved better
performance than frequency and collocation based
baselines and finally ranked the 8-th among 9
participanting teams.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem definition</title>
      <p>
        Unlike the static word embedding like Word2vec
        <xref ref-type="bibr" rid="ref14">(Mikolov et al., 2013)</xref>
        1, contextualized word
embeddings like ELMO
        <xref ref-type="bibr" rid="ref15">(Peters et al., 2018)</xref>
        and
BERT
        <xref ref-type="bibr" rid="ref3">(Devlin et al., 2018)</xref>
        generate word
representation based on the context of a word which
does in this way not have a unique mapping with
a fixed word vector.
      </p>
      <p>Let us denote a corpus with m sentences as C.
In this paper, C is related to a time span t because
of the task characteristics; however, the corpus can
be tailored to any specific aspect, e.g. a specific
domain such as news or books. For a word wi
appearing in C, its contextualized word
representation in the k-th sentence 2 is denoted by ei(;Ck). The
word representation in the corpus is a set
iC = fei(;C1); ei(;C2);
ei(;Ck);
; ei(;Cm) g
(1)</p>
      <p>To examine whether a word wi exhibits a
semantic change between two corpora C1 (in t1) and
C2 (in t2), we check the difference between two
sets iC1 and iC2 . Let li be a human-annotated
label indicating the semantic change degree; li
usually ranges from 0 to 1, where 1 denotes a full
semantic change. Let D be the dimension of the
word vector. We define the distance metric as a
function</p>
      <p>f : fRDgm; fRDgn ! R:
to obtain a semantic change degree based on the
representation of a word in two corpora denoted
as iC1 ; iC2 . When labels are binary, one may
simply use a threshold on the values of f ( ; ) to
predict the binary label. Let be a function to
generate a binary output, e.g., based on a hand-crafted
threshold. We can predict whether wi exhibits a
semantic change between C1 and C2 as follows
li = (f ( iC1 ; iC2 ))
where li is the predicted binary label.</p>
      <p>In conclusion, in our work the semantic change
detection task is formalized as follows
arg max X
f; wi
(f ( iC1 ; iC2 )) == li
(2)
(3)
(4)
1An overview on word vectors is in Wang et al. (2019).
2If a word appears in a sentence more than once, we take
the average.</p>
      <p>Since this is a closed task, we may not have
enough annotated samples to train a f using
gradient descent. Therefore, a well-selected f will be
crucial.
3
3.1</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <sec id="sec-3-1">
        <title>Contextualized Word Embedding</title>
        <p>Using contextualized word embeddings like
ELMO and BERT has be shown to improve
performance in various downstream tasks due to its
expressive power for words. In this paper, we use a
multilingual-BERT3. Uncased models are adopted
since we assume that semantic change detection is
insensitive to word case. All models are in base
settings with 12 layers, 12 heads, and a hidden
state dimension of 768. Only last-layer output of
BERT is used as word representation.
3.2
3.2.1</p>
      </sec>
      <sec id="sec-3-2">
        <title>Measuring Semantic Change Degree</title>
      </sec>
      <sec id="sec-3-3">
        <title>Distance-based methods</title>
        <p>In this section, we introduce various methods to
calculate the semantic change degree.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Average Geometric Distance. Average Geo</title>
        <p>
          metric Distance (AGD) (also can be seen in
          <xref ref-type="bibr" rid="ref11 ref16 ref19 ref5 ref6">(Kutuzov and Giulianelli, 2020; Giulianelli, 2019)</xref>
          ) is
defined as below:
        </p>
        <p>AGD( iC1 ; iC2 ) =</p>
        <p>1
mn</p>
        <p>
          X
x2 iC1 ;y2 iC2
d(x; y)
The distance function d( ; ) can be the Euclidean
Distance 4, the Canberra distance
          <xref ref-type="bibr" rid="ref12">(Lance and
Williams, 1966)</xref>
          5 or any distance function. In this
paper, we also use the negative cosine similarity as
a normalized distance metric.
        </p>
        <p>
          Hausdorff distance. Hausdorff distance
          <xref ref-type="bibr" rid="ref17">(Rockafellar and Wets, 2009)</xref>
          is denoted as HD in short
and is generally used to measure the distance
between two non-empty sets, namely,
        </p>
        <p>HD( iC1 ; iC2 ) = max( sup
x2 iC1 y2infiC2 jjx
x2 iC2 y2infiC1 jjx
sup
yjj2;
yjj2)
(5)
3https://storage.googleapis.com/bert_
models/2018_11_03/multilingual_L-12_
H-768_A-12.zip.</p>
        <p>4Euclidean Distance : d(x; y) = jjx yjj2
5Canberra distance is a normalized version of the
Manjxi yij
hattan distance, d(x; y) = PiD=1 jxij+jyij</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.2.2 Clustering-based Methods</title>
        <p>
          By clustering the union set between iC1 and iC2
in K clusters/categories, we obtained the
category distributions p; q for iC1 and iC2 ,
respectively. We adopted two commonly used
clustering methods: the K-means clustering method
and the Gaussian Mixture Model method. As
for the distance between distributions, we adopted
the Jensen–Shannon Divergence (JSD), which is
a symmetrized and smoothed version of the
Kullback–Leibler divergence:
where KL(p; q) = PiK=1 pi log pqii .
We took the top-K ranked target words of each
metric and aggregated them for the final
submission. The K was decided when the aggregated
target words reached the half of total words
numbers, since we assumed that the annotated labels
are balanced. See
          <xref ref-type="bibr" rid="ref18">(Schlechtweg et al., 2020)</xref>
          for
detailed discussions about thresholds.
DIACR-Ita is the first task on lexical semantic
change for Italian. DIACR-Ita aims to
automatically detect whether a word semantically change
over time. The task is to detect if a set of words,
called target words, change their meaning across
two periods, t1 and t2, where t1 precedes t2.
Participants are provided with two corpora C1 and C2
(corresponding to t1 and t2, respectively), and a
set of target words. For instance, the meaning of
the word ‘imbarcata’ has changed from t1 to t2;
originally, the word referred to an ‘acrobatic
manoeuvre of aeroplanes’, but it is nowadays used to
refer to the state of being deeply in love
          <xref ref-type="bibr" rid="ref1 ref2">(Basile
et al., 2020a)</xref>
          although the latter meaning is much
less used than the former meaning. The task is
formulated as a closed task, namely, models must
be trained solely on the provided data. The
occurrence about target words is reported in Table 1.
        </p>
        <p>Labels in this task are binary and the task is
considered as a binary classification problem. The
evaluation is based on accuracy:
T; F refers to ‘True’ and ‘False’, P; N refers to
‘positive’ and ‘negative’. For example, T P is the
number of Truly-predicted Positive samples.</p>
        <p>
          The task735680 organizers provided two
baselines: Frequencies: the absolute value of the
difference between the words’ frequencies is
computed; Collocations: for each word, it
computes the cosine similarity between two
Bag-ofCollocations (BoCs) vector representations related
to C1 and C2. In both baseline models, a threshold
is used to predict if the word has changed its
meaning.
In this section, we will provide a bi-dimensional
visualization of word representation to intuitively
understand how the contextualized word vectors
work. For each word, we get all contextualized
word vectors (with a dimension of 768) based on
its context. To visualized word in a 2D plane,
we used a typical dimension reduction algorithm
called T-SNE
          <xref ref-type="bibr" rid="ref13">(Maaten and Hinton, 2008)</xref>
          to
reduce word vectors from 768 to 2. Red and blue
points denote the low dimensional representation
of vectors when considering the two time-stamped
corpora C1 (blue) and C2 (red).
        </p>
        <p>For example, ‘rampante’ and ‘palmare’ are the
predicted positive samples while ‘cappuccio’ and
‘campanello’ are predicted negative samples. As
shown in Figure 1, the predicted
semanticallyshifted words exhibit a clear difference between
red points an blue points with respect to two
timestamped corpora. For the predicted
semanticallyunshifted words (see Figure 2), it looks slightly
indistinguishable.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Limitations</title>
      <p>
        In
        <xref ref-type="bibr" rid="ref18">(Schlechtweg et al., 2020)</xref>
        , semantic
representations are mainly divided to two categories:
average embeddings (‘type embeddings’) and
contextualized embeddings (‘token embeddings’).
Schlechtweg et al. (2020) illustrated the
performance of token-based models are much lower than
type-based embedding models. In this section,
we will discuss some limitations of currently-used
contextualized embedding based methods for
semantic change detection.
      </p>
      <p>
        There are typically two kinds of methods to use
contextualized embeddings for semantic change
detection: embedding-based distance metrics and
clustering-based distance metrics
        <xref ref-type="bibr" rid="ref11 ref18 ref19 ref5 ref6">(Schlechtweg
et al., 2020; Vani et al., 2020; Giulianelli et al.,
2020; Giulianelli, 2019)</xref>
        . The former are directly
calculated on the raw contextualized word
embeddings while the latter are based on the clustering
results of contextualized word embeddings.
5.1
      </p>
      <sec id="sec-4-1">
        <title>Embedding-based Distance Metrics</title>
      </sec>
      <sec id="sec-4-2">
        <title>Can distance metrics distinguish semantic shift</title>
        <p>
          patterns? Many typical patterns of semantic
shifts have been investigated
          <xref ref-type="bibr" rid="ref1 ref14 ref2 ref7">(Grossmann and
Rainer, 2013; Basile et al., 2020a)</xref>
          : 1)
pejoration or amelioration (when word meanings
become more negative or more positive); 2)
broadening or narrowing (when it evolves as a
generalized/extended object or a restricted or specialized
one); 3) adding/deleting a sense; 4) totally shifted.
The patterns of semantic change are multifaceted
and we are questioning that a single distance
metric could precisely distinguish all the above typical
semantic shift patterns.
        </p>
        <p>Normalization. Most of distance metrics are not
normalized except for negative cosine similarity.
Absolute values of unnormalized distance metrics
may differ a lot among individual words; they are
sometimes unexpectedly affected by the number
of samples, leads to that the values of metrics may
not be comparable among words.</p>
        <p>Outliers. Some distance metrics (e.g.,
Hausdorff distance) are sensitive to outliers. For
example, since the calculation of Hausdorff distance is
based on infimum and supremum, an outlier point
may largely affect the final Hausdorff distance. As
seen in Table 3, frequently-appearing words e.g.,
‘campionato’ and ‘unico’ have the highest
Hausdorff distance between C1 and C2, this is
probably biased by the fact that the two words appear
frequently (see Table 1) and therefore likely have
more unexpected outliers.</p>
        <p>Model Fine-tuning. The contextualized word
embedding that is based on pre-trained language
models like BERT achieved much better results
compared to static word embedding with a
twostage training paradigm, where the two stages are
pre-training in language model (e.g., mask
language model) and fine-tuning in downstream tasks
(e.g., classifications). However, in the semantic
change detection task, fine-tuning in downstream
tasks is currently impossible because the
annotated labels are insufficient to this aim; to some
extent, the lack of fine-tuning stage may harm the
performance of the pre-trained language models.
5.2</p>
      </sec>
      <sec id="sec-4-3">
        <title>Clustering-based Distance Metrics</title>
        <p>After clustering, we used the Jensen–Shannon
divergence (JSD) which is affected by the issues
mentioned in Section 5.1 like other distance
metrics. Plus, the clustering algorithm may introduce
some errors of semantic change detection. First,
typical clustering algorithms may not necessarily
converge to an identical clustering result when the
seed centroids are changed. Moreover, the number
of clusters is crucial since the optimal number of
clusters cannot easily be decided before clustering.
This paper formalizes semantic change detection
as a distance metric between two variable-sized
sets of vectors. The final prediction is based on an
ensemble of different distance metrics. The
proposed method outperformed weak frequency and
collocation baselines, but it performed less well
than SOTA baselines. As a future work, this task
may be largely improved via a supervised task
in a unified multi-lingual framework; thus, any
human-annotated labels in other languages could
be used in this task since currently the number of
annotated semantically-shift words in a single
language is limited.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work is supported by the Quantum Access
and Retrieval Theory (QUARTZ) project, which
has received funding from the European Union‘s
Horizon 2020 research and innovation programme
under the Marie Skłodowska-Curie grant
agreement No. 721321.</p>
      <p>A</p>
    </sec>
    <sec id="sec-6">
      <title>Appendix</title>
      <p>word
matematica
dettagliato
sanita`
senatore
istruzione
egemonizzare
lucciola
campanello
trasferibile
brama
polisportiva
palmare
processare
pilotato
cappuccio
pacchetto
ape
unico
discriminatorio
rampante
campionato
tac
piovra</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          , Annalina Caputo, Tommaso Caselli, Pierluigi Cassotti, and
          <string-name>
            <given-names>Rossella</given-names>
            <surname>Varvara</surname>
          </string-name>
          .
          <year>2020a</year>
          .
          <article-title>DIACR-Ita @ EVALITA2020: Overview of the EVALITA 2020 Diachronic Lexical Semantics (DIACR-Ita) Task</article-title>
          . In EVALITA 2020,
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro (Eds.). CEUR.org, Online.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          .
          <year>2020b</year>
          .
          <article-title>EVALITA 2020: Overview of the 7th Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>In Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ), Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro (Eds.). CEUR.org, Online.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pretraining of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Haim</given-names>
            <surname>Dubossarsky</surname>
          </string-name>
          , Simon Hengchen, Nina Tahmasebi, and
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Time-out: Temporal referencing for robust modeling of lexical semantic change</article-title>
          . arXiv preprint arXiv:
          <year>1906</year>
          .
          <volume>01688</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Mario</given-names>
            <surname>Giulianelli</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Lexical semantic change analysis with contextualised word representations</article-title>
          .
          <source>Unpublished master's thesis</source>
          , University of Amsterdam, Amsterdam (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Mario</given-names>
            <surname>Giulianelli</surname>
          </string-name>
          ,
          <source>Marco Del Tredici, and Raquel Ferna´ndez</source>
          .
          <year>2020</year>
          .
          <article-title>Analysing Lexical Semantic Change with Contextualised Word Representations</article-title>
          . arXiv preprint arXiv:
          <year>2004</year>
          .
          <volume>14118</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Maria</given-names>
            <surname>Grossmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>Franz</given-names>
            <surname>Rainer</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>La formazione delle parole in italiano</article-title>
          . Walter de Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>William L Hamilton</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jure Leskovec</surname>
            , and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change</article-title>
          .
          <source>In ACL</source>
          .
          <volume>1489</volume>
          -
          <fpage>1501</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Renfen</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Shen</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Shichen</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Diachronic sense modeling with deep contextualized word embeddings: An ecological view</article-title>
          .
          <source>In ACL</source>
          .
          <volume>3899</volume>
          -
          <fpage>3908</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Yoon</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <surname>Yi-I Chiu</surname>
            , Kentaro Hanaki, Darshan Hegde, and
            <given-names>Slav</given-names>
          </string-name>
          <string-name>
            <surname>Petrov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Temporal Analysis of Language through Neural Language Models</article-title>
          .
          <source>ACL</source>
          <year>2014</year>
          (
          <year>2014</year>
          ),
          <fpage>61</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Andrey</given-names>
            <surname>Kutuzov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mario</given-names>
            <surname>Giulianelli</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>UiO-UvA at SemEval-2020 Task 1: Contextualised Embeddings for Lexical Semantic Change Detection</article-title>
          . arXiv preprint arXiv:
          <year>2005</year>
          .
          <volume>00050</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Godfrey N Lance and William T Williams</surname>
          </string-name>
          .
          <year>1966</year>
          .
          <article-title>Computer programs for hierarchical polythetic classification (“similarity analyses”)</article-title>
          .
          <source>Comput. J. 9</source>
          ,
          <issue>1</issue>
          (
          <year>1966</year>
          ),
          <fpage>60</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Laurens van der Maaten and Geoffrey</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Visualizing data using t-SNE. JMLR 9</article-title>
          ,
          <string-name>
            <surname>Nov</surname>
          </string-name>
          (
          <year>2008</year>
          ),
          <fpage>2579</fpage>
          -
          <lpage>2605</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>arXiv preprint arXiv:1301.3781</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Matthew E Peters</surname>
            , Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
            <given-names>Kenton</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>and Luke</given-names>
          </string-name>
          <string-name>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In NAACL</source>
          .
          <volume>2227</volume>
          -
          <fpage>2237</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>Martin Po¨msl and Roman Lyapin</source>
          .
          <year>2020</year>
          . CIRCE at SemEval
          <article-title>-2020 Task 1: Ensembling ContextFree and Context-Dependent Word Representations</article-title>
          . arXiv preprint arXiv:
          <year>2005</year>
          .
          <volume>06602</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>R Tyrrell</given-names>
            <surname>Rockafellar and Roger J-B Wets</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Variational analysis</article-title>
          . Vol.
          <volume>317</volume>
          . Springer Science &amp; Business Media.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barbara</surname>
            <given-names>McGillivray</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Simon</given-names>
            <surname>Hengchen</surname>
          </string-name>
          , Haim Dubossarsky, and
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          .
          <year>2020</year>
          . SemEval
          <article-title>-2020 Task 1: Unsupervised Lexical Semantic Change Detection</article-title>
          . arXiv preprint arXiv:
          <year>2007</year>
          .
          <volume>11464</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>K</given-names>
            <surname>Vani</surname>
          </string-name>
          , Sandra Mitrovic, Alessandro Antonucci, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Rinaldi</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>SST-BERT at SemEval-2020 Task 1: Semantic Shift Tracing by Clustering in BERT-based Embedding Spaces</article-title>
          . arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>00857</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Benyou</given-names>
            <surname>Wang</surname>
          </string-name>
          , Emanuele Di Buccio, and
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Melucci</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Representing Words in Vector Space and Beyond</article-title>
          .
          <source>In Quantum-Like Models for Information Retrieval and DecisionMaking</source>
          . Springer,
          <fpage>83</fpage>
          -
          <lpage>113</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>