<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Comparative Study of Approaches for the Diachronic Analysis of the Italian Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierluigi Cassotti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierpaolo Basile</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco de Gemmis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Semeraro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of Bari Aldo Moro Via E. Orabona</institution>
          ,
          <addr-line>4 - 70126 Bari</addr-line>
          ,
          <country country="IT">ITALY</country>
        </aff>
      </contrib-group>
      <fpage>130</fpage>
      <lpage>140</lpage>
      <abstract>
        <p>In recent years, there has been a signi cant increase in interest in lexical semantic change detection. Many are the existing approaches, data used, and evaluation strategies to detect semantic drift. Most of those approaches rely on diachronic word embeddings. Some of them are created as post-processing of static word embeddings, while others produce dynamic word embeddings where vectors share the same geometric space for all time slices. The large majority of the methods use English as the target language for the diachronic analysis, while other languages remain under-explored. In this work, we compare state-of-theart approaches in computational historical linguistics to evaluate the pros and cons of each model, and we present the results of an in-depth analysis conducted using an Italian diachronic corpus. Speci cally, several approaches based on both static embeddings and dynamic ones are implemented and evaluated by using the Kronos-It dataset. We train all word embeddings on the Italian Google n-gram corpus. The main result of the evaluation is that all approaches fail to signi cantly reduce the number of false-positive change points, which con rms that lexical semantic change is still a challenging task.</p>
      </abstract>
      <kwd-group>
        <kwd>Computational Historical Linguistics • Diachronic word embeddings • Lexical Semantic Change</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background and Motivations</title>
      <p>
        Diachronic Linguistics concerns the investigation of language change over time.
Language change involves all levels of linguistic analysis: phonology, morphology,
syntax and semantics [
        <xref ref-type="bibr" rid="ref5 ref6">6, 5</xref>
        ]. In this work, we focus on lexical semantic change.
Two recent surveys [
        <xref ref-type="bibr" rid="ref11 ref19">11, 19</xref>
        ] describe and compare several lexical semantic change
models that have been developed in the last years. Several datasets and tasks
are employed in the evaluation of those models. In [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], the authors use two
corpora of scienti c papers and a corpus of senate speeches, both written in
Copyright '2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
English. They compare Static Bernoulli Embedding [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], Procrustes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and
Dynamic Bernoulli Embeddings [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] using the held-out likelihood as evaluation
metric. In [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], the authors evaluate, by using a rank-based approach, word2vec
embeddings and a variant of Procrustes alignment to detect words that have
undergone a semantic shift. Solving temporal word analogies is a common task
used to evaluate models of lexical semantic change, which consists in detecting
words analogies across time slices. In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the authors exploit the datasets created
by [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] to compare Temporal Word Embeddings with Compass [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
LinearTrans-Word2vec [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], Procrustes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Dynamic Word Embeddings [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
and Geo-Word2vec [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, few standard resources for evaluating lexical
semantic change detection models are available. Currently, this gap is tackled by
several initiatives. In [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], the authors introduce a framework (DUREL) for the
annotation of lexical semantic change and at the same time they make available
the annotated data1. DUREL is also employed in the annotation process of
Semeval 2020 Task 1 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] that involves four languages: English, Sweden, German
and Latin, while the Italian language remains under-explored. Semeval 2020 Task
1 provides corpora in four languages and a gold standard of lexical semantic
changes for the evaluation of unsupervised systems. However, the Semeval 2020
Task 1 corpora can only be used to evaluate lexical semantic change across two
time periods. Therefore, it cannot be used to perform a more ne-grained analysis
of the results. In this work, we describe a systematic evaluation of models for
lexical semantic change detection with the Italian Google Ngram as the corpus
for training word embeddings and Kronos-it [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as the gold standard for the
evaluation. Kronos-IT is a dataset for the evaluation of semantic change point
detection algorithms for the Italian language automatically built by using a web
scraping strategy. In particular, it exploits the information presents on the online
dictionary \Sabatini Colletti"2 to create a pool of words that have undergone a
semantic change. In the dictionary, some lemmas are tagged with the year of the
rst attestation of its sense. In some cases, associated with the lemma there are
multiple years attesting the introduction of new senses for that word. Kronos-IT
uses this information to identify the set of semantic changing words.
      </p>
      <p>
        Previous works about the Italian Google Ngram corpus and Kronos-it are
described in [
        <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
        ], but they are limited to the Temporal Random Indexing model
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and simple baselines based on word frequencies and collocations ignoring
recent approaches based on word embeddings.
      </p>
      <p>The paper is structured as follows: Section 2 describes the approaches under
analysis, while Section 3 reports details about the evaluation pipeline used in
our work. Results of the evaluation are reported and discussed in Section 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Models</title>
      <p>Traditional approaches produce word vectors that are not comparable across
time due to the stochastic nature of low-dimensional reduction techniques or
1 http://www.ims.uni-stuttgart.de/data/durel/
2 https://dizionari.corriere.it/dizionario_italiano/
sampling techniques. To overcome this issue a widely adopted approach is to
align the spaces produced for each time step, based on the assumption that only
few words change their meaning. Words that turn out to be not aligned after the
alignment, changed their semantics. In this work, we investigate two approaches
for producing word embeddings that are comparable across time.</p>
      <p>
        The rst approach is based on the alignment of computed word embeddings
(bins). Word vectors are computed before the alignment, once we get the bin
(the embeddings matrix for a speci c time slice), the di erent spaces obtained
for each time slice are aligned. An example of this kind of approach is Procrustes
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which aligns word embeddings with a rotation matrix. The assumption is
that each word space has axes similar to the axes of the other word spaces, and
two word spaces are di erent due to a rotation of the axes:
      </p>
      <p>R = arg minQT Q=I</p>
      <p>QW t</p>
      <p>W t+1</p>
      <p>F
where W t and W t+1 are two word spaces for time slices t and t + 1,
respectively, and Q is an orthogonal matrix that minimizes the Frobenius norm of the
di erence between W t and W t+1.</p>
      <p>
        The second approach directly produces aligned word embeddings for each
time slice, as it jointly learns word embeddings and aligns them. Dynamic word
embeddings (DWE) [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] fall in this second type of approaches and it is based
on the positive point-wise mutual information (PPMI) matrix factorization. In
a unique optimization function, DWE produces embeddings and tries to align
them according to the following equation:
      </p>
      <p>1
min</p>
      <p>U(t) 2
where the terms are, respectively, the factorization of the PPMI matrix Y (t), a
regularization term and the alignment constraint that keeps the word
embeddings similar to the previous and the next word embeddings.</p>
      <p>
        The objective function of static Bernoulli embeddings is closely related to
that of the CBOW (Continuous Bag of Words) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] model, except that static
Bernoulli embeddings regularize the embedding placing priors on both the
embedding and context vectors. Dynamic Bernoulli Embeddings (DBE) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] extends
static Bernoulli embeddings including the time dimension. Context vectors are
shared across all the time slices while embedding vectors are only shared within
a time slice. Moreover, Dynamic Bernoulli Embedding uses a Gaussian random
walk for obtaining smoothly changing estimates of each term embedding. The
random walk penalizes the shifting of consecutive vectors.
      </p>
      <p>
        Finally, we investigate Temporal Random Indexing (TRI) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] that is able to
produce aligned word embeddings in a single step. Unlike previous approaches,
TRI is a count-based method. TRI is based on Random Indexing [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], where a
word vector (word embedding) svjTk for the word wj at time Tk is the sum of
random vectors ri assigned to the co-occurring words taking into account only
documents dl 2 Tk. Co-occurring words are de ned as the set of m words that
precede and follow the word wj. Random vectors are vectors initialized randomly
and shared across all time slices so that word spaces are comparable.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>The corpus pre-processing module receives as input a corpus annotated with
the time label of each document. The rst operation is the corpus splitting into
temporal slices. During the splitting, the dictionary is computing by keeping
track of each new token encountered and its occurrence. The nal dictionary is
built with all tokens present in each time slice and selecting the rst n tokens
sorted by the number of occurrences. In our evaluation, we consider n = 50; 000.
3 https://github.com/williamleif/histwords
4 https://github.com/mariru/dynamic_bernoulli_embeddings
5 https://github.com/yifan0sun/DynamicWord2Vec
6 https://github.com/pippokill/tri
3.2</p>
      <sec id="sec-3-1">
        <title>Bins building</title>
        <p>The second module takes as input tokenized documents for each time slice and
generates for each approach preliminary information useful for the next steps.
It has an execution mode for each approach namely Word2Vec, PPMI, Static
Bernoulli and Temporal Random Indexing. Word2Vec mode trains a Word2Vec
model on each sub-corpus using Gensim7, an open-source library for
unsupervised topic modelling and natural language processing. The PPMI mode
constructs a PPMI matrix for each time slice, which will then be used to create
Dynamic Word Embedding. The Bernoulli mode builds static Bernoulli
embedding for each time slice that will later be used to construct Dynamic Bernoulli
embeddings. The Temporal Random Indexing mode saves the occurrences of
words and contexts that we will later be used to create word embeddings.
3.3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Alignment</title>
        <p>The aim of the alignment module is the alignment of the bins produced as
output in the previous module, and it is composed of several sub-modules:
Procrustes Aligner, Bernoulli Aligner, Dynamic word embeddings construction and
the TRI sub-module. The Bernoulli Aligner constructs Dynamic Bernoulli
Embeddings starting from the static Bernoulli output. Procrustes Aligner is the
sub-module that takes each Word2Vec model and applies Procrustes to each
time slice. The Dynamic Word Embeddings sub-module takes the PPMI
matrices previously created for building the Dynamic Word embeddings model. The
TRI sub-module produces word vectors for each time slice by relying on the
co-occurrences information built in the previous step.
3.4</p>
      </sec>
      <sec id="sec-3-3">
        <title>Time-series and change point detection</title>
        <p>We compute time-series by exploiting the word embeddings created for each time
slice. A time-series for each word is built, this result in a matrix W V xT where
V is the dictionary size and T is the number of time slices.</p>
        <p>We explore two approaches for the computation of the time-series, namely
point-wise and cumulative. In the point-wise approach, the element i; j of W V xT
represent the cosine similarity
where wi is the i-th word in the dictionary and j is the j-th time slice. While, in
the cumulative approach, the element i; j of W is
7 https://radimrehurek.com/gensim/</p>
        <p>Wi;j = cos(vwji 1; vwji )
Wi;j = cos(</p>
        <p>Pj k 1
k=1 vwi ; vwji )</p>
        <p>j</p>
        <p>
          In order to detect change points, we use the algorithm proposed in [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
According to this model, we de ne a mean shift of a general time-series Wi
pivoted at time period j as:
        </p>
        <p>K(Wi) =</p>
        <p>1
l
j</p>
        <p>l
X
k=j+1</p>
        <p>Wi;k</p>
        <p>j
1 X
j
k=1</p>
        <p>
          Wi;k
(1)
To understand if a mean shift is statistically signi cant at time j we use a
bootstrapping [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] approach under the null hypothesis. The null hypothesis states there
is no change in the mean. We sample B bootstrap examples by permuting Wi;j .
For each bootstrap sample P, K(P ) is calculated to provide its corresponding
bootstrap statistic and statistical signi cance (p-value) of observing the mean
shift at time j compared to the null distribution. Finally, we estimate the change
point by considering the time point j with the minimum p-value score.
        </p>
        <p>Change points together with the year, the p-value and the word are stored
in a le used for the evaluation.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <sec id="sec-4-1">
        <title>Data</title>
        <p>
          For the training, we use the Google Ngram, a dataset of ngrams extracted by
305,763 Google Books. Google Ngram covers the period from 1500 to 2012. OCR
errors can occur more in older historical documents, then we extract a sub-corpus
concerning the period 1900-2010. We split Google Ngram corpus into ten slices
with a range of ten years, starting from 1900 to 2010. We chose a time span
of ten years for reducing the computational complexity since semantic changes
are not frequent and generally require a large time span to be observed. Since
the full text is not available in the Google Ngram, we use the method described
in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] for extracting co-occurrences between words. As gold standard, we use
Kronos-it [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], a dataset for the Italian lexical change detection task. Kronos-it
provides for each lemma a set of years indicating the semantic change for that
lemma. Kronos-it is extracted by the Sabatini Coletti, an Italian dictionary that
contains for some word meanings the year of the rst appearance. The Kronos-it
dataset contains 13,818 lemmas and 13,932 change points. Lemmas reported in
Kronos-it have, on average, one change point.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Hyper-parameters</title>
        <p>We use the same hyper-parameters values shared by two or more models. We
use the same values for the context-window and the dimension of the
embeddings. Table 1 reports training strategies and hyper-parameters values. We adopt
default values used by the authors of the models.</p>
        <p>In particular, in DWE we specify the number of iterations over the data, the
alignment weight , the regularization weights and . In TRI, we set the down</p>
        <p>DWE TRI DBE Procrustes
Parameter Value Parameter Value Parameter Value Parameter Value
dimension 300 dimension 300 dimension 300 dimension 300
window 4 window 4 window 4 window 4
iters 5 down-sampling 0.001 negatives 2 min-count 1
10 seeds 10 minibatch 1000 negatives 20
100 n epochs 4 sample 1e-5
50 iter 4
sampling factor, and the number of seeds. In DBE, we set the number of negative
samples, the minibatch size and the number of epochs. In Procrustes, we set the
minimum number of occurrences a token must have to appear in the dictionary
min-count, the number of negative samples, the downsampling parameter sample
and the number of iterations over the data.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Metrics</title>
        <p>We compute the performance of each approach by using Precision, Recall and
Fmeasure. In the evaluation, a true positive is a change point for a word reported
in the gold standard that belongs to the range of the ten years predicted by the
system for that word. Change points provided by the systems are compared to
the change points reported in the gold standard. The false negatives (FN) are
the number of change points in the gold standard minus the true positives. The
false positives (FP) are the number of change points provided by the system
minus the true positives.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Results</title>
        <p>Table 2 reports Precision (P), Recall (R) and F-measure (F) for each system.
We can observe that generally, we obtain a low F-measure. This is due to the
large numbers of change points detected by each system (false positive). We can
observe that the best approach is DWE point-wise. However, the results of DWE
point-wise are close to those obtained by Procrustes point-wise and TRI
cumulative. A remarkable aspect is the worse performance of DBE respect those of
TRI and DWE, the entries of DBE time-series are very close to 1, this highlights
a heavy alignment. This is maybe due to the choice of hyper-parameters used to
train the DBE. We use, as mentioned above, the default hyper-parameters and
the type of datasets used by the authors is di erent from Google Ngrams, mainly
due to the large amount of data in the Google Ngrams. This could have a ected
results obtained by DBE. The results of the evaluation prove that the task of
semantic change detection is very challenging, in particular, the large number
of detected change points (false positive) drastically a ects the performance.
Sometimes change points are detected before or after the change point reported
0;4</p>
        <p>1
y
t
ira 0;9
l
i
m
i
S
e
n
iso 0;8
C
0;7</p>
        <p>1
y
t
i
r
a
l
i
m
iS0;5
e
n
i
s
o
C
0</p>
        <sec id="sec-4-4-1">
          <title>TRI cumulative</title>
          <p>Procrustes point-wise</p>
        </sec>
        <sec id="sec-4-4-2">
          <title>DBE point-wise</title>
          <p>DWE point-wise
atomica</p>
          <p>1960
palmare
1920
1940
1980</p>
          <p>2000
1920
1940
1980</p>
          <p>2000
in the gold standard, this supports the hypothesis that the change of semantics
of a word is a continuous process, which involves long periods before reaching a
stabilization. More studies are necessary to understand which component a ects
the performance, such an in-depth and explicit analysis of time-series.
Moreover, it is important to underline that the year reported in the dictionary may
be incorrect.</p>
          <p>In Figure 2, we show some examples of time-series. For the word `atomica',
DWE cumulative is the only approach that ts the change point in the gold
standard, indicating the change point as the decade 1950-1959, after 1945, year
of Hiroshima and Nagasaki. We do not detect change points in the time-series
produced by Procrustes point-wise and DBE point-wise, while we nd a change
point in the TRI-cumulative time-series in the 1950-1959 decade. For the word
`palmare', in the DBE point-wise and Procrustes cumulative time-series, two
change points are detected that are too early compared to the change point in
the gold standard 1998. Procrustes provided the right range 1950-1959 for the
word `Oscar', years in which for the rst time an Italian lm director, Vittorio De
Sica, won the Oscar. TRI cumulative and DBE point-wise do not detect change
points, while in the DWE point-wise time-series a change point is founded in the
decade 1960-1969.
5</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>In this paper, we present a systematic evaluation of Dynamic Word Embeddings,
Dynamic Bernoulli Embeddings, Procrustes and Temporal Random Indexing for
the lexical semantic change detection for the Italian language. The results show
that detect lexical semantic change is a complex task. A large number of change
points is detected by systems, a ecting the performance. A qualitative analysis of
words time-series highlights that some change points are detected just before or
after the correct period. This behaviour requires some further linguistic analysis
for understanding the reasons behind.</p>
      <p>This work can be extended in two directions: 1) including some recent
models of lexical semantic change that involve contextual embeddings and a
hyperparameter search optimized on the Italian Google Ngram dataset; 2)
investigating other diachronic Italian corpora as training data. Moreover, we plan to
investigate further methods for detecting changes in time-series.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This research has been partially funded by ADISU Puglia under the
postgraduate programme "Emotional city: a location-aware sentiment analysis
platform for mining citizen opinions and monitoring the perception of quality of
life".</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bamman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>N.A.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of geographically situated language</article-title>
          .
          <source>In: Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          . pp.
          <volume>828</volume>
          {
          <issue>834</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luisi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semeraro</surname>
          </string-name>
          , G.:
          <article-title>Diachronic analysis of the Italian language exploiting google ngram</article-title>
          .
          <source>In: Proceedings of the Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ). p.
          <fpage>56</fpage>
          . CEUR.org (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semeraro</surname>
          </string-name>
          , G.:
          <article-title>Analysing word meaning over time by exploiting temporal random indexing</article-title>
          .
          <source>In: First Italian Conference on Computational Linguistics</source>
          CLiC-it (CLiC-it
          <year>2014</year>
          ). CEUR.org (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Semeraro</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caputo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Kronos-it: a Dataset for the Italian Semantic Change Detection Task</article-title>
          .
          <source>In: Proceedings of the 6th Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2019</year>
          ). CEUR.org (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Blank</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Why do new meanings occur? A cognitive typology of the motivations for lexical semantic change. Historical semantics and cognition (</article-title>
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bybee</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          :
          <article-title>Diachronic linguistics</article-title>
          .
          <source>In: The Oxford handbook of cognitive linguistics</source>
          . Oxford University Press (
          <year>2010</year>
          ), https://www.oxfordhandbooks.com/view/ 10.1093/oxfordhb/9780199738632.001.0001/oxfordhb-9780199738632-e-36
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Di</given-names>
            <surname>Carlo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Bianchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Palmonari</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Training temporal word embeddings with a compass</article-title>
          .
          <source>In: Proceedings of the AAAI Conference on Arti cial Intelligence</source>
          . vol.
          <volume>33</volume>
          , pp.
          <volume>6326</volume>
          {
          <issue>6334</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Efron</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tibshirani</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>An Introduction to the Bootstrap</article-title>
          . CRC Press (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Ginter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanerva</surname>
          </string-name>
          , J.:
          <article-title>Fast Training of word 2 vec Representations Using N-gram</article-title>
          <string-name>
            <surname>Corpora</surname>
          </string-name>
          (
          <year>2014</year>
          ), https://www2.lingfil.uu.se/SLTC2014/abstracts/sltc2014_ submission_27.pdf
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>W.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leskovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change</article-title>
          . In:
          <article-title>Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</article-title>
          . pp.
          <volume>1489</volume>
          {
          <issue>1501</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kutuzov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>vrelid</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>Szymanski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velldal</surname>
          </string-name>
          , E.:
          <article-title>Diachronic word embeddings and semantic shifts: a survey</article-title>
          .
          <source>In: Proceedings of the 27th International Conference on Computational Linguistics</source>
          . pp.
          <volume>1384</volume>
          {
          <issue>1397</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient Estimation of Word Representations in Vector Space (</article-title>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Rudolph</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Dynamic embeddings for language evolution</article-title>
          .
          <source>In: Proceedings of the 2018 World Wide Web Conference</source>
          . pp.
          <volume>1003</volume>
          {
          <issue>1011</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Rudolph</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruiz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandt</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Exponential family embeddings</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . pp.
          <volume>478</volume>
          {
          <issue>486</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Sahlgren</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>An introduction to random indexing</article-title>
          .
          <source>In: Methods and Applications of Semantic Indexing Workshop at the 7th International conference on Terminology and Knowledge Engineering</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Schlechtweg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGillivray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hengchen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dubossarsky</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tahmasebi</surname>
          </string-name>
          , N.:
          <article-title>SemEval-2020 Task 1: Unsupervised Lexical Semantic Change Detection</article-title>
          .
          <source>In: Proceedings of the 14th International Workshop on Semantic Evaluation. Association for Computational Linguistics</source>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Schlechtweg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , im Walde,
          <string-name>
            <given-names>S.S.</given-names>
            ,
            <surname>Eckmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Diachronic Usage Relatedness (DURel): A Framework for the Annotation of Lexical Semantic Change</article-title>
          .
          <source>In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>2</volume>
          (Short Papers). pp.
          <volume>169</volume>
          {
          <issue>174</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Szymanski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Temporal word analogies: Identifying lexical replacement with diachronic word embeddings</article-title>
          .
          <source>In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          . pp.
          <volume>448</volume>
          {
          <issue>453</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Tahmasebi</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jatowt</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Survey of computational approaches to lexical semantic change</article-title>
          . arXiv preprint arXiv:
          <year>1811</year>
          .
          <volume>06278</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , W.A.:
          <article-title>Change-point analysis: a powerful new tool for detecting changes</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Tsakalidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bazzi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cucuringu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basile</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGillivray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Mining the UK Web Archive for Semantic Change Detection</article-title>
          .
          <source>In: Proceedings of the International Conference on Recent Advances in Natural Language Processing (RANLP</source>
          <year>2019</year>
          ). pp.
          <volume>1212</volume>
          {
          <issue>1221</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ding</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rao</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiong</surname>
          </string-name>
          , H.:
          <article-title>Dynamic word embeddings for evolving semantic discovery</article-title>
          .
          <source>In: Proceedings of the eleventh ACM International Conference on Web Search and Data Mining</source>
          . pp.
          <volume>673</volume>
          {
          <issue>681</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>