<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>SEBD</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Contextualised Semantic Shift Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francesco Periti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Università degli Studi di Milano</institution>
          ,
          <addr-line>Via Celoria 18, 20133 Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>31</volume>
      <fpage>02</fpage>
      <lpage>05</lpage>
      <abstract>
        <p>Language continuously evolves influenced by social practices, events, and political circumstances; goal of Semantic Shift Detection is to detect, interpret, and assess changes of word meanings over time. The recent development of computational semantics pushed the emergence of approaches based on word embedding techniques for detecting semantic shift mainly at word-level. The Ph.D. research focuses on the problem of Contextualised Semantic Shift Detection (CSSDetection), which is the use of contextualised embeddings for capturing and interpreting “semantic shift” in the meaning(s) of words. In particular, the research aims to: (1) define a novel approach to trace the evolution of word meanings over time, and (2) extend CSSDetection to also capture and interpret semantic shift in the usage(s) of sentences.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computational Semantics</kwd>
        <kwd>Contextualised Word Embeddings</kwd>
        <kwd>Semantic Shift Detection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Language continuously evolves influenced by social practices, events, and political
circumstances. This means that studying how words and sentences (e.g., quotations, idioms) change
in meaning/usage over time can help to deeper understand the evolution of the political and
social landscape [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. The recent availability of large diachronic corpora and the development
of computational semantics pushed the emergence of approaches based on Natural Language
Processing (NLP) techniques for detecting semantic shift of word meanings. In particular, over
the past three years, significant advancements in the field of Semantic Shift Detection (SSD) have
been made almost exclusively based on contextualised word embedding models (e.g., BERT) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Inspired by recent studies in literature, the Ph.D. research proposes to use contextualised
word embeddings to capture how words and sentences shift semantically over time. While
computational approaches have been already proposed for SSD, this research is the first, to our
knowledge, that extends the notion of semantic shift from word-level to sentence-level by also
aiming to detect, interpret, and assess the possible change in usage context of sentences.</p>
      <p>The remainder of the paper is organised as follows. In Section 2, the research problem is
presented and the PhD. research questions are outlined. In Section 3, an original analysis of
the relevant literature is discussed. Finally, preliminary results, ongoing and future work are
illustrated in Section 4.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Research problem</title>
      <p>
        The research problem of the Ph.D. thesis consists in tracing both word-level and
sentencelevel semantic shift by leveraging contextualised embeddings. For clarification purposes, as a
word-level example, consider the word isolated, which changed from its “feeling detached”
connotation to “quarantine” during the course of the COVID-19 pandemic [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Similarly, a
sentence-level example is the quotation to John 15:13, namely “There is no other love rather
than if someone gives soul for their friends”, uttered by the Russian President Vladimir Putin in a
political and non-religious context while praising Russian military forces’ actions in the war
in Ukraine1. However, it’s worth noting that sentence-level semantic shift refers not only to
changes in the meaning of individual sentences, but also, in a broader sense, to changes in the
discursive, cultural, and historical contexts in which the text is situated.
      </p>
      <p>For the sake of readability, we formalise the problem of SSD at word-level. This simplification
enables to review approaches to SSD in a clear and concise fashion, while being easily extendable
to sentence-level SSD. Consider a diachronic document corpus</p>
      <p>=
 = ⋃︁  ,</p>
      <p>
        =1
where  denotes a set of documents of the time . Contextualised SSD (CSSDetection) consists
in assessing the change of meaning for a set of target words  occurring in  across the whole
time span [1 . . . ] by leveraging contextualised embeddings [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For a word  ∈  occurring
in , its contextualised word representation (i.e., embedding) in the -th sentence is extracted
by a contextualised language model and it is denoted by . This means that the representation
of  in the corpus is a set
      </p>
      <p>Φ = {1 , 2 , ..., , ..., } .</p>
      <p>Then, the semantic shift score of  between two sub-corpora 1 and 2 is assessed by using a
distance function  between two sets Φ1 and Φ2, with  defined as</p>
      <p>
        : {R}1 , {R}2 → R ;
where  is the dimension of the word vectors, and 1 and 2 are the frequency of  in 1
and 2, respectively [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <sec id="sec-2-1">
        <title>2.1. Research questions</title>
        <p>Specifically, the Ph.D. research aims to extend the state of the art by answering the following
research questions:
RQ1: How can the evolution of word meanings be traced over time and how can it be used
to describe and categorise word-level semantic shift? As a matter of fact, most of the existing
solutions to CSSDetection focus on quantifying the degree of semantic shift for a target word.
Although some approaches have been appearing to enable the interpretation of which word
1www.washingtonexaminer.com/news/putin-invokes-the-bible-to-justify-ukraine-invasion-at-moscow-rally
meaning(s) is lost, gained, or changed (e.g., broadening/narrowing, amelioration/pejoration),
they are typically designed to analyse a corpus spanning two time periods. As a result, the
evolution of word meanings over time (more than two time periods) cannot be easily traced.
Thus, RQ1 aims to face this issue by promoting the design of i) a novel CSSDetection approach
capable of tracing the word meaning evolution, and ii) novel analysis techniques to describe
diferent categories of word-level semantic shift.</p>
        <p>
          RQ2: How can sentence-level semantic shift be captured, interpreted, and traced over time?
Although computational approaches have been already proposed for detecting semantic shift
at word-level, we are not aware of approaches for detecting semantic shift at sentence-level.
We argue that sentence-level SSD is meaningful for linguistics, social, and historical analysis,
as it allows for a more comprehensive understanding of language evolution. For instance,
sentence-level SSD can help identify changes in the use of quotations, idioms, collocations, and
other multi-word expressions, which are not captured by word-level analysis alone [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Thus,
RQ2 aims to address the research gap by extending RQ1 for studying sentence-level semantic
shift.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Original analysis of the literature</title>
      <p>
        We recently proposed a comprehensive classification framework for CSSDetection [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which
distinguishes approaches based on three dimensions of analysis: meaning representation,
timeawareness, and learning modality.
      </p>
      <p>Meaning representation concerns the representation of word meanings: i) form-based
approaches focus on high-level properties of a word, such as its dominant meaning or its degree
of polysemy; on the opposite ii) sense-based approaches focus on low-level properties of a word,
i.e. its multiple and diferent meanings.</p>
      <p>Time-awareness focuses on how the time information of the documents is considered in the
embedding model: i) time-oblivious approaches do not consider the time at which a document
is inserted in the corpus; on the opposite ii) time-aware approaches use a specific mechanism to
encode the time information into the embeddings.</p>
      <p>Learning modality is about the possible use of external knowledge for describing and learning
the word meanings or the semantic shift to recognise: i) supervised approaches exploit external
knowledge such as a dictionary or human-annotated dataset; on the opposite ii) unsupervised
approaches derive word meanings/semantic shift from the text in the corpus by using unsupervised
learning techniques.</p>
      <sec id="sec-3-1">
        <title>3.1. Approaches to CSSDetection</title>
        <p>Usually, CSSDetection approaches follow a three-step scheme: i) extraction of embeddings for
each occurrence of a target word from a contextualised language model such as BERT, ELMo, or
XLM-R; ii) an optional aggregation of the embeddings by averaging and/or clustering; and iii)
the application of a semantic shift function like Cosine Distance or Jensen-Shannon Divergence.</p>
        <p>For the sake of clarity, in the following, we present the main solutions according to the
meaning representation of the considered target word, namely form- and sense- based approaches,
respectively.</p>
        <p>
          Form-based approaches. Word embeddings are optionally aggregated by averaging in a
single representation, and used as input of a semantic shift function. On the one hand,
formbased approaches that aggregate embeddings tend to detect the shift of the dominant meaning
of the word  (e.g., [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]). On the other hand, form-based approaches that do not aggregate
embeddings tend to detect the shift in the degree of the polysemy of the word (e.g., [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]).
        </p>
        <p>
          Most form-based approaches follow the general scheme and are time-oblivious. A few
timeaware approaches have been recently published and they are all characterised by the adoption
of a specific fine-tuning operation to inject time information into the embedding model before
assessing the semantic shift of a word (e.g., [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]).
        </p>
        <p>
          All existing form-based approaches leverage unsupervised learning modalities. As an
exception, a Word-in-Context model (WiC) is trained in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] to reproduce the behavior of human
annotators in the manual annotation task.
        </p>
        <p>
          Sense-based approaches. In sense-based approaches, word embeddings are usually
aggregated by clustering (e.g., [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]). All the documents of two time periods are considered as a
whole, and a single clustering activity is performed, generating clusters with documents of
diferent time periods. The idea is that each cluster denotes a specific word meaning that can be
recognised in the considered documents. In this way, it is possible to assess and quantify the
shift in meaning of a word by analysing the cluster membership of the documents.
        </p>
        <p>
          All existing sense-based approaches are time-oblivious and most leverage unsupervised
learning modalities. A number of unsupervised clustering algorithms (e.g., K-Means [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]) are
proposed to sidestep the need of lexicographic resources.
        </p>
        <p>
          Only a few approaches employ a lexicographic supervision. For instance, a supervised
clustering is enforced in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] by leveraging a reference dictionary (i.e., the Oxford English
dictionary) to list the possible lexicographic meaning of a word beforehand; thus it is hardly
applicable to low-resource languages.
        </p>
        <p>
          Contrary to form-based approaches, sense-based approaches enable the semantic shift
interpretation by performing an in-depth qualitative analysis of the resulting clusters. For instance, a
cluster is often inspected by selecting the documents associated with the top closest vectors to its
cluster centroid (e.g., [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]) or its most discriminating Tf-Idf keywords (e.g., [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]). However, when
more than two time periods are considered, clusters of word meanings need to be re-calculated,
meaning that scalability issues arise and that resulting clusters could dramatically change from
one time period to the next. Thus, the possible evolution patterns of a meaning across diferent
time periods cannot be captured. As a possible solution, some recent works propose to perform
clustering separately for each time period. In this case, the resulting clusters need to be aligned
in order to recognise similar word meanings in diferent, consecutive time periods (e.g., [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]).
To avoid continuously re-calculating and aligning clusters, we propose to use an incremental
clustering algorithm to trace the evolution of clusters/word meanings over time [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
4. Preliminary results, ongoing and future work
In this section, we describe preliminary research results achieved so far, ongoing and future
work for each research question formulated in Section 2.
        </p>
        <p>RQ1: How can the evolution of word meanings be traced over time and how can it be used to
describe and categorise word-level semantic shift?</p>
        <p>
          We have recently proposed and evaluated a novel approach, called WiDiD, based on
incremental clustering of contextualised embeddings [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. WiDiD works under the assumption
that the documents of the corpus  become available as a stream and they are segmented
in a sequence of time periods. In this case, 2 represents the set of documents collected
at time , while 1 represents the cumulative set of documents collected in the  −  time
periods preceding . At each time step , a contextualised model is exploited to extract the
corresponding word embeddings of the word . Thus, at each step, two sets of embedding
vectors are available: Φ1 , the set of embeddings produced in the previous iterations of the
WiDiD approach over the corpus 1; and Φ2, produced at the current time  for the corpus 2.
In order to group word embeddings representing similar word meanings, a novel incremental
clustering algorithm called A Posteriori afinity Propagation (APP) is adopted. Finally, a
distance measure between the sets Φ1 and Φ2 is computed to quantify the semantic shift of
the word  in the considered time interval.
        </p>
        <p>
          Preliminary results of WiDiD have been evaluated and compared against a reference
benchmark using multiple configurations characterised by diferent clustering algorithms and
embedding methods [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. In particular, our experiments include the use of a pre-trained BERT
model and a trained Doc2Vec model, which has been adapted to provide pseudo-contextualised
word embeddings. A subset of results of our evaluation is shown in Table 1; further results are
discussed in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. All in all, WiDiD performs well in SSD compared to the approach based on
the conventional Afinity Propagation clustering.
        </p>
        <p>Corpus
SemEval</p>
        <p>Latin
SemEval
English</p>
        <p>Clustering
baseline</p>
        <p>AP
WiDiD
baseline</p>
        <p>AP
WiDiD</p>
        <p>Model
trained Doc2Vec
pre-trained BERT
trained Doc2Vec
pre-trained BERT
trained Doc2Vec
pre-trained BERT
trained Doc2Vec
pre-trained BERT</p>
        <p>JSD
0.485*
0.394*
0.512*
0.361*
0.514*
0.356*
0.333*
0.302*</p>
        <p>PDIS
0.229
0.347*
0.337*
0.210
0.139
0.326*
0.077
0.512*</p>
        <p>PDIV
-0.023
0.236
0.328*
0.036
0.134
0.406*
-0.078
0.370*</p>
        <p>Ongoing and future work to address RQ1 are about the definition of cluster analysis
techniques to describe and categorise patterns of semantic shift by considering the evolution of
word meanings. For instance, we are currently working on defining a set of metrics to describe
stable, growing, and shrinking trend in the dominance of a specific word meaning. In addition,
since our WiDiD evaluation was executed on a benchmark corpus spanning two time periods,
we are currently evaluating the WiDiD approach on a benchmark spanning more than two
periods. Finally, we are currently working on a real-world application of WiDiD on a large
corpus of Italian parliamentary speeches spanning 18 diferent time periods (i.e., 18 legislatures).
RQ2: How can sentence-level semantic shift be captured, interpreted, and traced over time?</p>
        <p>To address this research question, we propose to extend the WiDiD approach, which actually
enforces word-level SSD, to deal with the sentence-level SSD. In particular, as case study, we
plan to focus on semantic shift of quotations, meaning that we would like to capture, interpret,
and trace how the context of a quotation  change over time.</p>
        <p>
          The Vatican corpus. As an extension of our previous work in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], we plan to use a
diachronic corpus of Vatican Publications as a case study for sentence-level SSD. This corpus
contains plenty of quotations to the Bible and thus, it represents a perfect case study for the
new SSD topic. Currently, the considered corpus of Vatican publications includes all the
web-available documents from the digital Vatican archive2 at the time of downloading. This
corpus represents a valuable source for experimenting with SSD techniques for three reasons.
Firstly, it is characterised by an exceptional historical depth. Secondly, the documents are
available in various languages including Italian, Latin, English, Spanish, and German. And
ifnally, the third reason is that, through the writings of its popes, the Catholic Church has
always dealt with the most relevant issues in the public debate of its time, alongside themes of
faith and worship. Thus, these writings constitute a primary historical source for understanding
an important part of human cultural history, where the focus of public discourse shifted over
time to diferent topics such as the environment, the role of science, and various historical events.
        </p>
        <p>
          Ongoing and future work is about the definition of a novel framework to evaluate the
WiDiD approach for sentence-level SSD. Inspired by the recent shared tasks for word-level SSD
(e.g., SemEval-2020 Task 1 [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]), we are currently defining a manual annotation task to collect
gold data for three distinct computational tasks:
• Binary Change Detection: classifying a quotation  as stable or changed in meaning/context
between two time-specific corpora 1 and 2;
• Graded Change Detection: quantifying the extent to which the meaning/context of  shifts
between 1 and 2;
• Sentence Sense Disambiguation: identifying the intended meaning/context of each
occurrence of  from a synchronic perspective.
Furthermore, we plan to organise a shared task (e.g., for the SemEval series3) with the aim of
benchmarking and comparing diferent approaches in the field of sentence-level SSD.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Azarbonyad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Beelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Arkut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marx</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          , Words Are Malleable:
          <article-title>Computing Semantic Shifts in Political and Media Discourse</article-title>
          ,
          <source>in: Proc. of CIKM</source>
          , ACM, New York, NY, USA,
          <year>2017</year>
          , pp.
          <fpage>1509</fpage>
          -
          <lpage>1518</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Castano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montanelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Periti</surname>
          </string-name>
          ,
          <article-title>Semantic Shift Detection in Vatican Publications: a Case Study from Leo XIII to Francis</article-title>
          , in
          <source>: Proc. of SEBD</source>
          , CEUR-WS, Pisa, Italy,
          <year>2022</year>
          , pp.
          <fpage>231</fpage>
          -
          <lpage>243</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Montanelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Periti</surname>
          </string-name>
          ,
          <article-title>A survey on contextualised semantic shift detection</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2304</volume>
          .
          <fpage>01666</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Harrigian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dredze</surname>
          </string-name>
          ,
          <article-title>The Problem of Semantic Shift in Longitudinal Monitoring of Social Media: A Case Study on Mental Health During the COVID-19 Pandemic</article-title>
          , in
          <source>: Proc. of WebSci</source>
          , ACM, New York, NY, USA,
          <year>2022</year>
          , pp.
          <fpage>208</fpage>
          -
          <lpage>218</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Buccio</surname>
          </string-name>
          , M. Melucci, University of Padova@ DIACR-Ita,
          <source>in: Proc. of EVALITA</source>
          , CEUR-WS, Marrakech, Morocco,
          <year>2020</year>
          , pp.
          <fpage>420</fpage>
          -
          <lpage>425</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Walton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Macagno</surname>
          </string-name>
          ,
          <source>Wrenching from Context: The Manipulation of Commitments, Argumentation</source>
          <volume>24</volume>
          (
          <year>2010</year>
          )
          <fpage>283</fpage>
          -
          <lpage>317</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Martinc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Kralj Novak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pollak</surname>
          </string-name>
          ,
          <article-title>Leveraging Contextual Embeddings for Detecting Diachronic Semantic Shift</article-title>
          ,
          <source>in: Proc. of LREC</source>
          , European Language Resources Association, Marseille, France,
          <year>2020</year>
          , pp.
          <fpage>4811</fpage>
          -
          <lpage>4819</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Giulianelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Del Tredici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fernández</surname>
          </string-name>
          ,
          <article-title>Analysing Lexical Semantic Change with Contextualised Word Representations</article-title>
          ,
          <source>in: Proc. of ACL</source>
          , ACL, Online,
          <year>2020</year>
          , pp.
          <fpage>3960</fpage>
          -
          <lpage>3973</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pierrehumbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schütze</surname>
          </string-name>
          , Dynamic Contextualized Word Embeddings,
          <source>in: Proc. of ACL</source>
          , ACL, Online,
          <year>2021</year>
          , pp.
          <fpage>6970</fpage>
          -
          <lpage>6984</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>542</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Arefyev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fedoseev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Protastov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Homiskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Davletov</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Panchenko,
          <article-title>DeepMistake: Which Senses are Hard to Distinguish for a Word-in-Context Model</article-title>
          ,
          <source>in: Proc. of Dialogue</source>
          ,
          <source>(online)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <article-title>Diachronic Sense Modeling with Deep Contextualized Word Embeddings: An Ecological View</article-title>
          ,
          <source>in: Proc. of ACL</source>
          , ACL, Florence, Italy,
          <year>2019</year>
          , pp.
          <fpage>3899</fpage>
          -
          <lpage>3908</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -1379.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Montariol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martinc</surname>
          </string-name>
          , L. Pivovarova,
          <article-title>Scalable and Interpretable Semantic Change Detection, in: Proc. of NAACL-HLT, ACL</article-title>
          , Online,
          <year>2021</year>
          , pp.
          <fpage>4642</fpage>
          -
          <lpage>4652</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Periti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ferrara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Montanelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ruskov</surname>
          </string-name>
          ,
          <article-title>What is Done is Done: an Incremental Approach to Semantic Shift Detection</article-title>
          ,
          <source>in: Proc. of LChange</source>
          , ACL, Dublin, Ireland,
          <year>2022</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>McGillivray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hengchen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dubossarsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Tahmasebi, SemEval2020 Task 1: Unsupervised Lexical Semantic Change Detection</article-title>
          ,
          <source>in: Proc. of SemEval</source>
          , ICCL, Barcelona,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>