<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NLP-CIC @ DIACR-Ita: POS and Neighbor Based Distributional Models for Lexical Semantic Change in Diachronic Italian Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexander Gelbukh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIC, Instituto Polite ́cnico Nacional Mexico City</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Carlos A. Rodriguez-Diaz CIC, Instituto Polite ́cnico Nacional Mexico City</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Jason Angel CIC, Instituto Polite ́cnico Nacional Mexico City</institution>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Sergio Jimenez Instituto Caro y Cuervo Bogota</institution>
          ,
          <country country="CO">Colombia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present our systems and findings on unsupervised lexical semantic change for the Italian language in the DIACR-Ita shared-task at EVALITA 2020. The task is to determine whether a target word has evolved its meaning with time, only relying on raw-text from two time-specific datasets. We propose two models representing the target words across the periods to predict the changing words using threshold and voting schemes. Our first model solely relies on part-of-speech usage and an ensemble of distance measures. The second model uses word embedding representation to extract the neighbor's relative distances across spaces and propose “the average of absolute differences” to estimate lexical semantic change. Our models achieved competent results, ranking third in the DIACR-Ita competition. Furthermore, we experiment with the k neighbor parameter of our second model to compare the impact of using “the average of absolute differences” versus the cosine distance used in (Hamilton et al., 2016).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Lexical semantic change has recently gained
interest in the intersection of natural language
processing and historical linguistics1, therefore
several datasets have been proposed for different
languages
        <xref ref-type="bibr" rid="ref6 ref7">(Schlechtweg et al., 2020a)</xref>
        . This work
take place in the context of DIACR-Ita (Basile
“Copyright © 2020 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).”
1see https://languagechange.org/
et al., 2020a) at EVALITA 2020
        <xref ref-type="bibr" rid="ref2 ref3">(Basile et al.,
2020b)</xref>
        , which sets the task for the Italian language
in a fully unsupervised fashion. From
DIACRIta we received 18 target words2, and two
timespecific and preprocessed Italian corpora, namely
T0 and T1, which include part-of-speech tagging
and lemmatization information.
      </p>
      <p>We present two perspectives to approach the
problem, regarding how we represent target words
and estimate the lexical-semantic change across
datasets. (1) uses the POS distribution of target
words as representation, and employee an
ensemble of distance measures for the estimation. (2)
uses the target words neighbor similarities as
representation and one (of two proposed) similarity
measure for estimation.</p>
      <p>The following three sections describe the
previous works, modeling, and results we obtained
using these approaches. Following that, section 5
(Discussion) focuses on examine the second
approach to illustrate the impact of the k parameter
in similarity measures and the discriminatory
performance of our embedding-based model.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>
        Previous works have employed similar approaches
to address the unsupervised
lexical-semanticchange task, mostly for the English language
        <xref ref-type="bibr" rid="ref1 ref6 ref6 ref7 ref7">(Schlechtweg et al., 2020a; Asgari et al., 2020;
Schlechtweg et al., 2020b)</xref>
        . Our first approach
follows the idea of “syntactic models”
        <xref ref-type="bibr" rid="ref5">(Kulkarni
et al., 2015)</xref>
        , which supposes that some semantic
changes could imply a new syntactic functionality,
such as acquiring a new part-of-speech category,
as Kulkarni et al. (2015) exemplify: the word
“ap2’egemonizzare’, ’lucciola’, ’campanello’, ’trasferibile’,
’brama’, ’polisportiva’, ’palmare’, ’processare’, ’pilotato’,
’cappuccio’, ’pacchetto’, ’ape’, ’unico’, ’discriminatorio’,
’rampante’, ’campionato’, ’tac’, ’piovra’
ple” increased his use as a proper name in the ’80s.
      </p>
      <p>
        On the other hand, our second approach follows
the idea of “embedding-based models”
        <xref ref-type="bibr" rid="ref4 ref5 ref8">(Kulkarni
et al., 2015; Hamilton et al., 2016; Shoemark et
al., 2019)</xref>
        , which compares word vector
representations from each period using an aligned space,
which can be computed either globally (for the full
model) or locally (only for a target words). A
common strategy for local aligning is to perform a new
transformation representing the target words (the
same from different spaces) through neighborhood
structures, under the assumption that independent
training of embedding algorithms on comparable
corpora will still produce similar neighborhood
structures
        <xref ref-type="bibr" rid="ref5">(Kulkarni et al., 2015)</xref>
        .
      </p>
      <p>Our second approach align the space locally
using the nearest neighbors of target words as shared
feature.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>In this section we provide a detailed description
of our systems, each of them composed of two
stages, the model and the voting scheme.
3.1</p>
      <sec id="sec-3-1">
        <title>Models</title>
        <p>We represented the target words as vectors for
each time of period using two perspectives that
originate our submitted systems: the POS-model
and the embedding-model. The word
representations are comparable across spaces, and serve to
estimate the lexical semantic change through
similarity and distance measures, from which we
finally predict the changing words using thresholds
and voting schemes.</p>
        <p>POS-model: we simply analyzes the
Part-OfSpeech distribution as the relative frequency over
the datasets taking the top 4 most common
POStags, namely ADJ, NOUN, PROPN and VERB.
The produced four-dimensional vector pairs are
then used to assess the lexical semantic change of
each target word from the perspective of their
Euclidean, Manhattan and Cosine distances3.</p>
        <p>
          Embedding-model: We lowercase and
concatenate each word form with its corresponding
POS to build embedding models for each dataset
T , namely T 0 and T 1. Specifically, we used
Word2Vec models
          <xref ref-type="bibr" rid="ref5">(Mikolov et al., 2013)</xref>
          with the
CBOW version from gensim4 with the following
parameters: size of 256, window of 5, min count
of 3. Then we take the common vocabulary of both
Vc = V(T O) \ V(T 1), and use it to constraint the
set of top k nearest neighbors of the target word
only from T 05, i.e., Nk = fn1; n2:::nkg; nk 2 Vc,
to build the representation of the target word for
each space based on its neighbor proximity, i.e.
W~ T = [cos sim(w~ ; n~k)jnk 2 T ], and estimate
the lexical semantic change using the following
two formulas6:
avg.abs.diff = Avg(j W~ T 0
W~ T 1j)
(1)
cosine similarity = cos sim( W~ T 0; W~ T 1) (2)
The average of absolute point-wise differences
(avg.abs.diff for short) works under the
assumption that the neighbors a non-changing word
preserves their relative distance each other across
diachronic representations. Therefore, the value of
this measure increases according to the lexical
semantic change a target word underwent. In our
submission we used k = 10.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Threshold and voting schemes</title>
        <p>Given that DIACR-Ita is an unsupervised task
we experiment with different threshold and voting
schemes to aggregate the measure ranks and
determine which target words have underwent a lexical
semantic change. As a result, we propose three
voting schemes from which we derive our results.</p>
        <p>System1: Upper-third of distance ranks
(used for POS model): we sorted the target words
in descending order and rank their positions
according to the Euclidean, Manhattan and Cosine
distances. We then sum all these ranks and sort in
descending order again. Finally we label the first
upper-third part of this list as changing words.</p>
        <p>System2: Half intersection (used for the
embedding model): We sort the target words in
descending and ascending order for the
linealdifference scores (1) and the cosine-similarity (2)
respectively. Then we take the top 50% of each
group, and intersect them to obtained the words
that we predicted as changing words.</p>
      </sec>
      <sec id="sec-3-3">
        <title>System3: Union of Upper-third and Half in</title>
        <p>tersection: This is just the union of results from
System1 and System2.</p>
        <p>3we noticed that at this point Kulkarni et al. (2015) uses
Jenssen-Shannon divergence measure
4https://radimrehurek.com/gensim/models/word2vec.htm
5Unlike Hamilton et al. (2016) that takes the top-k
neighbors from each model and union them (Nk = NkT 0 [ NkT 1).
6Hamilton et al. (2016) only uses cosine distance.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>In this section we employee the gold-standard
labels of the target words to analyze at Figure 1
the capabilities of our neighbor-based
embeddingmodel using several settings. To this end, we
divide the Figure 1 into vertical and horizontal
views. The vertical view defines 3 groups (from
top to bottom), that serves to compare the three
proposed measures to estimate the lexical
semantic change, namely the average of absolute
differences, cosine similarity and cosine distance. At
the same time, the horizontal view serves to
compare the strategy of only use T0 (at left), versus
the union of T0 and T1 (at right), to define the top
nearest neighbors Nk.</p>
      <p>7https://github.com/ajason08/
evalita2020_diacrita</p>
      <p>Next, each of the charts shows an analysis of
the model for the given measure across the k
parameter. The area charts represent by color
regions the ranges that discriminate the lexical
semantic change of target words: “changing words”
(orange region) and “non-changing words”
(purple region). The yellow region in the middle marks
the intersection of these ranges, thus, words falling
into the yellow region are difficult to estimate,
according to the used measure. We also identified
the threshold that best discriminate changing and
non-changing target words, and draw a dashed line
at that point. On the other hand, the line charts
throw light on all the possible performance that the
model could obtain by changing the k parameter
while using the best possible discriminator
threshold.</p>
      <p>These results suggest that the “average of
absolute difference” is the best proposed measure
because it obtains a better performance for a larger
number of k values as displayed in the line charts.
Moreover, the “average of absolute difference”
offers a larger range for possible discriminator
thresholds (as shown in the area charts), and it is
tolerant to the Nk election, since it remains almost
unchanged while using either the union of T0 and
T1, or only T0. One can also note that the area
charts for the cosine similarity versus cosine
distance mirror each other, as expected, and their
performance is the same when using Nk only from T0
(at left), but slightly differ when using Nk as the
union of T0 and T1 (at right).
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We tackle the problem of unsupervised lexical
semantic change on two time-specific datasets for
18 target words in Italian language. Our two
approaches focus on the representation of target
words across the provided diachronic datasets,
they use part-of-speech usage and nearest
neighbors respectively, and a number of measures
between these representation to estimate the lexical
semantic change. Then, this estimation serves to
decide which target words underwent a change by
the use of proposed threshold and voting schemes.
Afterward, in the last part of this work, we
analyzed the nearest neighbor model through the
impact of deciding the k parameter and the
similarity measure that estimates the lexical semantic
change. Our results for the DIACR-Ita datasets
suggest that the estimations of “the average of
absolute differences” measures have a better
performance for a larger number of k values than the
cosine similarity and the cosine distance used in
Hamilton et al. (2016).</p>
      <p>As for future work, we plan to investigate
different mechanism for deciding the threshold, and
explore other diachronic datasets for other languages
such as English, German and Spanish. We also
believe that further experiments on a larger
number of target words will benefit the reliability of
models to judge the lexical semantic change in an
unsupervised fashion.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The authors thank CONACYT for the computer
resources provided through the INAOE
Supercomputing Laboratory’s Deep Learning Platform for
Language Technologies.
Web), page 625–635, Florence, Italy, May.
International World Wide Web Conferences Steering
Committee, Republic and Canton of Geneva,
Switzerland.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Ehsaneddin</given-names>
            <surname>Asgari</surname>
          </string-name>
          , Christoph Ringlstetter, and Hinrich Schu¨tze.
          <year>2020</year>
          .
          <article-title>Unsupervised embedding-based detection of lexical semantic changes</article-title>
          . arXiv preprint arXiv:
          <year>2005</year>
          .07979.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Pierpaolo</given-names>
            <surname>Basile</surname>
          </string-name>
          , Annalina Caputo, Tommaso Caselli, Pierluigi Cassotti, and
          <string-name>
            <given-names>Rossella</given-names>
            <surname>Varvara</surname>
          </string-name>
          .
          <year>2020a</year>
          .
          <article-title>DIACR-Ita @ EVALITA2020: Overview of the EVALITA2020 Diachronic Lexical Semantics (DIACR-Ita) Task</article-title>
          . In Valerio Basile, Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of the 7th evaluation campaign of Natural Language Processing</source>
          and
          <article-title>Speech tools for Italian (EVALITA 2020), Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Valerio</given-names>
            <surname>Basile</surname>
          </string-name>
          , Danilo Croce, Maria Di Maro, and
          <string-name>
            <surname>Lucia</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Passaro</surname>
          </string-name>
          . 2020b.
          <article-title>Evalita 2020: Overview of the 7th evaluation campaign of natural language processing and speech tools for italian</article-title>
          .
          <source>In Valerio Basile</source>
          , Danilo Croce, Maria Di Maro, and Lucia C. Passaro, editors,
          <source>Proceedings of Seventh Evaluation Campaign of Natural Language Processing and Speech Tools for Italian. Final Workshop (EVALITA</source>
          <year>2020</year>
          ),
          <article-title>Online</article-title>
          . CEUR.org.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>William L. Hamilton</surname>
            , Jure Leskovec, and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Cultural shift or linguistic drift? comparing two computational measures of semantic change</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>2116</fpage>
          -
          <lpage>2121</lpage>
          , Austin, Texas, November. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Vivek</given-names>
            <surname>Kulkarni</surname>
          </string-name>
          , Rami Al-Rfou,
          <string-name>
            <given-names>Bryan</given-names>
            <surname>Perozzi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Steven</given-names>
            <surname>Skiena</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Statistically significant detection of linguistic change</article-title>
          .
          <source>In WWW '15: Proceedings of the 24th International Conference on World Wide Tomas Mikolov</source>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barbara</surname>
            <given-names>McGillivray</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Simon</given-names>
            <surname>Hengchen</surname>
          </string-name>
          , Haim Dubossarsky, and
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          . 2020a. Semeval
          <article-title>-2020 task 1: Unsupervised lexical semantic change detection</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Dominik</given-names>
            <surname>Schlechtweg</surname>
          </string-name>
          ,
          <string-name>
            <surname>Barbara</surname>
            <given-names>McGillivray</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Simon</given-names>
            <surname>Hengchen</surname>
          </string-name>
          , Haim Dubossarsky, and
          <string-name>
            <given-names>Nina</given-names>
            <surname>Tahmasebi</surname>
          </string-name>
          . 2020b. Semeval
          <article-title>-2020 task 1: Unsupervised lexical semantic change detection</article-title>
          . arXiv preprint arXiv:
          <year>2007</year>
          .11464.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Philippa</given-names>
            <surname>Shoemark</surname>
          </string-name>
          , Farhana Ferdousi Liza,
          <string-name>
            <given-names>Dong</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Scott</given-names>
            <surname>Hale</surname>
          </string-name>
          , and
          <string-name>
            <surname>Barbara McGillivray</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Room to Glo: A systematic comparison of semantic change detection approaches with word embeddings</article-title>
          .
          <source>In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          , pages
          <fpage>66</fpage>
          -
          <lpage>76</lpage>
          ,
          <string-name>
            <surname>Hong</surname>
            <given-names>Kong</given-names>
          </string-name>
          , China, November. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>