<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop on Computational Humanities Research, November</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>What about Grammar? Using BERT Embeddings to Explore Functional-Semantic Shifts of Semi-Lexical and Grammatical Constructions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lauren Fonteyn</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Leiden University Centre for Linguistics, Department of English Language and Culture</institution>
          ,
          <addr-line>Arsenaalstraat 1, 2311CT, Leiden</addr-line>
          ,
          <country country="NL">the Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <volume>1</volume>
      <issue>4</issue>
      <fpage>8</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>The aim of this short paper is to extend the application of embedding-based methodologies beyond the realm of lexical semantic change. It focuses on the use of unsupervised BERT-embeddings and uncertainty measures (Classification Entropy), and assesses whether (and how) they can be used to (semi-)automatically flag possible functional-semantic changes in the use of the construction [BE about] in the Corpus of Historical American English (COHA).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Distributional Semantics</kwd>
        <kwd>Corpus Linguistics</kwd>
        <kwd>Grammatical change</kwd>
        <kwd>embeddings</kwd>
        <kwd>BERT</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Given its long tradition in computational and statistical research, it comes as no surprise that
text-based humanities have embraced the use of distributional-semantic ‘vectors’ or
‘embeddings’ – i.e. (compressed) numeric vector representations of a word’s contextual distribution
that serve as a proxy of that word’s meaning [e.g. 3]. In particular, in fields such as Corpus
Linguistics – a subfield of Linguistics which grew around the computer-aided retrieval,
annotation, and later also categorization of textual data [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] – there seems to be an unprecedented
interest in vector-based distributional semantic models, which ofer a quantifiable and
datadriven means of studying meaning. This interest has been fueled further with the arrival of
models equipped to create contextualized token vectors, which have eliminated the problems
associated with polysemy/homonymy conflation [e.g 7].
      </p>
      <p>
        In recent years, we have also witnessed a growth in the number of studies that have utilized
either “count” or “predictive” vector models [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to study historical and diachronic corpus
data [also see 19]. Such studies, which often involve examination of nearest neighbours and
cosine similarities between type- and/or token-vectors over time, have provided the key to
a data-driven means of detecting and describing the diachronic trajectory of, predominantly,
lexical change [e.g. 11, 9, 10, 13].
      </p>
      <p>
        One consequence of this focus on lexical change is that, at present, the number of
computational distributional semantic studies that consider the functional-semantic properties of
more abstract, grammatical constructions seems disproportionate compared to the interest in
the phenomenon within the (Corpus) Linguistic community. Much like lexical semantics, the
function(s) and underlying meaning(s) of grammatical structures – which are often notoriously
polysemous and cover a broad range of nuanced, abstract meanings – are prone to change, and
many linguists believe that the continued discovery, description and analysis of such changes
plays an essential role in fleshing out our understanding of the mechanisms and motivations of
language change. A logical and necessary continuation in the pursuit to automated semantic
(shift) detection in large diachronic corpora would therefore involve further, more in-depth
explorations of the extent to which embeddings can be employed to capture the (changing)
functional-semantic properties of grammar. Given that that diferent components of language
have difering diachronic dynamics [
        <xref ref-type="bibr" rid="ref12 ref21">21, 12</xref>
        ], they may pose diferent challenges in the
development of unsupervised, embedding-based means of detecting diachronic change in large corpora
– and it is only by exploring the functional-semantic properties of grammar (at whatever
scale and level of detail is deemed reasonable) that these challenges and, consequently, their
solutions, can be discovered.
      </p>
      <p>
        The present study, then, sets out to do precisely that: explore some possible avenues of
(semi)automatically detecting diachronic functional-semantic changes of grammatical constructions.
In that sense, the aim of this study is to extend the application of embeddings-based
methodologies beyond the realm of lexical semantics, and further add to the budding research on
whether and how token vectors [e.g. 14, 29] or contextualized embeddings [e.g. 4] can be used
to study the functional-semantic properties of grammar. More specifically, the study, which
focusses on embeddings created by means of BERT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] makes the following contributions:
1. It demonstrates that unsupervised BERT-embeddings can successfully be employed to
identify the diferent functions or ’usage types’ of the grammatical construction [BE about]
in English. 2. It formulates three expected functional-semantic developments of [BE about]
based on linguistic literature, and assesses whether (and how) BERT-embeddings can be used
in combination with Entropy Diference measures and time-sensitive t-SNE plots to
(semi)automatically detect these changes in the Corpus of Historical American English [COHA, 5].
3. It discusses potential pitfalls associated with the ‘present-day’ bias of the explored methods.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Data and Methodology</title>
      <sec id="sec-2-1">
        <title>2.1. Corpus and Data</title>
        <p>
          As a case study, I focus on the recent diachronic development of [BE about]. The data for this
study has been gathered from COHA, a corpus containing over 400 million words of American
English text written between 1810 and 2009. The corpus is balanced for genre (Fiction,
NonFiction, Magazine, Newspaper (after 1860)) and subgenre (e.g. prose, poetry, drama, etc.
(Fiction)). Such (sub-)genre balance is said to ensure that any changes observed in the corpus
will not simply be “artifacts of a changing genre balance”[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>
          While [BE about] can be used with a wide range of more abstract, grammatical meanings
(which vary substantially in frequency as well as distinctiveness), the present study will focus
on its three most frequent usage types: the futurate (e.g. I am about to leave), approximative
(e.g. There were about ten cats in the room), and descriptive use (e.g. This song is about
love). These usage types align with the higher-order sense categories distinguished in the
Oxford English Dictionary (OED) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], which includes a group of tokens expressing
approximation, another signalling connections/relations (in descriptions), and another containing tokens
(predominantly) followed by a to-infinitive expressing intention or imminent future.
        </p>
        <p>Figure 1 shows a two-dimensional t-SNE future approx descriptive other
mapping of the embeddings of 1,000 examples 40
of [be about], collected from COHA (decade
2000-2009). The examples were manually an- 20
notated at the level of granularity outlined 0
above. The sample contained 448 futurate, 20
279 approximative, and 225 descriptive uses 40
of [BE about]. The other 48 examples include
irrelevant structures (e.g. due to mistakes in 60 40 20 0 20 40 60
COHA’s tagging of the possessive marker ’s
as a finite form of the verb be), as well as three
much more infrequent usages of [BE about]. Figure 1: t-SNE of [BE about] token embeddings.
These infrequent types consist of spatial uses There are three largely distinct usage types: futurate,
(e.g. She must be somewhere about), the fixed approximative, and descriptive. Less frequent senses
expression that’s about it, as well as an ac- (e.g. spatial) are marked as ‘other’.
tional use which can be paraphrased as
‘occupied/dealing with’ (chiefly found in more or less fixed expressions such as be about your
business and know what one is about).</p>
        <p>
          To assess whether BERT-embeddings can be used to identify and distinguish the three
usage types under scrutiny, we can conduct a simple ‘sense distinction task’. For this task,
the embeddings of the 952 ‘relevant’ examples and their accompanying usage type labels were
used as the training set. The procedure involved fitting a logistic regression classifier with L2
regularization (as implemented in Scikit-learn [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]) on the embeddings created for the labelled
training set. Subsequently, the classifier was applied to unseen test set of 200 examples from
each of the 20 decades covered in COHA (with the exception of the first decade (1810-1819),
which contained only 73 tokens of [BE about]). In sum, the test set includes 3,873 unseen
examples for which a usage type label was predicted. The predicted labels were then assessed
against the true usage types of the tokens.
        </p>
        <p>Based on the manual assessment of the pre- 1.00
dicted usage type labels, it appears that the 0.98
model performs quite well at distinguishing 00..9946
tFhoer tthherefinealmdaeicnaduesaogfeCtOypHeAs , oofnl[yBE5 oaubtouotf]. trrccoe000...989280
200 tokens had been mislabelled, indicating a 0.86
ccllaassssiifificcaattiioonn aaccccuurraaccyy ooff 0t.h9e75m. oNdoetlarbelmy,atihnes 0.84 1800 1825 1850 1875 1b9i0n0 1925 1950 1975 2000
high with older data, ranging between 0.904
and 0.975 (as is shown in Figure 2). At the Figure 2: Accuracy based on labelled test set of
same time it can, perhaps unsurprisingly, be 3,873 unseen examples from the 20 decades included
noticed that the accuracy (slightly) decreases in COHA (1810-2009). The classification accuracy
as the linguistic data ages. ranges between 0.904 in the earliest decade, and 0.975</p>
        <p>Overall, these accuracy scores are encour- in the most recent decade.
aging, in that they highlight the efficiency of
BERT-embeddings in identifying specific usage types of a grammatical construction, such as
[BE about], in a large diachronic corpus. While that is an interesting observation in itself,
the remainder of this paper will focus on assessing whether BERT-embeddings can be used to
(semi-)automatically detect functional-semantic changes of grammatical constructions.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Method</title>
        <p>
          This paper’s methodological set-up starts from the assumption that changes in a word or
construction’s distribution, which can be captured in compressed numerical representations such
as BERT-embeddings, may indicate changes in its functional or semantic range. Following
earlier proposals using embeddings to detect lexical-semantic change in large diachronic
corpora [e.g. 11], this study will investigate whether known changes that have afected the [BE
about] construction can be detected by means of an embedding-based methodology combined
with Entropy Diference measures. More specifically, the task is conceptualized as follows: in
the case of lexical items, expansions or reductions in their distributional properties are often
equated with expansions or reductions of their possible interpretations – or, in other words,
with increases or decreases uncertainty regarding the exact interpretation of the lexical item.
To measure whether the “uncertainty over possible interpretations varies across time intervals”,
then, one can “compute the diference in entropy between the two usage type distributions in
these intervals” [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. An increase in entropy over time could be used to signal that the number
of interpretations of a word has increased (e.g. due to the emergence of a new usage type),
whereas a decrease in entropy would signal that the opposite has occurred (e.g. loss of a usage
type). In principle, it is possible to extend this approach to the study of grammatical items.
        </p>
        <p>
          However, unlike the distributional changes that accompany lexical-semantic changes of
wellknown examples such as broadcast or gay, the distributional changes witnessed for [BE about]
– and many grammatical constructions like it – seem to proceed in a protracted sequence of
small steps, often spanning over several centuries [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. The question, then, is whether we can
manipulate the use of Entropy Diference measures to detect not only the emergence or loss of
an entire usage type, but also to detect any small-scale shifts within a usage-type.
        </p>
        <p>The approach tested here works assumes that the researcher is interested in determining
whether any of the usage types they distinguished has changed in the time span covered by
their corpus with respect to a single reference point (for instance: Present-day data). For [BE
about], the procedure involved fitting a logistic regression classifier on the embeddings created
for the 952 present-day tokens, and applying it to a test set of 200 examples from each decade
included in COHA (cf. Section 2.1). Subsequently, for each test token xi and each label y ∈ Y ,
the conditional probability p(y|xi) is computed to assess the uncertainty of the classifier in
labelling the unseen examples. The resulting conditional probability over each label for each
test token is then summarized in an entropy score, H:</p>
        <p>H(xi) = −
∑ p(y|xi) ∗ log p(y|xi)
y∈Y
(1)
If the entropy score changes over the 20 decades included in COHA, one could take this as an
indication that the distributional properties of the test tokens in a particular category have
shifted over time. Such shifts could be indicative of proper (subtle) functional-semantic change,
or of increased or decreased use of the construction under scrutiny in diferent genres or text
types. Conversely, if the distributional properties of a linguistic item or construction have
not changed over the time span covered by the corpus, we should not expect to witness any
changes in the certainty by which the model classifies tokens to its true usage type category
(i.e. entropy).</p>
        <p>Of course, it is difficult to assess the efectiveness of this method if we do not know whether
[BE about] has actually undergone any functional-semantic changes in the period covered by
COHA. As a means of assessment, then, I will first formulate the expected results for each of
the construction’s three main usage types based on prior literature.</p>
        <sec id="sec-2-2-1">
          <title>2.2.1. futurate [BE about]</title>
          <p>
            The first usage type under consideration is the futurate use of [BE about], illustrated in
examples (2)-(4). In present-day English reference grammars, the phrasal expression be about to
is commonly described as a so-called ’quasi-auxiliary’ which can be used to make a temporal
reference to the future [
            <xref ref-type="bibr" rid="ref28">28</xref>
            ]. As such, be about to can be considered near-synonymous to other
English (quasi-)auxiliaries such as will, shall, and be going to (see example (2)). Notably,
however the use be about to is commonly said to convey a strong sense of immediacy, and it
has been suggested that the construction may in fact have more affinity with “aspectualizing
expressions such as begin to/start to V than form expressing futurity” [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ].
          </p>
          <p>
            Sheen, yes THAT Charlie Sheen, is about to become the best-paid actor in a comedy on TV.
(2006, COHA)
Just as I am about to step into the shower, the phone rings. It is George Stephanopoulos.
(2002, COHA)
But when he saw the sneer on St. Exeter’s face, Logan knew things were about to get much
worse. (2004, COHA)
From the relatively sparse number of accounts on the diachronic development of [BE about], it
can be concluded that the construction had already grammaticalized into a marker of
(immediate) future by the 19th century [
            <xref ref-type="bibr" rid="ref16 ref23">23, 16</xref>
            ] (the suggested time of the first, full-blown futurate uses
ranging between the late 15th or 16th century [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] and the late 18th century [31], when the
construction started occurring with, for example, inanimate subjects and non-intention verbs,
as in (4)). As such, the semantic and distributional changes typical of a grammaticalizing
construction (e.g. bleaching [30] or host-class expansion [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]) most likely pre-date COHA.
          </p>
          <p>However, the [BE about] future did undergo a distributional change: in Present-day English,
the phrase almost exclusively occurs with a to-infinitive complement clause, whereas 19th
century and early 20th century texts also contain a variant with an ing-clause complement (as
in (5)-(6)).</p>
          <p>I really thought all my bones were disjointed, and that my soul was about taking a last farewell
of my poor body. (1812, COHA)
I was trying to sleep, and just as I was about succeeding Henderson called out: ‘[...]’. (1902,
COHA).</p>
          <p>If accurate, the method should detect that the distributional properties of futurate [BE about]
have narrowed slightly over the course of the 19-20th century.
2.2.2. approximative about
[BE about] often occurs as a marker of approximation. Note that the use of approximative about
is not restricted to contexts with the verb be. The current sample therefore only represents a
subset of the possible occurrences of approximative about, some examples of which are listed
in (7)-(9):</p>
          <p>When he was about 12, his parents left him and his siblings at an orphanage for five months.
(2006, COHA)
Though you hate to say anyone is recession-proof, U2 is about as close to that as you can get.
(2009, COHA)</p>
          <p>
            It was about at that time that Takemore disappeared from the township too. (2001, COHA)
With respect to its diachronic development, it has been shown that the spatial preposition
about approximative use of about – much like the near-synonymous use of around and various
other adpositions in other languages – developed into “approximative qualifiers of numerical
expressions and other amount expressions” [
            <xref ref-type="bibr" rid="ref27">27</xref>
            ] sometimes called “rounders” [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ]. This process
is shown to have started around the beginning of the Middle English period (1250-1500) [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ],
and the establishment of approximative about in the wide range of contexts in which it can
presently occur pre-dates COHA with a very large margin. The inclusion of approximative
[BE about] is therefore not so much motivated by the fact that it has been stable over the
course of the 19th–20th century. As such, no changes in classification uncertainty should be
attested.
          </p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.2.3. descriptive [BE about]</title>
          <p>In the third and final category, we find a group of what we could call ‘descriptive’ uses of [BE
about], in which the phrase [BE about] can be roughly paraphrased as ‘regards’, ‘is (primarily)
concerned with’. Its occurrences commonly involve clarifications of why situations are occurring
(e.g. (10)), as well as descriptions of the theme, topic (or plot) of conversations, books, and
iflms (e.g. ( 11)-(12)):
(10) ... he wonders what the fighting is about and and who is fighting whom. Is it North against</p>
          <p>South again? (2008, COHA)
(11) The latest news is about Amanda! Haven’t you heard? (2000, COHA)
(12) Mars and Venus Collide is not just about men understanding women. It is also about women
understanding themselves and learning how to ask efectively for the support they need. (2008,
COHA)
Unlike with the previous two usage types, the number of historical and diachronic accounts
that treat the descriptive use of [BE about] are not sparse, but virtually non-existent. Still,
a quick scan of the dated examples listed in the Oxford English Dictionary suggests that the
use of [BE (all) about] with an animate subject (e.g. a person, organization, or company)
to describe what the subject is ‘primarily concerned with’ or ‘fond of’ constitutes a relatively
recent (i.e. mid-20th century) phenomenon (e.g. (13)-(14)).
(13) ... give him your authenticity spiel and how radio should be all about the music. (2009, COHA)
(14) I’m all about the blindfold. There’s something intensely sensual about not knowing where
you’re going (2005, COHA).</p>
          <p>In other words, the types of elements that can occur in the subject slot in the descriptive
subtype of [BE about] appear to have expanded during the time period captured by COHA,
which may indicate that the descriptive use of [BE about] came to cover a broader semantic
range. Given that the distributional properties of the descriptive use in the early 19th century
were quite diferent from the present-day, we would again expect to find a shift classification
uncertainty.</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>2.2.4. Summary: expected shifts</title>
          <p>In sum, the (in)stability of the classification entropy should reflect the following: The
distributional properties of futurate [BE about] have changed slightly during the 20th century. Having
lost the ability to occur with ing-complements, it seems that the futurate use has narrowed. As
with the futurate use, the distributional properties of the descriptive use have changed. More
specifically, the descriptive use has expanded or broadened. By contrast, the distributional
properties of the approximative use have remained stable.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results: detecting changes in [BE about]</title>
      <p>
        As a starting point, it is worth considering real
the distributional range of the [BE about] con- 1.4 fadupetpsucrroreixptive
struction as a whole. It appears that entropy 1.2
1h.a0s9)in,daenedd, ianscrseuacshe,d iotvceorutlidmbee( fflargogmed0.a7s6 ato H1.0
case of semantic broadening [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. To under- 0.8
stand what has been captured here precisely, 0.6
iotf itshweodritfehrecnotnussidaegreintygptehsearcerloastsivteimfree:quweitnhcy 0.4 1800 1825 1850 1875 1b9i0n0 1925 1950 1975 2000
the descriptive category growing more
frequent, there are efectively there major usage Figure 3: Entropy based on logistic regression
clastypes by the end of the 20th century (whereas sifier, fitted on the embeddings created for a training
there were only two at the start of the 19th set of 952 labelled present-day tokens, and applied to a
century). test set of 200 examples from each decade in COHA.
      </p>
      <p>What such a general test does not reveal is For each test token and each label, the conditional
whether there have been any changes within probability has been computed to assess the
uncerthese major usage types. For instance, there tainty of the classifier in labelling the unseen
examis no immediate indication that anything has ples. The resulting conditional probability over each
changed about the descriptive use besides its leanbterolpfyorsceoarceh, wtehsitchtogkreanduisalltyhednecsruemasmesarfiozredthien
daenoverall frequency. scriptive use, and, to a lesser extent, for the futurate
use.</p>
      <sec id="sec-3-1">
        <title>3.1. changing usage types</title>
        <p>To assess whether embeddings can be used to detect changes within the three major usage
types, one could explore treating the semantic change detection task as a classification
experiment. As explained in Section 2.2, the goal of the classification entropy test is to establish
whether any of the usage categories of [BE about] has changed compared to a present-day
reference point. We expect to find Entropy Diferences over time in two of the three usage
types: the descriptive use, and, to a weaker extent, the futurate use. This expectation appears
to be borne out (Figure 3).</p>
        <p>A fair point that can be raised is that the attested decrease in uncertainty does not in fact
signal that the usage of a construction has changed, but rather reflects a decrease in the extent
to which embeddings created by unsupervised BERT (trained on predominantly Present-day
English data) as the linguistic material more generally becomes older (and, consequently, less
familiar). Still, the latter explanation loses at least some of its plausibility when considering the
expected, relative stability of the approximative use. It is furthermore reassuring that the slight
80
60
40
20
0
20
40
60
2000
1975
1950
1925
1900
1875
1850
1825
(a) futurate
(b) descriptive
increase in classifier certainty appears to coincide with the decline of [BE about Ving], which
renders futurate [BE about] more uniform and, consequently, more unambiguously recognizable.</p>
        <p>
          The suggested shifts are also evident from the visualizations in Figure 4a and 4b, which
present what one could call a ‘time-sensitive t-SNE’ representation of the futurate and
descriptive use of [BE about]. If one wishes to avoid the use of an indirect, present-day reference
point to query a corpus for potential constructional changes, it would of course also be possible
to examine token groupings (as apparent from a time-sensitive version of t-SNE [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] or another
type of dimension reduction technique) that appear to be specific for a particular time-period.
        </p>
        <p>Using Figure 4, for example, one can examine the token groupings which are dated towards
the beginning of the corpus (representing groupings of the [BE about Ving] pattern), or the
smaller, markedly recent grouping top center left (containing solely negative uses, expressing
absence of intent, e.g. Wang’s not about to forgive you (2007, COHA)). Note that the relative
size of the time-specific token groupings discussed here may afect the extent to which summary
statistics (such as the average pairwise distance between tokens or the silhouette score of
clusters over time) capture their emergence or disappearance.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and Conclusion</title>
      <p>All in all, the present assessment of the use BERT embeddings and uncertainty measures to
detect functional-semantic change in grammatical constructions seems largely positive. However,
there are a number of potential pitfalls that must be addressed.</p>
      <p>A first, smaller point that could be raised concerns the attested change in the [BE about]
futurate. One may argue that this is, in fact, a formal rather than a functional-semantic
change. What we witness is a reduction in the variability of (near-synonymous) complement
clause types following the futurate, but this distributional change does not (straightforwardly)
mark a shift in the futurate’s meaning. Still, what has been detected is of value to linguists,
as the model has picked up a reduction in semasiological variation between a current and a
currently obsolete complementation pattern.</p>
      <p>Second, one may, as pointed out earlier, not be fully convinced that uncertainty measures
determined by an unsupervised, present-day neural language model are the most reliable
measure to detect semantic shift. First, it is unclear to what extent the fact that BERT has
been pre-trained on present-day English material afects its performance with respect to the
gradually aging data. Second, if tokens from a decade d are flagged as yielding a high degree
of classifier uncertainty with respect to the reference point r, one should not be too quick to
assume semantic change proper has taken place: in fact, this may be due to a diference in
how the tokens are distributed across (sub)genres in d and r. In this study, I minimized both
problems somewhat, but the concerns are still legitimate. With respect to genre variation,
it helps to work with a carefully balanced corpus such as COHA (if available), or to try and
incorporate meta-information on (sub)genre in the model [e.g. 26]. With respect to the
possible ‘present-day bias’ of the pre-trained model, it is reassuring to see that attest stability
with usage types that are not known to have changed, and that similar conclusions on possible
distributional shifts can be arrived at by examining time-sensitive t-SNE plots.</p>
      <p>
        However, it is important to stress that the explored approach solely considers uncertainty
with respect to the present-day reference point: while the classification entropy test successfully
pointed out that the descriptive use of [BE about] had undergone some changes, the decrease
in entropy cannot be equated to semantic narrowing. Instead, given that the descriptive
use of [BE about] rather seems to have shifted and broadened, the classification entropy test
merely indicates that the tokens have become more like the present-day examples and linguistic
material the model has been trained with. Second, the explored approach is limited in the
sense that the number (and nature) of usage types is imposed anachronistically to
non-presentday data. Since the procedure relies on a single reference point, it will not straightforwardly
lfag any usage types that are absent in the training set, and it may erroneously impose the
pre-defined category labels onto tokens representing obsolete usages. A further indication that
the method may be problematically biased towards present-day language can be found when
the model’s actual classification errors are considered in more detail. In the case of the [BE
about] futurate, the overall classification accuracy is remarkably high at 0.985, with only 22 of
the 1517 examples not being recognized as futurates. On closer inspection of those 22 mistakes,
it appears that 19 of them involve the now obsolete [BE about Ving] pattern. Given that there
are 102 examples of [BE about Ving] in the test set, this amounts to an error rate of 18.6%.
Furthermore, the mistakes are of the type illustrated in (15), where a Present-day descriptive
interpretation (i.e. Napoleon was fond of going to war with England) is erroneously imposed:
(15) Jeferson obtained the consent of Congress to make an efort to buy New Orleans and West
Florida, and sent Monroe to aid our minister in France in making the purchase. When the ofer
was made, Napoleon was about going to war with England, and, wanting money very much, he
in turn ofered to sell the whole province to the United States. (1897, COHA)
Because the use of data-driven, automated methods of semantic annotation and analysis is
appealing to researchers precisely because it could help avoid such anachronistic interpretations
of historical language [e.g. 29], it is of course unfortunate that they still occur at a reasonably
high rate. Yet, it should still be acknowledged that the very fact that word and phrase
embeddings created by BERT did succeed in recognizing diferent grammatical usage types in
Present-day language inspires hope that these problems can be tackled when models such as
these are trained on contemporary linguistic material and (following proposals such as [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ])
made dynamic.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>I am grateful to Folgert Karsdorp for his advice on how to implement parts of the analysis.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] “about, adv.,
          <source>prep.1</source>
          ,
          <string-name>
            <surname>adj</surname>
            <given-names>.</given-names>
          </string-name>
          ,
          <source>and int.” In: Oxford English Dictionary Online</source>
          . Oxford University Press,
          <year>1990</year>
          . url: oed.com/view/Entry/527.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Baroni</surname>
          </string-name>
          , G. Dinu, and
          <string-name>
            <surname>G. Kruszewski.</surname>
          </string-name>
          “
          <article-title>Don't count, predict! A systematic comparison of context-counting vs. context-predicting semantic vectors”. en</article-title>
          . In:
          <article-title>Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</article-title>
          . Baltimore, Maryland: Association for Computational Linguistics,
          <year>2014</year>
          , pp.
          <fpage>238</fpage>
          -
          <lpage>247</lpage>
          . doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>P14</fpage>
          -1023. url: http://aclweb.org/anthology/P14-1023 (visited on 01/26/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>G. Boleda.</surname>
          </string-name>
          “
          <article-title>Distributional Semantics and Linguistic Theory”</article-title>
          . en.
          <source>In: Annual Review of Linguistics 6</source>
          .1 (
          <issue>Jan</issue>
          .
          <year>2020</year>
          ). arXiv:
          <year>1905</year>
          .
          <year>01896</year>
          , pp.
          <fpage>213</fpage>
          -
          <lpage>234</lpage>
          . issn:
          <fpage>2333</fpage>
          -
          <lpage>9683</lpage>
          ,
          <fpage>2333</fpage>
          -
          <lpage>9691</lpage>
          . doi:
          <volume>10</volume>
          .1146/annurev-linguistics-
          <volume>011619</volume>
          -030303. url: http://arxiv.org/abs/
          <year>1905</year>
          .01896 (visited on 04/15/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Budts</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Petré</surname>
          </string-name>
          . “
          <article-title>Putting connections centre stage in diachronic construction grammar”</article-title>
          . In:
          <article-title>Nodes and Networks in Diachronic Construction Grammar</article-title>
          . Ed. by
          <string-name>
            <given-names>L.</given-names>
            <surname>Sommerer</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          . Amsterdam: John Benjamins,
          <year>2020</year>
          , pp.
          <fpage>317</fpage>
          -
          <lpage>352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Davies</surname>
          </string-name>
          .
          <article-title>Corpus of Historical American English (COHA)</article-title>
          .
          <source>Version V1</source>
          .
          <year>2015</year>
          . doi:
          <volume>10</volume>
          .7910/DVN/8SRSYK. url: https://doi.org/10.7910/DVN/8SRSYK.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>De Smet</surname>
          </string-name>
          . “
          <article-title>The course of actualization”. en</article-title>
          .
          <source>In: Language 88.3</source>
          (
          <issue>2012</issue>
          ), pp.
          <fpage>601</fpage>
          -
          <lpage>633</lpage>
          . issn:
          <fpage>1535</fpage>
          -
          <lpage>0665</lpage>
          . doi:
          <volume>10</volume>
          .1353/lan.
          <year>2012</year>
          .
          <volume>0056</volume>
          . url: http://muse.jhu.edu/content/crossre f/journals/language/v088/88.3.de-smet.
          <source>html (visited on 01/26/</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>G. Desagulier.</surname>
          </string-name>
          “
          <article-title>Can word vectors help corpus linguists?” en</article-title>
          .
          <source>In: Studia Neophilologica 91.2 (May</source>
          <year>2019</year>
          ), pp.
          <fpage>219</fpage>
          -
          <lpage>240</lpage>
          . issn:
          <fpage>0039</fpage>
          -
          <lpage>3274</lpage>
          ,
          <fpage>1651</fpage>
          -
          <lpage>2308</lpage>
          . doi:
          <volume>10</volume>
          .1080/00393274.
          <year>2019</year>
          .
          <volume>1616220</volume>
          . url: https://www.tandfonline.com/doi/full/10.1080/00393274.
          <year>2019</year>
          .
          <volume>1616220</volume>
          (visited on 05/17/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          et al. “
          <article-title>BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”. en</article-title>
          .
          <source>In: Proceedings of NAACL-HLT</source>
          <year>2019</year>
          . Minneapolis, Minnesota,
          <year>June 2019</year>
          , pp.
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Dubossarsky</surname>
          </string-name>
          et al. “
          <string-name>
            <surname>Time-Out</surname>
          </string-name>
          :
          <article-title>Temporal Referencing for Robust Modeling of Lexical Semantic Change”</article-title>
          . In:
          <article-title>Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</article-title>
          . Florence, Italy: Association for Computational Linguistics,
          <year>July 2019</year>
          , pp.
          <fpage>457</fpage>
          -
          <lpage>470</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -1044. url: https://www.aclweb.org/ant hology/P19-1044 (visited on 06/28/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Eger</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Mehler</surname>
          </string-name>
          . “
          <article-title>On the Linearity of Semantic Change: Investigating Meaning Variation via Dynamic Graph Models”</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>2</volume>
          : Short Papers). Berlin, Germany: Association for Computational Linguistics, Aug.
          <year>2016</year>
          , pp.
          <fpage>52</fpage>
          -
          <lpage>58</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/P1 6-
          <fpage>2009</fpage>
          . url: https://www.aclweb.org/anthology/P16-2009
          <source>(visited on 07/21/</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Giulianelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Del Tredici</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>R.</given-names>
            <surname>Fernández</surname>
          </string-name>
          . “
          <article-title>Analysing Lexical Semantic Change with Contextualised Word Representations”</article-title>
          . In: arXiv:
          <year>2004</year>
          .14118 [cs] (
          <year>Apr</year>
          .
          <year>2020</year>
          ). arXiv:
          <year>2004</year>
          .14118. url: http://arxiv.org/abs/
          <year>2004</year>
          .14118 (visited on 06/28/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Greenhill</surname>
          </string-name>
          et al. “
          <article-title>Evolutionary dynamics of language systems”</article-title>
          . en.
          <source>In: Proceedings of the National Academy of Sciences 114.42 (Oct</source>
          .
          <year>2017</year>
          ),
          <fpage>E8822</fpage>
          -
          <lpage>E8829</lpage>
          . issn:
          <fpage>0027</fpage>
          -
          <lpage>8424</lpage>
          ,
          <fpage>1091</fpage>
          -
          <lpage>6490</lpage>
          . doi:
          <volume>10</volume>
          .1073/pnas.1700388114. url: http://www.pnas.org/lookup/doi/10.1 073/pnas.1700388114 (visited on 07/21/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>W. L.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          . “
          <article-title>Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change”</article-title>
          . In: arXiv:
          <fpage>1605</fpage>
          .09096 [cs] (Oct.
          <year>2018</year>
          ). arXiv:
          <volume>1605</volume>
          .09096. url: http://arxiv.org/abs/1605.09096 (visited on 06/28/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hilpert</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. Correia</given-names>
            <surname>Saavedra</surname>
          </string-name>
          . “
          <article-title>Using token-based semantic vector spaces for corpus-linguistic analyses: From practical applications to tests of theoretical claims”. en</article-title>
          .
          <source>In: Corpus Linguistics and Linguistic Theory</source>
          <volume>0</volume>
          .0 (
          <issue>Sept</issue>
          .
          <year>2017</year>
          ). issn:
          <fpage>1613</fpage>
          -
          <lpage>7027</lpage>
          ,
          <fpage>1613</fpage>
          -
          <lpage>7035</lpage>
          . doi:
          <volume>10</volume>
          .1515/cllt-2017-0009. url: http://www.degruyter.com/view/j/cllt.ahead-o f-print/cllt-2017-0009/cllt-2017
          <source>-0009.xml (visited on 05/16/</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N. P.</given-names>
            <surname>Himmelmann</surname>
          </string-name>
          . “
          <article-title>Lexicalization and grammaticization: opposite or orthogonal?”” en</article-title>
          . In:
          <article-title>What Makes Grammaticalization: A Look from Its Components</article-title>
          and
          <string-name>
            <given-names>Its</given-names>
            <surname>Fringes</surname>
          </string-name>
          . Ed. by
          <string-name>
            <given-names>W.</given-names>
            <surname>Bisang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. P.</given-names>
            <surname>Himmelmann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Wiemer</surname>
          </string-name>
          . Berlin: Mouton de Gruyter,
          <year>2004</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>S. Höche. “</surname>
          </string-name>
          <article-title>I am about to die vs. I am going to die: A usage-based comparison between two future-indicating constructions”</article-title>
          . In: Converging Evidence:
          <article-title>Methodological and Theoretical Issues for Linguistic Research</article-title>
          . Ed. by
          <string-name>
            <given-names>D.</given-names>
            <surname>Schönefeld</surname>
          </string-name>
          . Amsterdam: John Benjamins Publishing Company,
          <year>2011</year>
          , pp.
          <fpage>115</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B.</given-names>
            <surname>Jirsa</surname>
          </string-name>
          . “
          <article-title>Synchronic Applications for Diachronic Syntax: The Grammaticalization of to be about to in English”. en</article-title>
          . In: Colorado Research in Linguistics 15 (
          <year>1997</year>
          ). issn:
          <fpage>1937</fpage>
          -
          <lpage>7029</lpage>
          . doi:
          <volume>10</volume>
          .25810/5h5t-xg32. url: https://journals.colorado.edu/index.php/cril/artic le/view/231 (visited on 07/21/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          et al. “
          <article-title>Temporal Analysis of Language through Neural Language Models”</article-title>
          . en.
          <source>In: Proceedings of the ACL 2014 Workshop on Language Technologies and Computational Social Science. Baltimore</source>
          ,
          <string-name>
            <surname>MD</surname>
          </string-name>
          , USA: Association for Computational Linguistics,
          <year>2014</year>
          , pp.
          <fpage>61</fpage>
          -
          <lpage>65</lpage>
          . doi:
          <volume>10</volume>
          .3115/v1/
          <fpage>W14</fpage>
          -2517. url: http://aclweb.org/anthology/W14-2517 (visited on 06/28/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kutuzov</surname>
          </string-name>
          et al. “
          <article-title>Diachronic word embeddings and semantic shifts: a survey”</article-title>
          .
          <source>In: Proceedings of the 27th International Conference on Computational Linguistics. Santa Fe</source>
          , New Mexico, USA: Association for Computational Linguistics, Aug.
          <year>2018</year>
          , pp.
          <fpage>1384</fpage>
          -
          <lpage>1397</lpage>
          . url: https://www.aclweb.org/anthology/C18-1117 (visited on 07/20/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L. van der Maaten and G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          . “
          <article-title>Visualizing Data using t-SNE”</article-title>
          .
          <source>In: Journal of Machine Learning Research</source>
          <volume>9</volume>
          (
          <year>2008</year>
          ), pp.
          <fpage>2579</fpage>
          -
          <lpage>2605</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>C.</given-names>
            <surname>Mair</surname>
          </string-name>
          and
          <string-name>
            <surname>G. Leech.</surname>
          </string-name>
          “
          <article-title>Current Changes in English Syntax”</article-title>
          . en. In: The Handbook of English Linguistics. Ed. by
          <string-name>
            <given-names>B.</given-names>
            <surname>Aarts</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>McMahon. Malden</surname>
          </string-name>
          , MA, USA: Blackwell Publishing, Jan.
          <year>2006</year>
          , pp.
          <fpage>318</fpage>
          -
          <lpage>342</lpage>
          . doi:
          <volume>10</volume>
          .1002/9780470753002.ch14. url: http://doi .wiley.
          <source>com/10.1002/9780470753002.ch14 (visited on 09/19/</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>T.</given-names>
            <surname>McEnery</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Hardie</surname>
          </string-name>
          .
          <source>The History of Corpus Linguistics. en</source>
          . Oxford University Press, Mar.
          <year>2013</year>
          . doi:
          <volume>10</volume>
          .1093/oxfordhb/9780199585847.013.0034. url: http://oxfordh andbooks.
          <source>com/view/10</source>
          .1093/oxfordhb/9780199585847.001.0001/oxfordhb-97801995858
          <fpage>47</fpage>
          -e-
          <volume>34</volume>
          (visited on 06/29/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Mee</surname>
          </string-name>
          . “
          <article-title>The evolution of constructions: The case of be about to”. en</article-title>
          . MA dissertation. University of New Mexico,
          <year>2013</year>
          , p.
          <fpage>138</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>W.</given-names>
            <surname>Mihatsch</surname>
          </string-name>
          . “
          <article-title>The Diachrony of Rounders and Adaptors: Approximation and Unidirectional Change”</article-title>
          . en. In: New Approaches to Hedging. Ed. by G. Kaltenböck,
          <string-name>
            <given-names>W.</given-names>
            <surname>Mihatsch</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. Schneider. BRILL</surname>
          </string-name>
          , Jan.
          <year>2010</year>
          , pp.
          <fpage>93</fpage>
          -
          <lpage>122</lpage>
          . isbn:
          <fpage>978</fpage>
          -
          <lpage>90</lpage>
          -04-25324-7. doi:
          <volume>10</volume>
          .1163 /9789004253247_007. url: https://brill.com/view/book/edcoll/9789004253247/B97890 04253247-
          <fpage>s007</fpage>
          .
          <source>xml (visited on 07/20/</source>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          et al. “
          <article-title>Scikit-learn: Machine learning in Python”</article-title>
          .
          <source>In: Journal of machine learning research 12</source>
          .
          <string-name>
            <surname>Oct</surname>
          </string-name>
          (
          <year>2011</year>
          ), pp.
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>V.</given-names>
            <surname>Perrone</surname>
          </string-name>
          et al. “
          <article-title>GASC: Genre-Aware Semantic Change for Ancient Greek”</article-title>
          .
          <source>In: Proceedings of the 1st International Workshop on Computational Approaches to Historical Language Change</source>
          (
          <year>2019</year>
          ). arXiv:
          <year>1903</year>
          .05587, pp.
          <fpage>56</fpage>
          -
          <lpage>66</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W19</fpage>
          - 4707. url: http://arxiv.org/abs/
          <year>1903</year>
          .05587 (visited on 09/19/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>F.</given-names>
            <surname>Plank</surname>
          </string-name>
          . “
          <article-title>Inevitable reanalysis: From local adpositions to approximative adnumerals, in German and wherever”. en</article-title>
          .
          <source>In: Studies in Language 28.1</source>
          (
          <issue>2004</issue>
          ), pp.
          <fpage>165</fpage>
          -
          <lpage>201</lpage>
          . issn:
          <fpage>0378</fpage>
          -
          <lpage>4177</lpage>
          ,
          <fpage>1569</fpage>
          -
          <lpage>9978</lpage>
          . doi:
          <volume>10</volume>
          .1075/sl.28.1.07pla. url: http://www.jbe-platform.com/c ontent/journals/10.1075/sl.28.1.07pla (visited on 07/20/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>R.</given-names>
            <surname>Quirk</surname>
          </string-name>
          et al.
          <article-title>A Comprehensive Grammar of the English Language</article-title>
          . London: Longman,
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sagi</surname>
          </string-name>
          , S. Kaufmann, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Clark</surname>
          </string-name>
          . “
          <article-title>Tracing semantic change with Latent Semantic Analysis”. en</article-title>
          . In: Current Methods in Historical Semantics. Ed. by
          <string-name>
            <given-names>K.</given-names>
            <surname>Allan</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Robinson</surname>
          </string-name>
          . Berlin, Boston: DE GRUYTER,
          <year>Jan</year>
          .
          <year>2011</year>
          , pp.
          <fpage>161</fpage>
          -
          <lpage>183</lpage>
          . isbn:
          <fpage>978</fpage>
          -3-
          <fpage>11</fpage>
          - 025290-3. doi:
          <volume>10</volume>
          .1515/9783110252903.161. url: https://www.degruyter.com/view/boo ks/9783110252903/9783110252903.161/9783110252903.161.xml (visited on 01/26/
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>