<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detecting linguistic change based on word co-occurrence paterns</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carmen Klaussner</string-name>
          <email>klaussnc@tcd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carl Vogel</string-name>
          <email>vogel@tcd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arnab Bhattacharya</string-name>
          <email>bhattaca@tcd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Trinity College Dublin</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Diachronic linguistic analysis focuses on detecting elements of language change over time. This change can take diferent forms, as for instance certain words could show a slow increase or decrease in frequency over time as these become more popular or obsolete. We are interested in sudden change in words that are attested for in every time slice of the overall examined time period. In particular, we are trying to relate the change in frequency of words that are always there to words that emerge around certain points in time and only remain frequent for shorter periods of time, suggesting they are more prone to sudden changes in popular topics or could be influenced by historical events. This addresses the question of how the more regular word expressions' frequencies are influenced by new clusters of words appearing and disappearing. Although there might be links to collocation analysis, words that occur frequently next to each other are not of primary interest here, but rather words that are conceptually related, where one is causing or afecting the frequency of the other, which causes them vary in a similar fashion. We use statistical change point analysis for identification of significant change over time and seek to validate our findings by randomly extracting example sentences from the data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Findings from temporal studies ofer an important source of
enrichment and validation for non-temporal studies, especially when one
needs to disentangle (general) temporal efects from non-temporal
ones, as for instance in the realm of stylometry and authorship
attribution [2]. Change in linguistic variables can occur in
diferent shapes and forms: slow gradual change as opposed to sudden
and abrupt, short as well as long-term efects. Diferences could be
rooted in levels of linguistic abstraction, as for instance individual
words are likely to show more variation over time than entire word
classes, where smaller fluctuations would be averaged and only
larger trends pervading the entire group would be more easily
discernible. Features are usually classified along one or more diferent
dimensions, such as membership of either an open-or closed-class
or according to the frequency strata, i.e. frequent, medium-frequent
and rare, they belong to. However, even though these
classifications are often represented as categories suggesting there is a clear
boundary between, for instance what is frequent and infrequent, a
continuous representation, especially considering temporal efects
would occasionally seem more suited. This is especially true in the
realm of diachronic analyses, the analysis of texts over time, as most
items might appear across diferent frequency strata depending on
the exact time period examined.</p>
      <p>We aim to show that even the open-class/closed-class view of
features might be insuficient when observed through the lens of
temporal text representation. While something might bear the label
of common noun, it could in fact be closer to a function word,
behaving and being afected by similar factors, such as regular
occurrence in diferent contexts.</p>
      <p>In this work, we consider the analysis of the more regular and
possibly also frequent items, in particular those appearing in all
years of the time period examined. Typically, these features are
more general in meaning (e.g. temporal expressions) rendering them
suitable for a variety of language contexts other than strongly
topicrelated words, such as hurricane or computer. Our main hypothesis
is that these items change through other less frequent items that
are more prone to topic change over time, such as concepts relating
to the outbreak of a war or natural disaster. The type of change we
are looking for is a change in mean, whereby a feature changes its
relative frequency fairly abruptly at time t , rising or falling to a new
level and remaining there for at least some time. If one compares
the mean over the samples before time t to the mean taken over
the samples after time t , one obtains significantly diferent means.
This type of research has to be distinguished from two related areas
of research, i.e. semantic change in the form of neologisms and
collocation analysis. Semantic change analysis is diferent in that it
considers cases whereby a word acquires a new sense and possibly
also a second part-of-speech class and could subsequently be used
in diferent syntactic contexts, whereas here we consider changes
of word frequencies and their possible non-semantic change related
causes, as for instance a particular temporal expression used for
contrasting diferent situations, e.g. ‘If I had only known then, what I
know now.’ Also, conceptually, these regular and irregular appearing
words could be relatable through collocations or otherwise longer
n-gram sequences. While the method presented here could be used
for their detection as well, it is not limited to relationships between
words that occur close to each other, but also words or expressions
that only share a conceptual rather than spatial relationship, such
as the first ( ‘Detecting’) and last (‘patterns’) word of the title of
this work, that as 8-grams are rarely computed, is less likely to
be captured by collocation analysis, while the terms are clearly
conceptually related.</p>
      <p>The remainder of this work is organized as follows: section 2
discusses related research in the field; section 3 provides
information about the data and pre-processing steps; section 4 presents
the methods we employed. Before moving on to the experiments
in section 6, section 5 considers trends in larger groups of word
expressions to anchor our findings based on individual words.
Section 6 then presents our change point experiments, and empirical
validation of these through the actual data. We discuss these results
in section 7 and conclude the work in section 8.
2</p>
    </sec>
    <sec id="sec-2">
      <title>PREVIOUS RESEARCH</title>
      <p>Diferent areas of linguistic research consider the change of broader
categories of words, such as frequency efects in syntax [ 1], largely
distinguishing between type and token frequency of a particular
variable or category. Bybee and Thompson [1] discuss three
frequency efects that are important not only in shaping phonology
and morphology, but also syntax; two efects are caused by high
token frequency, which have adverse tendencies that can only be
explained by considering the influence of the third frequency efect
of high type frequency. A high token frequency of an item
promotes its reduction, as visible in conventionalized contractions in
English (I’m, can’t). In contrast, the ‘Conserving Efect’ is visible
with high token items, where the more the form is used the more it
is strengthened, compare normalization of the English past tense
of ‘weep’ from wept to weeped, compared to high frequency items,
such as ‘sleep’ (slept). A syntactic example of this is the fact that
pronouns, although derived from full noun phrases show much
more conservative behaviour (e.g. case marking) due to their higher
frequency. The type of change that is resisted in the high token
frequency items is change on the basis of combinatorial patterns or
constructions that are productive. “The more lexical items that are
heard in a certain position in a construction, the less likely it is that
the construction will be associated with a particular lexical item”
[1, p.384]. This is observable in the ditransitive construction, which
is only acceptable with very specific lexical verbs of high frequency,
compare: “ He told the woman the news” vs. “He whispered the
woman the news” [1, p.385], where the verb tell is a lot more
frequent than whispered. To a limited extent, this is also productive in
that the construction can apply to a few new high frequency verbs,
such as e-mailed or telephoned.</p>
      <p>Hamilton et al. [4] consider the function aspect of diachronic
change by taking a closer look at global and local shifts in a word’s
distributional semantics in historical texts from English, French and
German.1 For the local or cultural shifts, they use a local
neighbourhood measure and for the global measure they compute the cosine
distance between two word vectors capturing the co-occurrence
statistics at consecutive time points t and t+1. Based on previous
results in the literature, they predict that nouns are more likely to
undergo change because of cultural shifts, whereas verbs are more
likely to change because of regular semantic change. Across all
languages as predicted, the local neighbourhood measure assigns
higher rates of semantic change to nouns than verbs with the
opposite applying to the global measure. This also remains the case,
when adverbs and adjectives are included among the verbs,
supporting previous results in the literature suggesting that adverbial
and adjectival modifiers are often the target of regular or global
linguistic change [4].</p>
      <p>The research presented by Kulkarni et al. [7] considers change
point analysis in the context of investigating statistically significant
shifts of semantic change. They consider three diferent approaches,
1Local or cultural shifts are deemed less regular and stable than global shifts, as they
are caused by more changeable factors, such as new technologies, whereas global
shifts are associated to regular semantic change, such as grammaticalization.
one frequency based, whereby sudden changes in word usage are
captured. The second one involves a syntactic time-series analysis,
analyzing word’s part-of-speech tag distributions and finally they
construct a distributional time-series by considering contextual
cues from word co-occurrence statistics. Using human evaluators
to assess the performance of their models, they find the highest
amount of agreement between annotators and method with respect
to words that have undergone change is the distributional method
with c.53% average agreement compared to c.22% (syntactic) and
c.13% (frequency). Another change point oriented analysis was
addressed by Riba and Ginebra [9], which investigates a possible
change in authorship of Tirant lo Blanc, identifying a clear single
sudden change point that is supported by cluster analysis. Our
work examines possible changes in features that are both regular
in occurrence and highly frequent caused by features that are only
highly frequent over a short period of time, similar to semantic
cultural shifts, but for the diference that these words would not
necessarily take on a new meaning.
3</p>
    </sec>
    <sec id="sec-3">
      <title>DATA</title>
      <p>For this analysis, we consider a 100-year long extract from The
Corpus of Historical American English (COHA) [3].2 This is a 400-million
word corpus, which contains samples of American English from
1810–2009 balanced in size, genre and sub-genre in each decade
(1000–2500 files each). It therefore contains balanced language
samples from fiction , popular magazines, newspapers and non-fiction
books, which are again balanced across sub-genre, such as drama
and poetry.3</p>
      <p>For this study, we selected all data from the years of 1880-1979
covering all genre of news, magazine, fiction and non-fiction . For
most of the experiments, we only use the news section of the data,
as it is most likely to contain the types of change we are targeting,
though occasionally comparing to the other three genre. In order
to arrive at a relative frequency count for each feature, we combine
the individual files on a per year basis and relativize by the overall
token count for that year. 4 As features, we consider the set of word
bigrams marked for syntactic context, e.g. the word like has diferent
meanings depending on its context. It can be used as both a verb and
a preposition, which should subsequently be treated as two separate
items. We chose bigram size as it provides more context and is richer
in meaning allowing us to discern more specific items of change
than with unigram size. We decided against analyzing items of
higher rank and abstraction, such as part-of-speech sequences as
these are more dificult to evaluate, while word sequences ofer
more possibilities for human evaluation.</p>
      <p>In order to extract part-of-speech (POS) features needed for
syntactic word features, we used the TreeTagger POS tagger [8, 10].
Our new syntactic word features were then created by using the tag
sequence as a sufix to the original word in context that gave rise
to it. Thus, “He likes her” becomes “he.PP likes.VBZ her.PP”.5 Items
2free version accessible on: http://corpus.byu.edu/coha –last verified July 2017.
3There is an excel file with a detailed list of sources available on:
http://corpus.byu.edu/coha/–last verified July 2017.
4In the case of higher sequence features, such as word bigrams the unigram token
count is replaced by the unique bigram token count.
5In this, the diference between the original word in context and the lemma of the
word would primarily be reflected in verbs.
1.0
0.5
s
ixa2 0.0
c
p
−0.5
−1.0
young.JJ man.NN
few.JJ moments.NNS
young.JJ lady.NN
good.JJ deal.NN
great.JJ deal.NN</p>
      <p>first.JJ time.NN
many.JJnpeexot.pJlJe.yNeNarS.NN
llaasstt..JJJJ wyeeaerk.N.NNN
are then joined to bigram sequences and each two syntactic word
sequence is relativized by the total number of bigram sequences in
that year. For this work, we are primarily interested in changes in
common nouns, requiring us to extract these from all other types.
We only retain adjective-noun or noun-noun combinations, as we
expect the other types that can occur in noun phrases, i.e.
determiners, proper nouns and pronouns to follow a diferent frequency
distribution that might introduce noise.</p>
    </sec>
    <sec id="sec-4">
      <title>4 METHODS</title>
      <p>In this section, we describe the methods used for initial detection of
interesting constant features, our data exploration and the change
point analysis.</p>
    </sec>
    <sec id="sec-5">
      <title>4.1 Detecting changing features</title>
      <p>As we are interested in change in variables appearing in all time
instances of a temporally-ordered data series, we consider only
those bigram adjective-noun/noun-noun types that appear in all
time slices and discard all others. Even when reducing the set of
features to these constant noun types, some 350 sequences remain for
examination. In order to discover interesting (and possibly related)
features more easily, we first order them according to mean relative
frequency and then use principal component analysis (PCA) on sets
of 50 bigram features, as we have found estimation and later
interpretation of components to be better, when the document-feature
ratio is in favour of more samples. PCA is an unsupervised
statistical technique to convert a set of possibly related variables to a new
uncorrelated representation or principal components. This type of
analysis groups features according to common variance patterns
and can help to detect features that vary in a similar way. The results
of running PCA on the 50 most frequent noun-noun/adjective-noun
sequences are 50 new components that group related features
together.6 A feature can be negatively or positively related to a new
component. The components themselves account for decreasing
proportions of variance, e.g. in this case the first component
accounts for 25% and the second component for 12% of the variance
with the rest being more broadly spread out. Inspection of first
principal component allows for discovery of the three highest
associated items: last week, last year and next year with very similar
weights: 0.259558004, 0.256880159 and 0.252325884 respectively
(Figure 1).
0.0000 ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ●</p>
    </sec>
    <sec id="sec-6">
      <title>4.2 Type-token analysis</title>
      <p>In order to understand the underlying development in the data
better, yielding a more informed analysis, we explore some methods
6For this experiment only, we took logarithms of relative frequencies before applying
PCA.
inspired from data analysis in the financial sector, Rate of Change
(RoC). The methods described here are used in section 5. In
particular, we are interested in what way diferent groups of features
change over time with respect to diferent quantities. For this, we
consider two basic linguistic measures, type-token ratio (TTR)
calculated as the number of types in a particular category divided
by all the number of tokens and a type-vs-all-types ratio (TYPR),
whereby we compare the number of types in a particular category
to all types. Further we combine these two ideas with the RoC in
order to gain insights in how types of features not only change
over time, but are related to the respective previous time instance.</p>
      <p>The log RoC is defined as, the value of the variable Vt today
divided by the value yesterday Vt −1 as shown in eq. 1. Thus, this
is showing how a particular value changes, for instance from one
year to the next.</p>
      <p>RoCln = ln</p>
      <p>Vt −1</p>
      <p>For most of the analysis, we use simple TTR and TYPR, except
for figure 5, for which a slightly modified version of the RoC is used.
Rather than having the current value at time t in the numerator,
we consider the set of linguistic types that two time points have
in common , as shown in Eq. 2. Thus, |typesyt | ∩ |typesyt−1 | refers
to the size of the group of features found in year yt and yt −1 with
respect to the number of tokens in yt −1. The second version, shown
in eq. 3 relativizes with respect to the number of total types in yt −1.
In the following, we refer to these to measures as TTR’ and TYPR’ to
distinguish these from static type-token ratios. As is common with
ifnancial data to achieve symmetry between decrease and increase,
we take natural logarithms.</p>
      <p>Vt !
TT Ry′t = ln
T Y PRy′t = ln
|typesyt | ∩ |typesyt−1 | !</p>
      <p>|tokensyt−1 |
|typesyt | ∩ |typesyt−1 | !</p>
      <p>|typesyt−1 |</p>
      <p>These two measures allow us to observe how the broader
categories behave with respect to feature types and what proportion
these take of either all types or tokens.
4.3</p>
    </sec>
    <sec id="sec-7">
      <title>Change-point Detection</title>
      <p>
        Change point analysis is the analysis of a time-series with the aim
to detect specific points t in time that separate the points before
and after it with respect to some criterion. More formally, aspects of
change point analysis can be defined as follows: given a time-series
{yt : t ∈ 1, ...n}, a change point occurs if there exists a time k,
where 1 ≤ k ≤ n − 1, such that the distributions of {y1...yk } and
{yk+1...yn } are diferent with respect to some criterion, i.e. change
in mean, change in regression or change in variance. For this analysis,
we are primarily interested in changes in mean as these would
signal a higher or lower average usage of a feature with respect to an
earlier time period, while for instance a change in variance would
indicate greater or lesser variability in how a feature is used. As we
are interested in long-term change, that lasts at least 10 years or so
we require a change point detection technique that is less volatile
to short-term fluctuations in the data. For our experiments, we
chose the approach by James et al. [5], originally used for breakout
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
0.0020
) 0.0015
%
(
n
o
i
tr 0.0010
o
p
o
r
P
0.0005
detection in cloud data in the presence of anomalies.7 The proposed
approach (‘E-divisive with Medians’(EDM)) is a non-parametric
technique using medians and estimating the statistical significance
of a change point through a permutation test. We found this
technique to return fewer change points that were more spread out than
distribution based change point methods, rendering it even more
desirable as our data is not always normally-distributed. We found
that using transformations occasionally smooths over interesting
developments making these less desirable to use in this context.
5
      </p>
    </sec>
    <sec id="sec-8">
      <title>DATA EXPLORATION</title>
      <p>In the following section, we look at unigram and bigram instances of
both function and content types to gain an intuition about general
trends in the data. We begin by exploring changes in type and token
relations in unigrams of both a function (determiner) and a content
(noun) category. For the determiner category, we considered both
instances corresponding to ⟨ DT⟩ and ⟨ WDT⟩, thus both the.DT and
which.WDT in contexts, such as ‘Which/The book...’.</p>
      <p>Figure 3 shows the number of determiner types with respect to
all types/all tokens over time. The top line shows a sharp decrease
in 1920 for the determiner vs. all types ratio, with this being also
visible but less pronounced for the type-to-token ratio. One aspect
that needs to be taken into account in this context is the influence
of types attested for in all time instances. We refer to these as the
‘universally’ constant types to distinguish between these and the
types that are ‘partially’ constant appearing in a few consecutive
years but not in all.8 The shortest span of constancy is two instances
(years), which we refer to here as ‘pairwise’ constancy. The concept
of constancy in itself has to be distinguished from possible
associated frequency distributions. A feature could appear in all time
instances and be therefore constant, but might vary considerably
with respect to its relative frequency. With respect to determiners,
the proportion of constant determiners of all occurring determiners
is relatively high indicating that other constancy types shared less
7This is implemented in the R package ecp [6].
8By ‘universally’ the span of our entire data set is meant rather than any data and
time space that could be examined in this way.
0.12
) 0.10
%
(
n
o
it
ro0.08
p
o
r
P
0.06
0.04
of the variation observed in figure 3. Thus, as would be expected
with a true function category, most of its types account for a high
proportion of all of its types as well as variation in frequency over
time. With the noun category, in this case only considering singular
and plural common nouns the situation is somewhat reversed. All
common noun types account for c. 38% of the entire token variety,
but only 0.05% of these are types that are universally constant.
Similarly, the non-universally constant types also account for most of
the tokens, indicating that variety rather than constancy of types is
predominant here. So we expect the proportions of universally
constant features to behave diferently for nouns and determiners. We
test this by subtracting the proportion of constant determiner types
of all types from the proportion of all determiner types of all types
for each year and compare these diferences using the Wilcoxon
signed rank test to the same quantities for nouns. The diference
in means over these yearly diferences is significant, meaning that
universally constant types behave diferently in each group.
Conversely, comparisons based on the complement of those universally
constant features is also significant. In terms of frequency changes,
one can observe a sharp drop in tokens (TTR) and a slightly more
temperate downward curve in types (TTYR) after 1920.
● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ●</p>
      <p>As a final step, we consider word bigram sequences, in
particular the group of bigrams comprised of either adjective + common
noun (plural/singular) or bigrams of only common nouns
(plural/singular). Figure 4 shows the adjective-noun bigram types with
respect to all types/all tokens. Interestingly, although the
common noun unigram types are decreasing over time, these bigram
types are increasing in proportion to both all types and tokens.
The universally constant features in this set are very limited in
this subgroup and largely all variation is accounted for by partially
constant features.</p>
      <p>Figure 5 shows the pairwise ‘appearing’ features, i.e. those
features found at time t , that are not at time t − 1 with respect to all
types (TYPR’)/tokens (TTR’) at time t − 1. As is observable, there
is a comparatively large increase of new features in 1920 and a
somewhat smaller increase again at c. 1939. This indicates that
there might be new concepts emerging for this feature type that
have not been there previously. The variation found with respect
● ●</p>
      <p>● ●
to these groups might only be found in this particular domain
of news articles, where language might be more variable than in
other domains, such as fiction or non-fiction . Thus, one would not
necessarily expect these findings to translate to other genres.</p>
    </sec>
    <sec id="sec-9">
      <title>6 CHANGE-POINT EXPERIMENTS</title>
      <p>Based on the exploratory analysis described in section 4.1, we chose
‘last year’ as our variable of interest and determine by change point
analysis whether there is an abrupt sort of change as opposed to
a more gradual trend. Figure 2 shows two adjective-noun phrase
bigrams, ‘last week’ and ‘last year’. Both are universally constant
over both the news and magazine corpus, but not constant for fiction
and non-fiction .</p>
    </sec>
    <sec id="sec-10">
      <title>6.1 Discovery of Related Variables</title>
      <p>Given our variable of interest, we seek to find variables that display
similar change or as we hypothesise are somewhat responsible for
the change observed in that universally constant feature.</p>
      <p>Our set of suitable candidate variables comprises the set of
bigram adjective-noun/noun-noun combinations, that in contrast to
our main variable need not be universally constant, but might only
turn out to be partially constant over the entire time span. Each
variable’s frequency in the group is relativized with respect to the
entire token count for the respective year.</p>
      <p>The very first step in this is to run a change point analysis for
‘last year’ over the entire 100-year span in order to ascertain the
exact point of change. As could be observed from figure 2, a change
happened a little after 1920 with the period afterwards giving rise
to a higher frequency pattern than the years leading up to it. One
then chooses an interval of a certain length after the change point
to limit the number of candidate features for examination. The
rationale in this case being that given a rise in frequency after the
change point one would expect correlated features to be partially
constant for at least a certain period of time afterwards, e.g 10 years.
We therefore extract the partially constant features for this time
period only. We then take the remaining features and calculate their
individual change points over the entire 100-year period. Given all
change points over all features, these are divided into three diferent
groups of features, those whose change points occur before the
main feature’s change point (in this case ‘last year’), those whose
occur exactly at the same time and those whose occur after. Only
those features that change significantly with respect to their mean
within 10 years before or after the main feature’ change are retained,
the reason being that we deem it unlikely that those changes more
remote in time would be related. Using the present method for
detection, features usually do not have more than one change point
and in the cases that they have two, these are separated by a time
span of at least 20 years. A change point indicates a change in
mean and what follows could either be an increase or a decrease in
frequency.9 As we only focus on features with similar trends, we
discard those with an opposing trend to our candidate feature, by
calculating the correlation between ‘last year’ and each feature over
the interval covering 15 years on either side of a feature’s change
point and only retaining those features for which this correlation is
positive.10 In the present case, the specifications were set as follows:
the change point for ‘last year’ was estimated at 1923, so we choose
the interval spanning the years 1924-1934 to look for features that
are constant over this period of time. We would not expect the
exact time frame to be of high importance, as one would expect
most features to level of more gradually over time. After discarding
features not constant over this interval, 103 features are left, where
at least 22 of these are also temporal expressions. In fact, when
we consider the universally constant adjective-noun combinations
that are constant over the entire 100-span, the majority of these
turn out to be temporal expressions (12/16). The fact that not more
features are constant over the entire span hints at the domain being
somewhat volatile with respect to content sequences.</p>
      <p>Table 1 shows the highest pairwise correlations (either
negative or positive) between ‘last year’ and each of the 104 features
over smaller intervals of 10 years from 1920 to 1970, where the
universally constant features are marked in italics. The first
interval covers a few years before the change point and a few years
9We focus on synonymous changes and causes here, i.e. the parallel increase of two
features together, rather than assuming that a decrease in one feature causes an increase
in the other feature, although this would also be a valid scenario.
10We used the Spearman rank coeficient for this, as available from the core R package.
after that, so somewhat of a transition period where diferent
concepts have similar trends to ‘last year’. There are a few temporal
expressions and politically/industry-related terms, such floor leader ,
executive session,vice president and automobile industry and a few
expressions (possibly temporal), that would probably be anchored
more strongly in the business context, such as first quarter and
second quarter. The second time window spanning 1930-1940,
features various concepts related to the stock exchange and business,
such common stock, stock market, business conditions and income
tax as well as a few temporal expressions possibly used in this
context, such as first quarter and second quarter. Interestingly, over
the next time span covering 1940-1950, for instance stock market
goes from being reasonably positively correlated (0.63) to being
negatively correlated (−0.42). and other concepts, such as european
countries and oil production take precedence instead. In the next
time window (1950-60), the highest rated concepts are negatively
correlated with ‘last year’, this efect becoming even stronger in the
very last time frame of 1960-70. Overall, we interpret this to mean
that very diferent concepts come to be used with ‘last year’ than
provided the basis for this set of correlated features. Certain events,
such the surprising wall street crash in 1929 could have caused
temporal expressions to gain more prominence and created an
atmosphere of immediacy that at least in the news world made the
use of temporal expressions more likely. With WWII and the cold
war shortly following, this might have kept the temporal dimension
palpable. When we examine the list of change points, including the
ones more than 10 years after ‘last year’, it is noticeable that a few
expressions’ points of change lie very close together, for instance
stock market, preferred stock, financial position , first quarter and third
quarter all change in either the year 1915 or 1916 and in 1945 or
1946. Figure 6 depicts this overlap in increase after the first change
point and return to initial mean frequency pattern after the second
change point.</p>
      <p>Another aspect that is noticeable in the results is that various
temporal expressions appear in the list of features highly correlated
with ‘last year’. This suggests that temporal expressions in general
increased in usage over time with respect to this genre. Figure 7
shows a few of the expressions from table 1. All seem to increase
in frequency over time. However, correlation analysis might be
a little volatile in that smaller spans of the entire period are not
representative of the overall correlation. For this reason, we seek to
validate our results further, which will be done in the next section.</p>
    </sec>
    <sec id="sec-11">
      <title>6.2 Validation of Results</title>
      <p>As the final part of this analysis, we seek to further validate our
results. One part of this is to see whether this efect also exists in less
changeable genre, such as fiction . Thus, we repeat the exact same
experiment, but using the fiction corpus as a basis rather than the
news corpus. We first estimate possible change points on the basis
of the new corpus. Interestingly, the change point for ‘last year’ in
the fiction genre happens earlier, around 1917. There seems to be
a lot less variety in adjective-noun combinations as the partially
constant features over 1918-1928 only add up to 27. Of these 27,
only 9 are positively correlated to ‘last year’ based on a span of ±
15 around their individual change point.</p>
      <p>However, only good evening, little girl, good night, good time and
very well are actually positively correlated with ‘last year’ over an
0.00015
y
c
n
e
u
req0.00010
F
e
v
ilt
a
eR0.00005
●
●
●
0.00000 ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ● ●
interval of ± 15 years around its own change point, with the highest
correlation being around 0.4. Figure 8 shows them side-by-side with
‘last year’. This suggests that the temporal aspect has not grown
as much in importance in this genre and is less closely linked to
adjective-noun types as it seems to be the case in the news domain.</p>
      <p>In order to validate this in the news data, we consider actual
language samples for co-occurrence of items highly correlated with
‘last year’. We randomly extract sentences containing ‘last year’
and observe what concepts co-occur in the same sentences. We
take ten samples each from before 1923 (1910-1920), immediately
after (1924-1934) and again at a later stage (1950-1960).</p>
      <p>Table 2 shows salient concepts occurring in the same sentence as
‘last year’ for all three time periods. The number in bracket indicates
in how many sentences of the ten selected ones the term occurred.
The terms occurring in the first time span are mostly related to
elections and governments with some more general political topics,
such as company and wages entering into it as well. The second
time span set around the change in ‘last year’ seems to contain
almost exclusively stock exchange related news items. The final
period, set after the end of WWII contains very mixed samples from
sports, to international politics, companies and space programs.
Although extracting a few random samples from a large set of
texts cannot provide very fixed conclusions, these results seem
to support our earlier findings of a strong correlation between
stock exchange related items and the temporal expression ‘last
year’ during a particular time period, where this seems to have
dominated the news. In order to see to what extent this efect
generalizes to other temporal expressions, we need to analyze these
separately.</p>
    </sec>
    <sec id="sec-12">
      <title>7 DISCUSSION</title>
      <p>
        We have reported an exploratory analysis to investigate the
relationship between temporal expressions, such as ‘last year’ and
temporally less stable word expressions that appear and
disappear over time. We hypothesized that these fluctuating words that
are more strongly connected to current events would somewhat
influence the rise in frequency of more stable concepts, such as
salient words occurring with ‘last year’
(primary) election(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), party(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), board of education(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), mexican bullets (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), company (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ),
director(s)(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), railroad(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), wages(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), shareholders(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), submarine(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), national committee(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
adjustment bond(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), common stock(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), stock (dividend)(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), (cash) investment (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), congress (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), sales(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), share(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ),
corporation(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), net profit(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), minor purchases(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), liquidation(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), dividend rate(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), president(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), preferred dividends(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
tournament(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), basketball coach(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), chicago medical society(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), tax bill(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), (space) administration(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), international agreement(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ),
wage(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), arbitration(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), net income(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), auto companies(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), production schedules(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), national aeronautics(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), russians pioneer(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
temporal expressions. Our results suggest that there might indeed
be a connection between ‘last year’ and clusters of words linked to
historical events, such as the stock marked crash. However, while
stock market related words are only constant and very frequent
for a limited time frame, ‘last year’ and other temporal expressions
remain frequent. We believe that this could be due to temporal
aspect in news language having become more important after 1923,
having gathered momentum through events, such as the stock
market crash and then remained to stay. Our parallel analysis of fiction
data at the same time seems to confirm this insofar as this efect
is not found with the same strength in fiction data. Based on our
language sample analysis that appears to support our change point
and correlation analysis, ‘last year’ is continued to be used in
various diferent concepts, possibly more varied than before 1923. Our
analysis using change points adds to a simpler relative frequency
detection approach by considering the uncertainties associated with
our predictions. Although without having conducted a semantic
change analysis, we cannot be entirely certain that this change is
not caused by a shift in semantics, however, the possible semantic
space of temporal expressions could be seen as more limited than
that for regular common nouns or adjectives. In fact, these temporal
expressions might semantically be closer to function types than to
content types, in spite of belonging to the latter word class. In a
sense, temporal adverbs are similar to prepositions, only anchored
in time rather than in space and consequently there might be less
room for reinterpretation of their meaning.
      </p>
      <p>Our results also need validation from historians, especially with
respect to events in 1923 that could have caused temporal
expressions to become more frequent. The type of analysis we have done
here shows changes in words’ relative frequency patterns that could
reflect political or cultural changes. In this, we are at the mercy
of the sampling of our newspaper corpus that although balanced
over diferent sources is not impervious to other external factors
that could influence the language samples. For instance, by the
mid-1920s, the businessman William Randolph Hearst had acquired
28 newspapers, that consequently have been subject to same
editorial decisions, distorting our perception of what language was
representative for that time.</p>
    </sec>
    <sec id="sec-13">
      <title>8 CONCLUSION AND FUTURE WORK</title>
      <p>In essence, this work has been exploratory trying to connect groups
of words that might not occur close to each other in space making
their relatedness less tangible. Although, additional work is needed
to further support our findings, our results tentatively suggest that
words or expressions that are stable in occurrence, might be rather
volatile with respect to their relative frequency distribution. As
temporal expressions have fewer semantic associations, they might
depend more strongly on features that do.</p>
      <p>The results we have obtained are tentative and in order to claim
an increase of temporal expressions possibly related to certain
historical events, one needs to show this efect to hold for other
temporal expressions as well as exclude any possible semantic shift.
We also need validation from historians to interpret and relate our
results to historical and cultural changes in or around 1923.
Particular language usage and change therein can reflect shifts in society
and general opinion, adding a more subtle basis for interpretation
of past events.</p>
      <p>Acknowledgement
We would like to thank our anonymous reviewers for their helpful
suggestions on how to improve the earlier version of this paper. This
research is supported by Science Foundation Ireland (SFI) through
the CNGL Programme (Grant 12/CE/I2267 and 13/RC/2106) in the
ADAPT Centre (www.adaptcentre.ie)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Joan</given-names>
            <surname>Bybee</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sandra</given-names>
            <surname>Thompson</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Three frequency efects in syntax</article-title>
          .
          <source>In Annual Meeting of the Berkeley Linguistics Society</source>
          , Vol.
          <volume>23</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Walter</given-names>
            <surname>Daelemans</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Explanation in computational stylometry</article-title>
          .
          <source>In Computational Linguistics and Intelligent Text Processing</source>
          . Springer,
          <fpage>451</fpage>
          -
          <lpage>462</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Davies</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>The Corpus of Historical American English: 400 million words</article-title>
          ,
          <year>1810</year>
          -
          <fpage>2009</fpage>
          . http://corpus. byu. edu/coha/.
          <volume>24</volume>
          (
          <year>2010</year>
          ),
          <year>2011</year>
          . (last verified:
          <volume>24</volume>
          .
          <fpage>08</fpage>
          .
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>William</surname>
            <given-names>L Hamilton</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jure Leskovec</surname>
            , and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Cultural Shift or Linguistic Drift? Comparing Two Computational Measures of Semantic Change</article-title>
          .
          <source>In Empirical Methods in Natural Language Processing (EMNLP).</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Nicholas</surname>
            <given-names>A James</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Arun</given-names>
            <surname>Kejariwal</surname>
          </string-name>
          , and David S Matteson.
          <year>2016</year>
          .
          <article-title>Leveraging cloud data to mitigate user experience from Breaking Bad</article-title>
          .
          <source>In Big Data (Big Data)</source>
          ,
          <source>2016 IEEE International Conference on. IEEE</source>
          ,
          <fpage>3499</fpage>
          -
          <lpage>3508</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nicholas</surname>
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>James and David S Matteson</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>ecp: An R package for nonparametric multiple change point analysis of multivariate data</article-title>
          .
          <source>arXiv preprint arXiv:1309.3295</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Vivek</given-names>
            <surname>Kulkarni</surname>
          </string-name>
          , Rami Al-Rfou,
          <string-name>
            <given-names>Bryan</given-names>
            <surname>Perozzi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Steven</given-names>
            <surname>Skiena</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Statistically Significant Detection of Linguistic Change</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web (WWW '15)</source>
          .
          <source>International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, Switzerland</source>
          ,
          <fpage>625</fpage>
          -
          <lpage>635</lpage>
          . https://doi.org/10.1145/2736277.2741627
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Meik</given-names>
            <surname>Michalke</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>koRpus: An R Package for Text Analysis</article-title>
          . http://reaktanz.de/ ?c=
          <source>hacking&amp;s=koRpus (Version</source>
          <volume>0</volume>
          .
          <fpage>05</fpage>
          -
          <lpage>4</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Riba</surname>
          </string-name>
          and
          <string-name>
            <given-names>Josep</given-names>
            <surname>Ginebra</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Diversity of vocabulary and homogeneity of literary style</article-title>
          .
          <source>Journal of Applied Statistics</source>
          <volume>33</volume>
          ,
          <issue>7</issue>
          (
          <year>2006</year>
          ),
          <fpage>729</fpage>
          -
          <lpage>741</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In Proceedings of international conference on new methods in language processing</source>
          , Vol.
          <volume>12</volume>
          .
          <string-name>
            <surname>Manchester</surname>
          </string-name>
          , UK,
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>