<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dynamics of core of language vocabulary</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Solovyev V. D.</string-name>
          <email>maki.solovyev@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bochkarev V.V.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shevlyakova A.V.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kazan Federal University</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <abstract>
        <p>Studies of the overall structure of vocabulary and its dynamics became possible due to creation of diachronic text corpora, especially Google Books Ngram. This article discusses the question of core change rate and the degree to which the core words cover the texts. Different periods of the last three centuries and six main European languages presented in Google Books Ngram are compared. The main result is high stability of core change rate, which is analogous to stability of the Swadesh list.</p>
      </abstract>
      <kwd-group>
        <kwd>core of vocabulary</kwd>
        <kwd>language dynamics</kwd>
        <kwd>Google Books Ngram</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In this paper, we investigate the dynamics of the overall structure of the language
vocabulary from a cognitive point of view. Traditionally, two components of the
language vocabulary are distinguished: the center and periphery. The former contains
highly stable words of maximum frequency (go, read, etc.) and provides stability to
the language; the periphery contains the words that have become outdated or, on the
contrary, have just appeared in the language, and thus, guarantees greater flexibility to
it. We will present some quantitative characteristics of the dynamics of the center.</p>
      <p>
        To do it, we should answer the following questions. How to determine the core?
What is the size of the core? What is the rate of change of the core? What is the
overall frequency of the core words? We will refer to Google Books Ngram corpus to
answer these questions (https://books.google.com/ngrams). Similar problems were
considered in [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. The frequency approach is a standard approach used to study
core formation. In this paper, we consider two kinds of frequency: the word
occurrence frequency in the corpus and the share of books in which the word occurs.
Though these approaches are rather close, yet there are some differ-ences.
      </p>
      <p>
        The first question to answer is how to determine the core. It’s impossible to define
a clear boundary of the core. For example, the known Swadesh wordlists contain 40,
100 or 200 items. In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] the core contains 100 words. It appears to be too limited. Let
us note that Basic English contains 850 words, and the basic set of root words of
Esperanto contains 900 items. The Voice of America’s Special English [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and
Wikipedia in Simple English use, corresponding-ly, about 1500 and 2000 words. The basic
vocabularies for foreigners [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], creole [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and pidgin languages [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] contain 1.5 to 3
thousand words.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] the core is composed of 1000 most frequent words (the first 100 words
constitute what is called the head, and words 101 to 1000 form the body), and the
periphery consists of the following in frequency 6000 words. In [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] the size of core
vocabulary that provides a speci-fied percentage of word usage based on the Google Books
Ngram data is calculated. Thus, 2300 most frequently used English words have the
total relative frequency of 75 %.
      </p>
      <p>We carry out calculations not only for one fixed core, but for consecutive variants:
for 1000, 2000, …, 8000 most frequent words, covering the whole range described
above.</p>
      <p>
        The following data preprocessing which allowed reducing the number of mistakes
in the used data base was performed in this work. Only lexical 1-grams were selected
which consisted only of the corresponding alphabet letters and one apostrophe in
some cases. To normalize and calculate the relative frequencies, the number of lexical
1-grams was calculated for each year (as distinct from the Google Books Ngram
Viewer where the normalization is made for the total number of all 1-grams). Parts of
speech are marked in the 2012 version of the corpus. But parts of speech are marked
wrongly in many cases which can result in incorrect conclusions based on these data.
We used the method explained in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], i.e. if the number of word forms correspond-ing
to some part of speech doesn`t exceed 1 % of total frequency of the given word form,
such word forms were marked and not used during further analysis.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Rate of change of the core.</title>
      <p>When considering the rate of change of the core, we calculate the share of words of
the core excluded from it during a given period. Figure 1 shows the relevant data for
an interval of 50 years in English language. Changes of word frequencies can be due
to both language evolution and random factors. To eliminate these factors,
frequencies of word usage were studied throughout rather long 50-year intervals: 1676-1725,
1726-1775, … 1976-2008. Then, the words were ranked in decreasing frequency
order and the percentage of words, which dropped out from the core of the successive
50-year interval were calculated. For example, the columns of the diagram marked
“1825” show the percentage of core words for the period 1776-1825 which dropped
out from the core in 1826-1875.</p>
      <p>We observe a rather steady rate of updating of the core in the last 300 years: an
aver-age of 13-15% of the words drop out of the core in 50 years. Of course, it does
not mean that these words disappear from language, only their frequency decreases,
and they are forced out from the core by other words. There is not enough data in
Google Books Ngram for the previ-ous period (1500–1700), and therefore they are
not provided here. Curiously, the updating rates of the core decrease during the
Victorian era and increase in the first half of the 20th century. Also, it should be noted that
the found mean value of 13-15% almost does not depend on the core size in the range
from 1 to 8 thousand.</p>
      <p>When the core is defined through the share of books, the following changes occur
in its content. If we select all English words that are found at least in one out of two
books, we obtain a wordlist of 2302 items. We can construct, for comparison, a list
with the same quantity of most frequent words. In spite of the fact that the share of
books in which a word is used corre-lates poorly with its frequency (the correlation
coefficient for all words of English language is just 0.15, for one thousand of the most
frequent words it is 0.25), both lists overlap by 79%. At the same time, the differences
between the lists are quite essential – there are 482 words that appear just in one list.</p>
      <p>Words included in list 1 seem to be, according to the intuitive perception of the
lan-guage, the most suitable for the core group of words. List 2 contains words that
can hardly be attributed with certainty to the core vocabulary. These words
correspond, first of all, to geo-graphical names and vocabulary with related meaning (for
example, Africa, African, Rome, Berlin, Japan, Japanese, Spain, Spanish, India,
Indians, Canada, California, Virginia, Asia), prop-er names/appellations (Wilson,
Richard, Louis, Oxford), parts of words/letters that entered the list accidentally (ff),
abbreviations (cf, vol., al, ibid.), articles and prefixes in loanwords and found foreign
vocabulary (der, des, du, le, les, un, el), words belonging more to professional
vocabulary than to common (carbon, oxygen, copper, equation, electron, protein), loanwords
(bureau), words connected chiefly with political actions (socialist, colonial, empire,
queen). However, according to the intuitive notion of language core, it is difficult to
ascribe the words from the specified groups to the core, but we should not deny their
importance for English-speaking society. In the culturological context, the words
Oxford and queen for British people are undoubtedly important, as well as the words
California, Africa and Virginia for Americans; additionally, professional words come
into broad use together with the growth of public aware-ness.</p>
      <p>As for the dynamics of the core (updating by 13-15% in 50 years), it practically
does not change, regardless of these two ways of determination.</p>
      <p>
        Let us consider the structure of the core from the perspective of the parts of speech.
In the latest version of Google Books Ngram, English words have been marked as
parts of speech with 95% accuracy [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In figure 2 we can see the share of each part
of speech around the year 1800 and today.
      </p>
      <p>X stands for abbreviations, foreign words or words whose membership to a part of
speech has not been determined. In 200 years the share of nouns and verbs has
diminished. Figure 3 shows the dynamics of the parts of speech. The algorithm for marking
the parts of speech works with higher accuracy in the case of modern words; this is
why the share of X is the one declining most rapidly.</p>
      <p>As one would expect, the parts of speech with the highest content, i.e. nouns and
verbs (about 45%), drop out at the highest rate, while auxiliary parts of speech,
articles, conjunctions, etc. (about 15 to 20%), do it at the lowest rate.</p>
      <p>Similar data are obtained for the main European languages (fig. 4) representing
three different branches of Indo-European languages: Slavic, Romance and German,
which separated just a few thousand years ago. This is somewhat similar to Swadesh
results. Russian rather stands out from the general picture. The social upheavals in the
beginning of the 20th century (the socialist revolution, which led to radical economic,
political, cultural changes) were re-flected in the vocabulary core.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Degree of covering of texts by the core</title>
      <p>The important characteristic of core words is to what extent they are efficient for
communica-tion. Formally, this can be presented by percentage of core words in the
texts, in other words by the degree to which the core words cover these texts. Let us
analyze now the change of the to-tal frequency of words of the core, that is the degree
of covering of texts by these words. If one considers the core for the language state in
1800 (for a higher stability in calculations one takes the interval 1795–1805 and
defines the core in the whole interval), it is evident that some words from the core will
become outdated, and the overall frequency will fall over time. The exact quantitative
characteristics of this process are given in figure 5 (the left window).</p>
      <p>
        For a 1000-word core the overall frequency falls in 200 years approximately from
0.7 to 0.6. Frequency curves for cores of bigger sizes look similarly. This effect may
be explained not only by the obsolescence of the words of the core (their removal
from the core), i.e. by the up-dating of the language, but also by the extension of the
vocabulary, which in general grants greater expressive opportunities to the language
and, naturally, leads to the reduction of the share of old words. According to data
provided in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the number of words in English language grew from 544,000 in 1900
up to 1,022,000 in 2000, i.e. almost twice.
      </p>
      <p>If one considers the modern core (years 2000–2008), the dynamics of its frequency
looks as follows (fig. 5, the right window). Here two tendencies confront. On the one
hand, it is evi-dent that two hundred years ago the frequency of modern words was
lower (up to 0), and it seems that one should expect a growth in the frequency of these
words. But, on the other hand, as we see in the previous diagram, the frequency of
words of the core in general falls. And these two tendencies approximately
counterbalance each other. The overall frequency of words for a core with 4 thousand words
remains at the level of approximately 0.8, for a core of 1000 words it slightly falls
from 0.67 to 0.65. The next graph (fig. 6) explains the essence of the pro-cesses
taking place. Here we can see separately the words that are present in the core both in
1800 and in 2000, and also the words present in one of them but not in the other.</p>
      <p>The overall frequency for the words remaining in the core during these two
centuries decreases from 0.7 to 0.6. The frequency of the words that drop out of the core
decreases, and that of the words entering the core increases, and this augmentation is
more intensive than the loss of frequency of the previous group.</p>
      <p>
        These data must be taken into account when analyzing the frequency dynamics of
dif-ferent groups of rather-high-frequency vocabulary. The frequency dynamics of
basic emotions are studied in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Data for English are presented in figure 7 (taken
from [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]). One can see that the overall frequency of emotive vocabulary considerably
decreases from 1800 to 2000. A pri-ori this can be explained either by a reduction of
emotionality of people (or at least that of texts) during this period, or by a general
reduction of the frequency of all the words of the core, which includes also the
considered emotive words. Comparison of the frequencies shows that the main acting
factor is the first one. The frequency of emotive vocabulary decreased approx-imately
by 50%, while the overall frequency of the words of the whole core decreased just by
15%. Thus, the reduction of the frequency of emotive vocabulary cannot be explained
only by the reduction of the frequency of the whole core.
      </p>
      <p>In the article, the lexicon structure is considered from cognitive point of view
distinguishing the centre (core – the most frequently used lexis) and periphery. The core
size is evaluated differently in different papers – from 1 to 8 thousand words. In our
paper, the calculations are performed for all core sizes in this range. The core change
data are presented for the first time. It turned out that the core has steadily changed
during the last 300 years – approximately 15% of words is substituted every 50 years.
The result is obtained for different languages (which are presented in Google Books
Ngram) and is, to some extent, analogous to the results obtained by Swodesh
concerning the stability of words from his list. The size of texts covered by the core words is
counted (or the total frequency of core words). It was found that the core (for the
contemporary language) consisting of 1 thousand words covers two thirds of texts. If we
regard the core words in 1800, the share of texts covered by them decreases from 0.7
to 0.6 for the last 200 years. This effect can be explained not only by core words
obsolescence (removing from the core), i.e. by language updating but also by lexicon
expansion which offers significant expression opportunities to a language and results
in decreasing of old words percentage.</p>
      <p>Acknowledgements. This research was supported by the Russian Foundation for
Basic Research (grant № 15-06-07402).
5</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Perc</surname>
          </string-name>
          , Matjaz:
          <article-title>Evolution of the most common English words and phrases over the centuries</article-title>
          .
          <source>J. R. Soc. Interface. 9</source>
          , pp.
          <fpage>3323</fpage>
          -
          <lpage>3328</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cocho</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flores</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gershenson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pineda</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sánchez</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Rank Diversity of Languages: Generic Behavior in Computational Linguistics</article-title>
          .
          <source>PLoS ONE</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ):
          <fpage>e0121898</fpage>
          . (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1371/journal. pone.
          <volume>0121898</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Beare</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Voice of America Special English Dictionary</article-title>
          .
          <article-title>English as 2nd Language</article-title>
          . http://esl.about.com/cs/reference/a/aavoa.htm.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Takala</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Estimating students' vocabulary sizes in foreign language teaching</article-title>
          .
          <source>In: Practice and Problems in Language Testing</source>
          , , vol.
          <volume>8</volume>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>165</lpage>
          . Afinla. https://www.jyu.fi/hum/laitokset/solki/afinla/julkaisut/arkisto/40/takala (
          <year>1985</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          : Haitian Creole: Grammar, Texts, Vocabulary. American Folklore Society, Philadelphia (
          <year>1953</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Romaine</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Pidgin and
          <string-name>
            <given-names>Creole</given-names>
            <surname>Languages</surname>
          </string-name>
          . Longman, London (
          <year>1988</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Michel</surname>
          </string-name>
          , J.-B.,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>Y.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aiden</surname>
            ,
            <given-names>A.P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veres</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          , et al.:
          <article-title>Quantitative analy-sis of culture using millions of digitized books</article-title>
          .
          <source>Science</source>
          <volume>331</volume>
          :
          <fpage>176</fpage>
          -
          <lpage>182</lpage>
          . (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Bochkarev</surname>
            ,
            <given-names>V. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solovyev</surname>
          </string-name>
          , V. D.:
          <article-title>Quantitative analysis of trends in the use of words with negative and positive connotations in Russian and English languages. (in Russian</article-title>
          ) In
          <source>: Proceedings of the VI International Conference on Cognitive Science</source>
          . Kaliningrad State University. (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Bochkarev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solovyev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wichmann</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Universals versus historical contingencies in lexical evolution</article-title>
          .
          <source>J. R. Soc. Interface</source>
          .
          <volume>11</volume>
          ,
          <issue>20140841</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Michel</surname>
          </string-name>
          , J.-B.,
          <string-name>
            <surname>Aiden</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orwant</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brockman</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Syntactic Annotations for the Google Books Ngram Corpus</article-title>
          .
          <source>In: Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics</source>
          Volume
          <volume>2</volume>
          :
          <string-name>
            <given-names>Demo</given-names>
            <surname>Papers</surname>
          </string-name>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>