<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Behaviour of Russian Competing Verbs: a Computer-Assisted Approach</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Kazan Federal University</institution>
          ,
          <addr-line>Kremlyovskaya str. 18, 420008 Kazan</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1887</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The article studies morphological variability of Russian verbs. The distributional and quantitative analyses of these verbs were performed based on the extra-large diachronic corpus Google Books Ngram. The obtained frequency data were interpreted in terms of language norm and evolution. The accurate time of the norm change was identified for each pair of verbs. It was found that distribution of the competing verbs can both coincide and be markedly different. Four main trends in the frequency behaviour of the competing pairs of verbs were revealed. The analysis of the variability type and frequency of the variants showed that usage of a particular variant form is largely context dependent and is not determined only by a speaker`s individual preferences. Each of the variant forms has its niche in the language. It was also revealed that the observed active return of unproductive forms of verbs to the Russian language indicates the general stability of the verb system of the Russian language and tendency to unification by productive type.</p>
      </abstract>
      <kwd-group>
        <kwd>variability</kwd>
        <kwd>Russian verb</kwd>
        <kwd>linguistic norm</kwd>
        <kwd>language evolution</kwd>
        <kwd>Google Books</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The problem of linguistic variation has been extensively studied in the past
half-century, and it has now become a highly productive subfield of research in corpus and
computer linguistics. Variability is an inherent feature of any human language caused
by the asymmetric dualism of a linguistic sign when any content can be expressed by
different means [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Variations can occur at all levels within language and can be due to different factors.
At the beginning of the 20th century, the founder of the Kazan linguistic school
Baudoin de Courtenay [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] wrote that variability is the driving force of language evolution,
which involves constant oscillations and fluctuations in the structure of any language.
Variation analysis is one of the most fruitful areas in studies of language change and
evolution because change involves competition.
      </p>
      <p>
        Changes of language variants may not be obvious. However, thorough diachronic
analysis can trace their behavior within time and indicate how language evolves.
Several models of change have been proposed that describe the way linguistic changes start
and their stages [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ].
      </p>
      <p>Frequency of variant occurrence is a critical issue in such studies since analysis of
fluctuation between variants can show tendency in which variants have a greater or
lesser likelihood of occurring under certain conditions and help to observe language
change in progress.</p>
      <p>The beginning of the 21th century has witnessed remarkable growth in the
quantitative study of linguistic variation due to new research opportunities associated with
creation of extra-large text corpora. One of such text databases is the electronic library
Google Books. It contains a great number of texts that can be used to estimate frequency
of certain language phenomena and objectively deduce regularities of their use.</p>
      <p>
        Different types of language variation were studied using the Google Books corpus.
The dependence of rate of change of semantic meaning and connotative characteristics
of a word on its frequency usage was studied by Hamilton, Leskovec &amp; Jurafsky (2016).
Two quantitative patterns of semantic changes were identified for a period of 200 years
based on data from six historical text corpora written in four languages (including the
Google Books corpus): 1) the law of correspondence – the rate of se-mantic changes is
inversely proportional to the word frequency; 2) the law of innovation – regardless of
frequency, the more meanings the polysemantic word has, the higher the rate at which
it acquires new semantic meanings is [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Google Books Ngram data on morphological variability are increasingly interpreted
using native speakers` behavioral models when they choose a linguistic form [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and
text style (in which the variant is used) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Such data are also used to study how context
influence the use of one or another variant, as well as cognitive factors in general [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
      </p>
      <p>
        In our work, we study morphological variability of Russian verbs, which was first
studied by Smirnitsky [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and Vinogradov [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Variability of grammatical forms is
especially relevant for the Russian language due to its inflectional nature.
      </p>
      <p>As it is known, the Russian language has shown the following tendency for at least
several centuries. Some verbs forming a relatively small and (almost) unproductive
class, such as iskat` (‘look for’) transfer to the superproductive class of verbs, such as
igrat` (‘play’).</p>
      <p>The work objective was to study pairs of verbs and verbal forms that partially differ
in shape but have close meaning. The first task was to analyse the distribution of these
words and find some regularities of their use. The second task was, provided that the
distribution of some of the words are almost identical (muchit / muchaet – ‘torture’,
muchaetsia / muchitsia – ’anguish’), to analyze frequency of their use over time in terms
of language norm and evolution.</p>
    </sec>
    <sec id="sec-2">
      <title>Materials and Methods</title>
      <p>
        We studied pairs of verbs, which had the same semantics and similar form. They were
selected in the following way. Some of the verbs were taken from the book by
Graudina [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the rest were obtained using the Google Books Ngram corpus. More than
1000 verbs with alternating consonants at the end of the stem before the ending were
extracted automatically from the Google Books Ngram corpus (miau ...ch-et/...ka-et;
bryz ...zh-et/...ga-et (‘meows’, ‘splashes’). Then, pairs of verbs that met the
requirements were selected manually from the list. The list of the studied verb pairs included
122 verbs (each verb has two verb paradigms).
      </p>
      <p>Then, the Ngram Viewer service was used to conduct a distributive analysis and
study the contexts of use of the verb pairs. Ngram Viewer allows one to see the words
that are most often used with a given word. Word distribution makes it possible to draw
conclusions about differences in semantics of words, whether these words are
completely interchangeable or not, and what patterns of usage they have.</p>
      <p>Having performed the distributional analysis, we studied frequency of words used
in the most similar contexts in the course time to see if one verb is supplanted by another
and which of the forms dominates.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>The most common contexts of use of each pair of verbs and verbal forms were analyzed.
It was found that these contexts coincide in some cases and are markedly different in
other cases. For example, the distribution of words slomannyi/slomlennyi (‘broken’) is
significantly different.
The word slomannyi often collocates with concrete nouns in the meaning of ‘object’ or
‘body part’ in the form of the nominative case. The word slomlennyi collocates with
the word chelovek or abstract nouns dukh, gore, zhizn', pytki, ustalost', bolezn' (‘spirit’,
‘grief’, ‘life’, ‘torture’, ‘fatigue’, 'illness’) in the form of the ablative case (in the
meaning of cause, method, means or stimulus).</p>
      <p>Lexical and syntactic variability is also observed for many forms of verb pair
paradigms, for example:
krasotoi
‘by beauty’
noviznoi, umom, original'nost'iu
‘by novelty, by intelligence, by originality’
The participles blistaiushchii (‘shining’) in combination with the words mir, svet, mech,
zolotom, ogniami, chistotoi (‘world’, ‘light’, ‘sword’, ‘gold’, ‘lights’, ‘clean’) and
bleshchushchii (‘glittering, shining’) sneg, almaz, zdorov'em, umom, ostroumiem
(‘snow’, ‘diamond’, ‘health’, ‘intelligence’, ‘wit’) differ in the same way.
zdorov'em, sneg, ostroumiem, almaz
‘with health, snow, with wit, diamond’
The participle dvizhushchii (guiding) acts as an agreed definition in combination with
abstract nouns such as faktor, motiv, printsip, moment, impul's (‘factor', ‘motive’,
‘principle’, ‘moment’, ‘impulse’), as well as nerv (‘nerve’) and mekhanizm (‘mechanism’),
possible used in figurative meaning.
vpered, nogami, dushi
‘forward, on feet, souls’
faktor, motiv, mekhanizm, printsip, moment, impul's, nerv,
stimul, napor
‘factor, motive, mechanism, principle, moment, impulse,
neur, stimulus, pressure’
In this case, the difference in the contexts does not arise any questions as it is known
that the language tends to linguistic economy and two words which forms are different
but the meanings are the same can rarely be used in the texts of the same style for a
long time because one word often displaces the other. Though sometimes the process
of phraseologisation can take place and historical and archaic words and grammatical
phenomena can be saved and used in phraseological units due to linguistic memory
(zhit' pripevaiuchi – ‘live happily ever after’, sidet' slozha ruki – ‘sit back’, pritcha vo
iazytsekh – ‘a byword’). However, it should be noted that words with an excess
paradigm exist in the language due to various distribution. This reflects the cognitive
mechanisms of the language functioning.</p>
      <p>Besides, the study of distribution of some words (muchaet/muchit – ‘tortures /
tortures’, lazaet/lazit – ‘climb/climbs’) showed that the contexts of their use are identical.
Therefore, the semantics of these words has no differences. The use of words with the
same meanings but slightly different in form can be explained by difference in style,
the difference between bookish and colloquial speech.</p>
      <p>After the distributive analysis, the frequency analysis was performed to find how the
words under study behave and whether one word displaces another one with time and
becomes normative. The Google books data were used to build graphs of frequency
change of the finite and infinite forms of verbs with an excessive paradigm.</p>
      <p>We obtained 232 graphs described diachronic changes of verb pairs in 3Sg and 3Pl.
The verb pairs were classified according to the frequency dynamic of their use.</p>
      <p>The following results were obtained.
1. In 50% of cases, the unproductive form dominates the productive one, the norm
change is not expected. For example, kolyshet &gt; kolykhaet (‘to wave’), zhazhdu &gt;
zhazhdaiu (‘to yearn’), mechus', &gt; metaius' (‘flounce’), khnych' &gt; khnykai
(‘whimper’), pashut &gt; pakhaiut (‘plough’).
2. In 37% of cases, the productive form dominates the unproductive one. For example,
fyrkaesh' &gt; fyrchish' (‘to snort’), mykajutsja &gt; mychutsja (‘to torment’),
muchish'&gt;muchaesh' (‘to torture’), mykaet&gt; mychet (‘to wander’).
3. In 10% of cases, an unsuccessful attempt to change the norm was observed. For
example, the forms of the verbs dvigat' (‘to move’) and sekat'sia (‘to split’).
4. In 13% of cases, there was a long-term competition between the two forms. For
example, between the verbs zametat'sia (‘to sweep’), klikat'sia (‘to shriek’)
5. In almost 3% of cases, only a form with the stem of one type was found. For
example, the verb nianchit' (‘to nurse’).
6. Significantly lower number of graphs (8%) showed that productive forms are
displaced by unproductive forms (the form kaplet ‘to trickle’ is displaced by the form
kapaet). For example, the verb slomat'sia (‘to break’).</p>
      <p>The productive declination class is much more common in the other forms. In
approximately 64% of cases, the declination type of the 3Sg and 3Pl forms (kudakhchet,
kudakhchut – ‘to cackle’) does not coincide with the type of declination in other forms
(kudakhtaiu, kudakhtaia, kudakhtaiushchii, kudakhtai).</p>
      <p>Four tendencies were identified during the frequency analysis of competing forms
of words:
1. Absence of competition between the verb forms (only one form is found or dominate
another one throughout the target period).
2. Both forms have almost the same frequency.
3. Norm change (frequency of the less widespread form is increasing, and frequency
of the more widespread form is decreasing (the frequency curves tend to become
closer).
4. One form displaces another one (X-shaped chart).</p>
      <p>Accurate time of the norm change concerning each pair of verbs was also revealed. It
was found that the norm change most often occurs within two twenty-year periods:
1860-1880, 1910-1930 (see the Tables 5 and 6).</p>
      <p>Year
1830
1855
1870
1875
1910
1915
1915
1915
1920
1920
1940
1940
1950
1950</p>
      <p>Annual average relative frequency
1.600E-08
5.932E-09
1,754E-08
1.584E-08
3.202E-08
3.202E-08
5.045E-08
1.265E-07
4.592E-08
5.630E-08
4.903E-08
1.238E-07
9.401E-08
7.144E-08
It was also found that more frequent verbs require more time for changing by this or
that type. Verbal paradigms changed in the middle of the 20th century are 4-6 times
more frequent than those that changed in the middle of the 19th century (Fig. 1 and 2).
The large corpus data allowed us to determine which verb forms are often used in the
bookish texts and which forms are less frequent.</p>
      <p>As it was expected, the most frequent verbs were verbs in the forms 3Sg and 3Pl.
Present passive participle and present active participle were the least frequent.
It should be noted that the most "conservative" form is 2Pl (for example, mashete,
makhaete – ‘to wave’).
The “conservativeness” (tendency to save the original form) of the verbs turned out to
be directly proportional to their frequency: the more often a verb is used, the more often
it saves its original form.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>Differences in historically changing variant forms are often difficult to determine and
appropriately characterize. They relate to the sphere of native speakers` communicative
habits.</p>
      <p>
        However, modern corpus-based studies serve as a valuable tool for distributive and
statistical studies of word usage which can reveal regularities of their use [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. In
this work, the frequency characteristics of the studied excess verbs were first obtained.
      </p>
      <p>Besides, their usage was studied using the distributive semantics approach. These
results can be further used to deeply research cognitive processes relating to Russian
morphology and inflexion types. Study of such processes can allow one to solve some
problems of Russian grammar: mechanisms of development of grammatical semantics,
factors that influence appearance of irregular and non-standard inflection models of
words referring to different parts of speech in Russian.</p>
      <p>
        The obtained results are of great value for descriptive morphology of the modern
Russian language and can be useful for teaching practice of Russian as a foreign
language [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>The methods used to describe frequency dynamics of the variant verb forms can be
used in other cases. Moreover, these methods can allow one not only to describe and
explain linguistic phenomena, but also to make reasonable quantitative predictions of
the development of linguistic forms.</p>
      <p>The applied approach and obtained results can be used to compile dictionaries of
collocations and cognitive dictionaries.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The data on characteristic properties of use of the excessive verbs were first obtained
in this work.</p>
      <p>The distributive analysis of the studied pairs of words showed that the distribution
of the competing forms can be the same or have significant differences which are due
to different factors.</p>
      <p>The frequency analysis of the competing forms revealed 4 main trends in the
frequency behaviour of the competing pairs of verbs. The accurate time of the norm
change was identified for each pair of verbs.</p>
      <p>The active return of unproductive forms of verbs to the Russian language was
revealed which indicates the general stability of the verb system of the Russian language
and tendency to unification by productive type. 3Sg. and 3Pl. forms are at the top of the
frequency rating of the verb paradigm. They are more resistant to changes and
unification.</p>
      <p>Thus, the importance of considering the quantitative data of competing verb forms
in combination with the dynamics of their frequency is substantiated.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgment</title>
      <p>The work is performed at the expense of grant, diverted within the state support of the
Kazan (Volga Region) Federal University aimed at improving of its competitiveness
among the leading world scientific and educational centers, with funding from the
Russian Foundation for Basic Research (project No.17-29-09163 "Quantitative models of
diachronic changes and synchronic variability in the Russian language and linguistic
databases based on extra-large corpora").</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Elizarenkova</surname>
          </string-name>
          , T.:
          <article-title>O fakul'tativnosti i ee osobennostiakh v drevneindiiskom iazyke [About optionality and its features in the ancient Indian language]</article-title>
          .
          <source>In: Vostochnoe iazykoznanie: fakul'tativnost'</source>
          . Moscow, Russia (
          <year>1982</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Boduen de Kurtene, I.:
          <article-title>Izbrannye trudy po obshchemu iazykoznaniiu [Selected works on general linguistics]</article-title>
          .
          <source>Saint-Petersburg, Russia</source>
          (
          <year>1912</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bailey</surname>
          </string-name>
          ,
          <string-name>
            <surname>C-J.</surname>
          </string-name>
          :
          <article-title>Variation and linguistic theory</article-title>
          . Washington, DC, USA (
          <year>1973</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Pintzuk</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Variationist approaches to syntactic variation</article-title>
          . In Joseph,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Janda</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.)
          <article-title>The handbook of historical linguistics</article-title>
          ,
          <fpage>509</fpage>
          -
          <lpage>528</lpage>
          . Blackwell, Malden/Oxford (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leskovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jurafsky</surname>
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Diachronic Word Embeddings Reveal Statistical Laws of Semantic Change</article-title>
          .
          <source>In: Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <fpage>1489</fpage>
          -
          <lpage>1501</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Amato</surname>
          </string-name>
          , R.:
          <article-title>Human collective behavior: language, cooperation and social conventions</article-title>
          . Universitat de Barcelona,
          <string-name>
            <surname>Spain</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Klaussner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : Elements of Style Change. University of Dublin, Ireland (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Brand</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>The role of cognitive factors on the development of the vocabulary</article-title>
          . Lancaster University, United Kingdom (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Nesset</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuznetsova</surname>
          </string-name>
          , J.:
          <article-title>Stability and Complexity: Russian Suffix Shift over Time</article-title>
          .
          <source>Scando-Slavica</source>
          ,
          <volume>57</volume>
          (
          <issue>2</issue>
          ),
          <fpage>268</fpage>
          -
          <lpage>289</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Smirnitskii</surname>
          </string-name>
          , A.:
          <article-title>K voprosu o slove (problema tozhdestva slova) [To the question of the word (the problem of the word identity</article-title>
          ).
          <source>]. Trudy Instituta iazykoznaniia AN SSSR</source>
          ,
          <volume>4</volume>
          ,
          <fpage>3</fpage>
          -
          <lpage>9</lpage>
          (
          <year>1954</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Vinogradov</surname>
          </string-name>
          , V.:
          <article-title>Russkii iazyk. Grammaticheskoe uchenie o slove [The Russian Language. Grammatical studies of the word]</article-title>
          . Moscow, Russia, (
          <year>1972</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Graudina</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Grammaticheskaja pravil'nost' russkoj rechi. Stilisticheskij slovar' variantov [Grammatical correctness of Russian speech</article-title>
          .
          <source>Stylistic dictionary of variants]</source>
          .
          <source>Astrel'</source>
          , Moscow (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Zakharov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azarova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <source>Semantic Structure of Russian Prepositional Constructions. Lecture Notes in Computer Science</source>
          ,
          <volume>11697</volume>
          ,
          <fpage>224</fpage>
          -
          <lpage>235</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mukhamedshin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suleymanov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nevzorova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Choosing the Right Storage Solution for the Corpus Management System</article-title>
          .
          <source>Smart Innovation, Systems and Technologies</source>
          ,
          <volume>146</volume>
          ,
          <fpage>105</fpage>
          -
          <lpage>114</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Gracheva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kopylova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Konsonantnye cheredovaniia pri slovoizmenenii sovremennykh russkikh glagolov [Consonant alternations in the inflection of modern Russian verbs]</article-title>
          .
          <source>Vestnik Volgogradskogo gosudarstvennogo universiteta, Seriia</source>
          <volume>2</volume>
          : Iazykoznanie,
          <volume>5</volume>
          (
          <issue>29</issue>
          ),
          <fpage>153</fpage>
          -
          <lpage>158</lpage>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>