<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Internal Dynamics of Text: Parts of Speech Distribution in Verse</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vadim Andreev</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>vadim.andreev@ymail.com</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Herzen State Pedagogical University of Russia</institution>
          ,
          <addr-line>Russian Federation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Larisa Beliaeva</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Smolensk State University</institution>
          ,
          <addr-line>Smolensk</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>The research is aimed at the study of the degree of regularity in the relationship between the frequencies of di erent parts of speech in a verse text, in particular between verbs and nouns. The data-base for the analysis includes 20 sonnets of famous Russian poets of the Silver Age of Russian poetry. The results demonstrate regularity in the distribution of parts of speech frequencies. The exponential function provides a good t.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Parts of speech (PoS) are often used in research in the sphere of quantitative linguistics,
stylometry and others [Best, 1994; Stamou, 2008]. Analysis of the frequencies of di erent PoS
in a text and proportions between them allow to solve important problems in \linguistics of
verse" which has been intensively developing in Russia in the 20th century [Gasparov, 2012].</p>
      <p>Depending on the peculiarities of individual styles the frequency of PoS vary to some
extent, and sometimes rather considerably [Cech, Altmann, 2013]. Nevertheless it is possible
to raise a question of the possible limits of variation and if there are any tendencies of keeping
certain proportions between PoS, if there is any general regularity in their frequencies common
for all the speakers of the same language. In some studies the results obtained demonstrated
the existence of certain order in PoS distribution in speech [Andreev, Popescu, Altmann,
2017]. The present research is aimed at exploring the possibility of such general tendencies
in the distribution of parts of speech in verse texts written by authors, di ering in style and
creativity manner.</p>
      <sec id="sec-1-1">
        <title>Teny proshlogo</title>
      </sec>
      <sec id="sec-1-2">
        <title>K portretu R. D. Balmonta</title>
      </sec>
      <sec id="sec-1-3">
        <title>Zhenshchine</title>
      </sec>
      <sec id="sec-1-4">
        <title>Kleopatra</title>
      </sec>
      <sec id="sec-1-5">
        <title>Sonet o poete</title>
        <p>Konstantin Balmont</p>
      </sec>
      <sec id="sec-1-6">
        <title>Bretan'</title>
      </sec>
      <sec id="sec-1-7">
        <title>Propovednikam</title>
      </sec>
      <sec id="sec-1-8">
        <title>Proklyatiye gluposty</title>
      </sec>
      <sec id="sec-1-9">
        <title>Razluka T10</title>
      </sec>
      <sec id="sec-1-10">
        <title>Put' pravdy</title>
        <p>Maksimilyan Voloshin
T1
T2
T3
T4
T5
T6
T7
T8
T9
T11
T12
T13
T14
T15
T16
T17
T18
T19
T20
Vyacheslav Ivanov</p>
      </sec>
      <sec id="sec-1-11">
        <title>Polyet</title>
      </sec>
      <sec id="sec-1-12">
        <title>La superba</title>
      </sec>
      <sec id="sec-1-13">
        <title>La pineta</title>
      </sec>
      <sec id="sec-1-14">
        <title>Nostal'giya</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Data-base</title>
      <p>Verse is characterized by a much bigger number of restrictions and rules in choosing words
than prose whereas sonnets is a poetic genre with one of the most formalized structure and a
big number of strict schemes.</p>
      <sec id="sec-2-1">
        <title>The data-base includes 20 sonnets by famous Russian authors V. Brusov, K. Balmont, V.</title>
        <p>Ivanov and M. Voloshin, written by these poets in the rst part of their creative activities.
All these poets belong to the period known as the Silver Age of Russian poetry (the beginning
of the 20th century) when much attention was paid exactly to this strictly structured genre
of sonnet. Below the names (titles) of these sonnets are given with their text numbers in the
data-base.</p>
        <p>Valery Brusov
\Starinnym zolotom i zhelchyu napital. . . "
\Zdes' byl svyaschenny les. Bozhestvenny gonets. . . "
\Ravnina vod kolishitsa shiroko. . . "
\Nad zibkoy ryab'u vod vsayet iz glubiny. . . "
\Mare internum"</p>
      </sec>
      <sec id="sec-2-2">
        <title>Na mig (\Den' purpur tsarstvenny dayet. . . ") 2</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methods and feature set</title>
      <p>The following parts of speech were counted in the sonnets: nouns ( N ), adjectives ( A ), verbs
( V ), adverbs (ADV), personal pronouns (PRNP), other types of pronouns which can be used
in attributive function (PRNA), participles (PTL), adjectivized participles (PTLA), category
of state words { adjectives, used as predicates in non personal sentences (STW).</p>
      <sec id="sec-3-1">
        <title>After counting the PoS in the sonnets quantitative data were obtained, specifying their</title>
        <p>frequencies. Thus in the sonnet Zhenshchine by Brusov (3) the following numbers of PoS
were obtained: 27 nouns, 15 personal pronouns, 9 verbs, 5 adjectives, 4 adjectivized
participles, 3 participles and 2 adverbs. (Category of state words and pronouns-adjectives were not
registered).</p>
        <p>The frequencies of all morphological classes in the samples were ranked in decreasing order
so that the most frequent PoS is ranked higher than all the others, the second frequency PoS
receives Rank 2, etc. To t the distribution of such ranked PoS frequencies the formula of the
exponential function which was suggested for such purposes in [Andreev, Mistecky, Altmann,
2018] was used:</p>
        <p>y = a exp( b x) ;
where a and b are parameters.</p>
      </sec>
      <sec id="sec-3-2">
        <title>If some PoS class was not found in the sample it was omitted (no zero classes were used).</title>
        <p>Consider, for example, the above-mentioned sonnet (3). PoS counting in it brought about
the following numbers, represented in Table 1. The rst column of this table shows rank
numbers, the second represents the PoS classes, the third is their observed frequencies in the
sample, the forth column shows theoretically expected frequencies which should be according
to the formula. Besides, at the bottom of the table the values of a and b parameters and of
the determination coe cient R2 are presented. The coe cient of determination is a measure
of goodness of t and provides information on whether a statistical model tted to empirical
data is successful. R2 ranges between 0 and 1. When R2 &gt; 0:8 the model ts well.</p>
        <sec id="sec-3-2-1">
          <title>In our case the value of R2 is over 0.99 which implies a very good t. In Figure 1 this is</title>
          <p>demonstrated graphically.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Along the x-axis we have ranks and along the y-axis|the values of the observed (empirical) and theoretically expected frequencies. As shown in Figure 1 the observed frequencies (OBSF) are very close to those expected on the curve (EXPF).</title>
      </sec>
      <sec id="sec-3-4">
        <title>For all the sonnets the results of tting were obtained, they are shown in Table 2.</title>
        <sec id="sec-3-4-1">
          <title>All the values of the determination coe cient R2 are very high. Thus even the lowest R2</title>
          <p>value for tting the distribu-tion of PoS in 4 (R2 = 0:8897) and in 7 (R2 = 0:8924) should
be considered as a proof of a good t. Since the sonnets were written by di erent poets whose
style and manner of writing as well as the topics were di erent, it should be recognized that
the distribution of PoS does not depend on such individual matters, but displays some kind
of regularity.</p>
        </sec>
      </sec>
      <sec id="sec-3-5">
        <title>From the point of view of how description takes place in sonnets one can group di erent PoS</title>
        <p>into two classes. One of these classes actualizes a static vision of the poetic world [Naumann,
Popescu, Altmann, 2012]. In this case the author depicts the world attributing to the themes,
expressed by nouns, some features which are viewed as more or less permanent qualities. This
is achieved by using such PoS as A , PRNA, PTLA, STW. The other class, on the contrary,
gives a description which can be called dynamic, because the features ascribed to the themes
in the sonnet are represented as a process or action. This class includes V and PTL. Further
on we shall analyze how these two classes of PoS interrelate with one another [Martynenko,
2004].</p>
        <p>Replacing the PoS of two classes by the name of the class to which they belong: S for the static
description and D for the dynamic one, and omitting all other PoS, one obtains sequences which
characterize the level of homogeneity of description.</p>
        <p>Let us consider, as example, T3 again. After marking up the two above-mentioned classes we get
the following sequence:</p>
        <p>D</p>
        <p>D</p>
        <p>S</p>
        <p>S</p>
        <p>S</p>
        <p>S</p>
        <p>S</p>
        <p>D</p>
        <p>D</p>
        <p>D</p>
        <p>D</p>
        <p>D</p>
        <p>S</p>
        <p>D</p>
        <p>S</p>
        <p>S</p>
        <p>D</p>
        <p>S</p>
        <p>D</p>
        <p>D</p>
        <p>D;
This sequence consists of a number of strings formed by repeated elements which further on will be
called \runs" [Andreev, Mistecky, Altmann, 2018, p. 50{52]. Here the following runs can be singled
out:
Little number of runs, including big chains of similar members, indicates intensi ed monotony of
description, big number of runs with few elements in them, on the contrary, suggests something
like variability in depicting poetic world. In this example the total number of elements in all the
runs equals 21, the number of runs equels 9. Thus it follows that the index of the homogeneity
of description is Itotal = 21=9 = 2:33. Measuring homogeneity of static and dynamic descriptions
separately we get the following:
(1) static Ic = 2; 25(9=4);
(2) dynamic Iv = 2; 4(12=5).
of variability. It should be noted that sonnets are rather brief limited to 14 lines, and nevertheless
in a number of cases the di erences are apparent. This can be shown graphically. In Figure 2 the
scatterplot demonstrates the relations of two indices| Ic and Iv in 20 sonnets. On the horizontal
axis the values of Ic are set, the y-axis sets the values of Iv in the sonnets.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion &amp; Discussion</title>
      <p>The scatterplot shows that priority should be given to y-axis coordinate. X-axis does not provide
a basis for classi cation, except for 1 outlier (T20) all the other texts form a rather dense group.
On the other hand, index Iv , showing the dynamic homogeneity, divides the sonnets into 2 groups.
The rst one consists of 2, 3, 5, 9 17. It should be noted that 2 and 9 completely overlap and are
marked by a common dot. All the other sonnets form another group. At the level of Iv = 1:5 it is
split into two subgroups of equal number of texts. The homogeneity index of dynamics in description
Iv = 1:5 is observed in three texts (15, 18, 20) thus forming a basis for the splitting of the whole
group. Above and below this borderline there are 6 texts in each subgroup.</p>
      <p>On the whole it is possible to conclude that this research demonstrated some order in all PoS
distribution in sonnets and the relations between words which depict static and dynamic description
of the poetic world.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Best</source>
          , 1994]
          <string-name>
            <given-names>Best</given-names>
            <surname>K.-H.</surname>
          </string-name>
          (
          <year>1994</year>
          ).
          <article-title>Word class frequencies in contemporary German short</article-title>
          prose texts // Journal of Quantitative Linguistics.
          <year>1994</year>
          . Vol.
          <volume>1</volume>
          . P.
          <volume>144</volume>
          {
          <fpage>147</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Stamou</source>
          , 2008]
          <string-name>
            <surname>Stamou</surname>
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2008</year>
          )
          <article-title>Stylochronometry: Stylistic Development, Sequence of Composition,</article-title>
          and Relative Dating // Literary Linguistic Computing,
          <volume>23</volume>
          (
          <issue>2</issue>
          ).
          <source>P</source>
          <volume>181</volume>
          {
          <fpage>199</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Gasparov</source>
          , 2012]
          <string-name>
            <surname>Gasparov</surname>
            <given-names>M. L.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Exact methods of grammar analysis in verse</article-title>
          . In.:
          <string-name>
            <surname>M.L Gasparov</surname>
          </string-name>
          .
          <article-title>Selected works. V. 4</article-title>
          . Moscow: Languages of Slave culture,
          <year>2012</year>
          . P.
          <volume>23</volume>
          {
          <fpage>35</fpage>
          . (In Russian) =
          <article-title>Tochniye metody analiza grammatiki</article-title>
          v styhe // M.L. Gasparov.
          <source>Izbranniye trudy. .4</source>
          .
          <year>2012</year>
          . S.
          <volume>23</volume>
          {
          <fpage>35</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Cech, Altmann, 2013]
          <string-name>
            <surname>Cech</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altmann</surname>
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>Descriptivity in Slovak lyrics</article-title>
          .
          <source>Glottotheory</source>
          .
          <year>2013</year>
          . Vol.
          <volume>4</volume>
          (
          <issue>1</issue>
          ). P.
          <volume>92</volume>
          {
          <fpage>104</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Andreev, Popescu, Altmann, 2017] Andreev,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Popescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.-I.</given-names>
            ,
            <surname>Altmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Skinner's hypothesis applied to Russian adnominals</article-title>
          .
          <source>In: Glottometrics 36</source>
          . RAM-Verlag. P.
          <volume>32</volume>
          {
          <fpage>69</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Andreev, Mistecky, Altmann, 2018]
          <string-name>
            <surname>Andreev</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mistecky</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altmann</surname>
            <given-names>G</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Studies in quantitative linguistics - 29</article-title>
          . Lu^denscheid: RAM-Verlag,
          <year>2018</year>
          . { 130 p.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Naumann, Popescu, Altmann, 2012]
          <string-name>
            <surname>Naumann</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popescu</surname>
            <given-names>I.-I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Altmann</surname>
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2012</year>
          ). Aspects of nominal style // Glottometrics.
          <year>2012</year>
          . V. 23. P.
          <volume>23</volume>
          {
          <fpage>55</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Martynenko</source>
          , 2004]
          <string-name>
            <surname>Martynenko G. Ya.</surname>
          </string-name>
          (
          <year>2004</year>
          )
          <article-title>Rhythmic-semantic dynamics of the Russian classical sonnet</article-title>
          . Saint-Petersburg: SPb University,
          <year>2004</year>
          . { 30 p.
          <article-title>(In Russian) = Ritmiko-smyslovaya dynamika russkogo klassicheskogo soneta</article-title>
          .
          <source>SPb: SPbGU</source>
          ,
          <year>2004</year>
          . { 30 s.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>