<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Dynamics of Extensive Text Variables in Russian Short Stories</article-title>
      </title-group>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The research presented in this paper is aimed at the analysis of dynamic organization of a literary text. Using the statistical time series method, the dynamics of the main extensive text variables - the mean paragraph length and the mean sentence length - is considered. The material for this study was the annotated subcorpus from the Corpus of the Russian Short Stories of 19001930, which consists of 310 stories written by 300 Russian writers. It was narrative fragments of texts (the narrator's speech) that were subjected to analysis, dialogical fragments were not taken into consideration. As a result, the most frequent dynamic profiles of paragraph length and sentence length were obtained, which reflect the most typical structures of the dynamic organization of short literary texts.</p>
      </abstract>
      <kwd-group>
        <kwd>dynamic text structure</kwd>
        <kwd>quantitative literary studies</kwd>
        <kwd>paragraph length</kwd>
        <kwd>sentence length</kwd>
        <kwd>time series</kwd>
        <kwd>stylometrics</kwd>
        <kwd>corpus linguistics</kwd>
        <kwd>text composition</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The research described in this paper is aimed at studying dynamic organization of
literary texts expressed in the categories of paragraph length and sentence length.
These quantitative measures are traditionally the focus of style and language studies
[
        <xref ref-type="bibr" rid="ref1 ref14 ref18 ref2 ref28 ref5 ref6 ref8">1, 2, 5, 6, 8, 14, 18, 28</xref>
        ]. Scientific interest in investigating these extensive text
variables has been intensified in recent years and may be explained by their importance for
solving many urgent applied tasks related with texts classification and
attribution [4, 7; 21, 25, 26]. As a rule, these variables are calculated on average over the
text, their measures of variation being not taken into account. However, the question
– How stable these variables are on the time scale of text composition? – remains
actual [11, p. 219].
      </p>
      <p>
        On the other hand, no less urgent is the task of comparing the dynamics of these
variables with the study of text composition or its plot structure. In recent years, the
interest in computer analysis of fiction texts has sharpened significantly, that may be
explained by the modern technological development of society, technologies of
computational linguistics and digital humanities, as well as the needs for the development
of artificial intelligence systems [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref16 ref9">9, 11–13, 16</xref>
        ]. A vivid example of a dynamic
apCopyright © 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
proach to the analysis of fictional texts from the point of view of the development of
text tonality (emotional trajectories) is given in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        Developing the methodology for this study, the authors relied on the method
proposed by Gregory Martynenko, who is the founder of the St. Petersburg stylometric
school [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data and Method</title>
      <sec id="sec-2-1">
        <title>Material</title>
        <p>
          The material for this study is the annotated subcorpus of the Corpus of Russian Short
Stories of the First Third of the 20th Century [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. This corpus is designed as a
literary resource which should become the research site for various linguistic and stylistic
studies, which implies the necessity to include literary texts of the maximum number
of writers who wrote in a given historical period [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Apart from texts written by
well-known and outstanding writers, the corpus contains short stories of the large
number of ʻsecond-rateʼ authors, which are also involved into consideration. Thereby
better literary representation of different aspects of social and cultural life, as well as
of language and stylistics diversity is achieved [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>
          From the corpus, a subset of 300 randomly chosen short stories was created, with
100 stories for each subperiod (1900–1913, 1914–1922, 1923–1930), one per the
author. Finally, 10 texts of the authors that wrote through all three subperiods were
added [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Method</title>
        <p>The texts were cleaned from dialogues and checked for the presence of at least
10 paragraphs or sentences. Thus, 305 and 308 texts were included in the final subset
for paragraph analysis and for sentence analysis respectively.</p>
        <p>
          In order to create dynamics contours, the general approach proposed in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] was
used. It was implemented in the following way:
1. It was hypothesized that there is a correlation between two extensive variables –
mean sentence length and mean paragraph length – and a plot of a text. Namely,
action text fragments were expected to have shorter paragraph and sentence length
than descriptive or reflective episodes.
2. Each text was tokenized with R software [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. Thereafter, two variables were
measured: paragraph length (in sentences) and sentence length (in words).
3. Texts were tested for whether the changes in paragraph and sentence length are
significant.
4. The sequence of paragraphs was divided into 10 groups, regardless of their size.
        </p>
        <p>
          Then, for each group mean length was measured. The same was done for
sentences.
5. The results were presented as time series, with the group as the independent
variable and mean paragraph or sentence length – as an independent one [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
6. Time series were smoothed with the moving average method.
7. The final result was visualized as a line graph.
        </p>
        <p>
          For example, the application of this method to the short story “Resort husband”
(“Kurortny muzh”) written by Alexander Amfiteatrov in 1911 leads to the following
results (see Table 1).
Then, these numbers are smoothed with the moving average method [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], where the
absolute numbers are replaced with the mean for a certain interval (three groups in
this experiment). In this way, the means for each group are replaced with their
smoothed values – for example, second and third ones are calculated with the
following formulas:
The result for the story “Resort husband” is presented in Table 2 and Figure 1.
Means for the first and the final groups are smoothed by separate formulas:
 2 =  1+  32+  3,
 3 =  2+  33+  4.
 1 = 2 1+  2−  4,
        </p>
        <p>2
 10 = 2 10+  9−  7 .</p>
        <p>3
4,5
3,5
4
3
2
1
2,5
1,5
0,5
0</p>
        <p>a)</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The Results</title>
      <sec id="sec-3-1">
        <title>Mean paragraph length</title>
        <p>Table 3 contains some of the most frequent figures for paragraph length dynamic,
which cover about 50% of the text sample.</p>
        <p>As this table shows, the most frequent patterns are those, in which a text begins
with “heavy” paragraphs that gradually shorten towards the rising action. The
following changes depend on the type of a figure. For instance, types 1 and 2 have a similar
dynamic – the rise of paragraph length and its subsequent fall; the key difference here
lies in the nature of the rise. In type 1, a growth of paragraphs for the most part
happens at the climax or near the falling action – these elements tend to be the heaviest
ones. In type 2, on the other hand, the rise leans closer to the rising action and climax.
Moreover, type 2 is distinguished by the lower extremity of the initial fall – mean
paragraph length in the rising action rarely becomes less than the mean value across
all intervals.</p>
        <p>Another variation of the change is represented by groups 3 and 4, in which the
paragraph size in the later text parts remains relatively small. Texts of both types still
retain the increase of mean paragraph length, but in the first case it is closer to the
falling action and is usually insignificant, while in the second case the rise is followed
by the fall typical for types 1 and 2.
“A complicated case” (“Zaputanny sluchay”, 1927)
by Arkady Bukhov, “An accident”
(“Proisshestviye”, 1924) by Boris Lavrenyov, “Six days” (“Shest
dney”, 1925) by Nikolay Nikitin
To conclude, in the standard pattern the exposition is always composed of large
paragraphs – it is probably related to the fact that this part of the text gives the reader a
background of the main conflict and requires a detailed explanation. The rising action,
on the other hand, should be more dynamic – short, abrupt paragraphs are more
fitting. The subsequent climax has different patterns – this can happen due to two
possible factors: the difference in climax nature (more “thoughtful” episodes describing the
characters’ feeling at the higher point of conflict (types 1 and 2) vs more dynamic,
emotional ones (types 3 and 4)) or the presence of dialogues (for the cases where
narration followed characters’ speech and served as short remarks for it; the deletion
of dialogue could lead to the drop in the mean paragraph length in such episodes).
Finally, the falling action and resolution require neither a big amount of information
nor a detailed description – here, smaller paragraphs can be rather used.</p>
        <p>There are, however, several figures that differ from the “standard” pattern. For
example, types 5 and 7 follow the initial “large exposition – small rising action – large
climax” pattern but have a rise in paragraph length in the falling action and resolution.
Texts of these groups probably required more explanations in the end, having either
characters’ reflection on the event of the ending or a more detailed description of the
events following the climax.</p>
        <p>Another non-standard variation is the type where only the “long climax – short
resolution” pattern is retained, while the exposition is composed of shorter
paragraphs. The stories of this type usually have a more dynamic beginning – perhaps,
descriptions in them are replaced with the action or dialogues.</p>
        <p>Finally, among frequent figures, there is a type opposite to types 1 and 2: a
dynamic exposition, a “heavy” rising action, a subsequent dynamic climax, and a large
resolution. The explanatory parts of these texts are evidently moved towards the rising
action and the resolution, while the rising action and the climax are richer in action or
dialogue.</p>
        <p>This pattern is, however, the only one completely different from the standard: all
the other types keep the features of the common type in one or another way. Based on
this, it can be concluded that literary texts mostly lean towards the “heavy” exposition
and climax and the more dynamic rising and falling action.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Mean sentence length</title>
        <p>Table 4 contains some of the most frequent figures of the dynamic of mean sentence
lengths, which cover about 44% of the sample.</p>
        <p>Sentences generally tend to behave similarly to paragraphs – in most cases their
length diminishes along the text. The possible fluctuations are also close to the ones
found in paragraph lengths, with the pattern of the rise of volume in the climax being
the most frequent. The intensity of these fluctuations also has some differences: for
instance, group 1 is characterized by the milder change in mean sentence length and
its earlier rise, while group 4 has a more extreme initial drop and a rise closer to the
resolution. Moreover, there remains a type with a sharp drop in the rising action and
only a slight rise in the climax.</p>
        <p>At the same time, the type with longer rising action and resolution becomes more
frequent. Coupled with the opposite tendency in paragraph length, it can be explained
as the preference for big paragraphs filled with short sentences in the exposition and
climax – and for paragraphs with few long sentences in the ending.</p>
        <p>There is also an increase in frequency for some figures that are rare for paragraphs,
such as type 6, a more radical case of “long beginning – short ending” pattern, and
type 5, an U-shaped figure that might emerge either due to the need for more dynamic
action episodes and more explanations in the ending or due to the presence of large
dialogues, similarly to the case of paragraphs.
“Petushkov Rocket” (“Raketa Petushkova”, 1924)
by Gleb Alekseyev, “A glass of champagne” (“Bokal
shampanskogo”, 1911) by Kazimir Barancevich, “It
blowed gently” (“Potyanulo”, 1910) by Vasily</p>
        <p>Bashkin
To conclude, sentences by most part have little difference from the paragraphs.
However, there is an increase in the frequency of the type opposite to the “standard” one.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Comparing the dynamics of plot and extensive variables</title>
      <p>
        The study of plot dynamics on the material of the corpus of the Russian story was
previously described in [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. In this study, we will check how promising is the
comparison of compositional elements and the dynamics of extensive variables.
      </p>
      <p>Let us consider in what way the discovered dynamic is tied to the plot of short
stories, on the example of two texts: “The mistery” (“Tayna”) by Aleksey P. Dementyev
written in 1915 and “Strained relationships” (“Obostrenie otnosheniya”) by Euhene
V. Chirikov written in 1903.</p>
      <p>Table 5 and Figure 2 show the means and dynamic contours for “The mystery” by
Aleksey P. Dementyev.
This short story can be considered a “standard” one: both paragraph and sentence
length change according to the most frequent patterns.</p>
      <p>The text opens with long exposition and rising action (groups 1–3) that are rich
with descriptions of nature and the deacon’s thoughts. Closer to the episode of the
deacon coming to pope Mikhail, paragraph and sentence lengths decrease – this part
of the text is comprised mostly of dialogues and does not require long explanations.</p>
      <p>As the plot progresses and the deacon’s “mystery” is revealed in the climax
(groups 4–6), both lengths grow again. This happens, on one hand, due to the
presence of long reflexive paragraphs that build the anticipation of the revelation of the
“mys- tery” and, on the other hand, due to the detailed description of the deacon’s
hunt and his feelings related to this experience. Here, the pacing of the narrative slows
down: more importance gains not the action itself, but the sense of excitement that the
main character feels from the hunt. Because of that, the climax of the story is
somewhat similar to the exposition, both in its content (mostly descriptive and reflective
para- graphs) and size.</p>
      <p>In the falling action and resolution (groups 7–10), a drop of both quantitative
variables takes place – these parts of the text contain a relatively short dialogue between
the deacon, his wife, and pope Mikhail. The ending is given mostly through a
sequence of paragraphs with lots of short sentences that describe how other characters
and the deacon himself view his hunts.</p>
      <p>A different structure can be seen in the story “Strained relationships” by
Euhene V. Chirikov. Table 6 and Figure 3 show the means and dynamic contours for this
text.</p>
      <p>Like Dementyev’s story, the story by Chirikov begins with relatively long
exposition and rising action (groups 1–4). In the exposition the reader is given the reason for
Misha’s quarrel with parents and his refusal to eat with them – in other words, the
basis of the main conflict. The rising action is built in the same way: it contains
mostly reflective paragraphs flowing into the action-oriented ones.</p>
      <p>However, in the rising action and the climax (groups 5-8), the dynamic begins to
deviate from the “standard” variation: the author still uses large paragraphs but fills
them with short sentences. This difference reflects the distinction in the mood of both
stories – while “The mystery” had a more melancholic, slow-paced feel to it,
“Strained relationships” is more energetic. The plot in it is developed first through
Misha’s chaotic planning depicted by short, abrupt sentences and then through the
market episode comprised mostly of action and dialogues.</p>
      <p>The opposite occurrence can be seen in the falling action and the resolution
(groups 9-10) where paragraphs become shorter and sentences – longer. This, again,
is related to the fact that the consequences of Misha’s lie can be described shortly: the
family got concerned about his health, and the “strained relationships” thus were
resolved. Moreover, the length of the sentences is accomplished through the listing of
multiple actions – this conveys the fuss and concern that overcame the mother and
sister of the main character. The ending, on its end, concentrates on Misha’s feelings
after the resolution of the conflict and describes them in a few large sentences.
Mean
sentence length,
smoothed
value
14.43
10.5
8.85
8.05
8.13
8.05
8.31
9.05
10.5
12.99
4,5
3,5
4
3
2,5
2
As it was mentioned before, the dynamic of extensive variables in question has
several similarities: both paragraph and sentence length lean towards the downward trend
with the large exposition and falling action; in addition, they tend to increase in
volume in the climax and resolution. There are, however, other variations of dynamic
contours that appear due to the nature of the plot and the structure of a specific text.</p>
      <p>
        Resulting figures coincide with those obtained in [
        <xref ref-type="bibr" rid="ref10 ref11">10–11</xref>
        ] – most notably, among
the most frequent were the ones similar to types 1 and 6 for paragraphs and type 3 for
sentences. Moreover, the results of the experiment prove that the pattern of large
sentences in the beginning, their subsequent drop in the rising action, and the growth in
climax with the final drop in the ending can be considered standard for the description
of the narrative.
      </p>
      <p>At the same time, in the research by Gregory Martynenko, some figures were
found to be infrequent by the results of the current experiment. This difference can be
explained by the characteristics of Martynenko’s subset: it included only 20 texts
from 4 writers. The expansion of the subset and the increase in its variety thus might
specify the exact frequencies of the plot development figures.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>The presented study confirmed the results obtained earlier in the analysis of small
Russian prose that: 1) extensive text variables, such as the average sentence length
and the average paragraph length are not constant values, but change with the
development of plot narration, and 2) the change of the extensive text characteristics is not
a chaotic process; general dynamic patterns (or trends) of their changes in time can be
traced. Moreover, it can be argued that there are some invariant typical structures that
occurs more often than others in literary texts of short prose.</p>
      <p>
        It seems appropriate to continue the research by involving data on another
important extensive text characteristic – the total size of literary text measured in words,
as well as additional information about text internal structure (its division into
chapters, sections, etc.), the topic(s) of the text [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] and other content-based characteristics
related to literary annotation of data. Thus, more accurate information about the
frequency and implementation features of certain typical dynamic profiles will be
obtained.
7
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The research is supported by the Russian Foundation for Basic Research, project
# 17-29-09173 “The Russian language on the edge of radical historical changes: the
study of language and style in prerevolutionary, revolutionary and post-revolutionary
artistic prose by the methods of mathematical and computer linguistics (a
corpusbased research on Russian short stories)”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Admoni</surname>
          </string-name>
          , V.G.:
          <article-title>Razmer predlozheniya i slovosochetaniya kak yavlenie sintaksicheskogo stroya [The Length of Sentences and Phrases as a Phenomenon of Syntactic Structure]</article-title>
          .
          <source>Voprosy yazykoznaniya [Topics in the Study of Language]</source>
          ,
          <source>1966(4)</source>
          ,
          <fpage>111</fpage>
          -
          <lpage>118</lpage>
          (
          <year>1966</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Akimova</surname>
            ,
            <given-names>G.N.</given-names>
          </string-name>
          :
          <article-title>Razmer predlozheniya kak faktor stilistiki i grammatiki [Sentence Length as a Factor of Stylistics and Grammar]</article-title>
          .
          <source>Voprosy yazykoznaniya [Topics in the Study of Language]</source>
          ,
          <source>1973(2)</source>
          ,
          <fpage>67</fpage>
          -
          <lpage>79</lpage>
          (
          <year>1973</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Coghlan</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>A Little Book of R for Time Series, Release 0.2</article-title>
          . Wellcome Trust Sanger Institute, Cambridge (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Grieve</surname>
          </string-name>
          , J.:
          <source>Quantitative Authorship Attribution: An Evaluation of Techniques. Literary and Linguistic Computing</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ),
          <fpage>251</fpage>
          -
          <lpage>270</lpage>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Huxtable</surname>
          </string-name>
          , R.:
          <article-title>Sentence length</article-title>
          .
          <source>Science</source>
          ,
          <volume>197</volume>
          (
          <issue>4300</issue>
          ),
          <volume>208</volume>
          (
          <year>1977</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kelih</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grzybek</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antić</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stadlober</surname>
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Quantitative Text Typology: The Impact of Sentence Length</article-title>
          . In: Spiliopoulou,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Kruse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Borgelt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Nürnberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Gaul</surname>
          </string-name>
          , W. (eds.)
          <article-title>From Data and Information Analysis to Knowledge Engineering, Proceedings of the 29th Annual Conference of the Gesellschaft für Klassifikation e</article-title>
          .V., University of Magdeburg, March 9-
          <issue>11</issue>
          ,
          <year>2005</year>
          ,
          <fpage>382</fpage>
          -
          <lpage>389</lpage>
          . Springer, Berlin (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Lagutina</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagutina</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boychuk</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vorontsova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shilakhtina</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belyaeva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paramonov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demidov</surname>
            ,
            <given-names>P.G.</given-names>
          </string-name>
          :
          <article-title>A Survey on Stylometric Text Features</article-title>
          . In: Balandin,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Niemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ,
            <surname>Tuytina</surname>
          </string-name>
          , T. (eds.).
          <source>Proceedings of the 25th Conference of Open Innovations Association FRUCT</source>
          , Helsinki, Finland,
          <fpage>184</fpage>
          -
          <lpage>195</lpage>
          . Institute of Electrical and Electronic Engineers, New York (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lesskis</surname>
            ,
            <given-names>G.A.</given-names>
          </string-name>
          :
          <article-title>Nekotorye statisticheskie zakonomernosti kharakteristiki prostogo i slozhnogo predlozheniya v russkoj nauchnoj i khudozhestvennoj proze XVIII-XX vv. [Some Statistical Laws of the Characteristics of Simple and Compound Sentences in Russian Scientific and Fiction Texts of 18-20th centuries]. Russkij yazyk v nacionalnoj shkole [Russian language in the national school</article-title>
          ],
          <source>1968(2)</source>
          ,
          <fpage>67</fpage>
          -
          <lpage>80</lpage>
          (
          <year>1968</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Manovich</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Software Takes Command</article-title>
          . Bloomsbury Academic, New York (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Martynenko</surname>
          </string-name>
          , G.Y.:
          <article-title>Vvedenie v chislovuyu garmoniyu teksta [The Introduction to Numeral Harmony of the Text]</article-title>
          .
          <source>St</source>
          . Petersburg State University, Saint-Petersburg (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Martynenko</surname>
          </string-name>
          , G.Y.:
          <article-title>Metody matematicheskoy lingvistiki v stilisticheskikh issledovaniyakh [Computational Linguistics Methods in the Stylistics Research]</article-title>
          . Nestor-Istoriya,
          <source>SaintPetersburg</source>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Martynenko</surname>
            ,
            <given-names>G.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherstinova</surname>
          </string-name>
          , T.Y.:
          <article-title>Chislovoj profil syujeta [The Numeric Profile of the Plot]</article-title>
          . In:
          <article-title>Proceedings of IV Congress of Russian language researchers 'Russkij yazyk: istoricheskie sudby i sovremennost' [Russian Language: Historical Fates</article-title>
          and Modern Age],
          <fpage>524</fpage>
          -
          <lpage>525</lpage>
          . Moscow State university, Moscow (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Martynenko</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherstinova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Emotional waves of a plot in literary texts: new approaches for investigation of the dynamics in digital culture</article-title>
          . In: Alexandrov,
          <string-name>
            <given-names>D.A.</given-names>
            ,
            <surname>Boukhanovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. V.</given-names>
            ,
            <surname>Chugunov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.V.</given-names>
            ,
            <surname>Kabanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Koltsova</surname>
          </string-name>
          ,
          <string-name>
            <surname>O</surname>
          </string-name>
          . (eds.)
          <source>Digital Transformation and Global Society. DTGS 2018. Communications in Computer and Information Science</source>
          ,
          <volume>859</volume>
          ,
          <fpage>299</fpage>
          -
          <lpage>309</lpage>
          . Springer, Cham (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Martynenko</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherstinova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Analytical Distribution Model for Syntactic Variables Average Values in Russian literary Texts</article-title>
          . In: Alexandrov,
          <string-name>
            <given-names>D.A.</given-names>
            ,
            <surname>Boukhanovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.V.</given-names>
            ,
            <surname>Chugunov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.V.</given-names>
            ,
            <surname>Kabanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Koltsova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Musabirov</surname>
          </string-name>
          , I. (eds.)
          <source>Digital Transformation and Global Society. DTGS 2019. Communications in Computer and Information Science</source>
          ,
          <volume>1038</volume>
          ,
          <fpage>719</fpage>
          -
          <lpage>731</lpage>
          . Springer, Cham (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Martynenko</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherstinova</surname>
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Linguistic and Stylistic Parameters for the Study of Literary Language in the Corpus of Russian Short Stories of the First Third of the 20th Century</article-title>
          . In: Ronzhin,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Noskova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Karpov</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . (eds.) R.
          <article-title>Piotrowski's Readings in Language Engineering</article-title>
          and Applied Linguistics,
          <source>Proc. of the III International Conference on Language Engineering and Applied Linguistics (PRLEAL-2019)</source>
          , Saint Petersburg, Russia, November
          <volume>27</volume>
          ,
          <year>2019</year>
          , CEUR Workshop Proceedings,
          <volume>2552</volume>
          ,
          <fpage>105</fpage>
          -
          <lpage>120</lpage>
          . RWTH Aachen University, Aachen (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Martynenko</surname>
            ,
            <given-names>G.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherstinova</surname>
          </string-name>
          , T.Y.,
          <string-name>
            <surname>Popova</surname>
            ,
            <given-names>T.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melnik</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zamirajlova</surname>
            ,
            <given-names>Y.V.:</given-names>
          </string-name>
          <article-title>O printsipakh sozdaniya korpusa russkogo rasskaza pervoy treti XX veka [On the Principles of Creation of the Russian Short Stories Corpus of the First Third of the XX Century]</article-title>
          .
          <source>In: Proceedings of the 15th TEL International Conference on Computational and Cognitive Linguistics (TEL-2018)</source>
          ,
          <volume>1</volume>
          ,
          <fpage>180</fpage>
          -
          <lpage>197</lpage>
          .
          <string-name>
            <surname>Izdatelstvo</surname>
            <given-names>AN RT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kazan</surname>
          </string-name>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Moretti</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          : Distant Reading. Verso, London (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Olmsted</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>On some axioms about sentence length</article-title>
          .
          <source>Language</source>
          <volume>43</volume>
          (
          <issue>1</issue>
          ),
          <fpage>303</fpage>
          -
          <lpage>305</lpage>
          (
          <year>1967</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Martynenko</surname>
            ,
            <given-names>G.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherstinova</surname>
          </string-name>
          , T.Y.,
          <string-name>
            <surname>Melnik</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popova</surname>
            ,
            <given-names>T.I.</given-names>
          </string-name>
          :
          <article-title>Metodologicheskie problemy sozdaniya Kompyuternoy antologii russkogo rasskaza kak yazykovogo resursa dlya issledovaniya yazyka i stilya russkoy hudozhestvenny prozy v epokhy revolutsionnykh peremen (pervoy treti XX veka) [Methodological problems of creating a Computer Anthology of the Russian story as a language resource for the study of the language and style of Russian artistic prose in the era revolutionary changes (first third of the 20th century)]</article-title>
          . In:
          <article-title>Kompjuternaya lingvistika i vychislitelnye ontologii [Computational Linguistics and Computational Ontologies]</article-title>
          .
          <source>Issue 2 (Proceedings of the XXI International Conference “Internet i sovremennoe obshchestvo” [Internet and Modern Society], IMS2018, St. Petersburg, 30 May</source>
          <year>2018</year>
          -2
          <article-title>June 2018</article-title>
          .
          <source>Collection of scientific articles)</source>
          ,
          <fpage>99</fpage>
          -
          <lpage>104</lpage>
          . ITMO University, St.
          <source>Petersburg</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Reagan</surname>
            , A.J., Mitchell,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danforth</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dodds</surname>
            ,
            <given-names>P.S.:</given-names>
          </string-name>
          <article-title>The emotional arcs of stories are dominated by six basic shapes</article-title>
          .
          <source>EPJ Data Science</source>
          <volume>5</volume>
          ,
          <issue>31</issue>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Rudnicka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Variation of sentence length across time and genre. Influence on syntactic usage in English</article-title>
          . In: Whitt,
          <string-name>
            <surname>R.J</surname>
          </string-name>
          . (ed.)
          <article-title>Diachronic Corpora, Genre, and Language Change (Studies in Corpus Linguistics</article-title>
          ,
          <volume>85</volume>
          ),
          <fpage>219</fpage>
          -
          <lpage>240</lpage>
          . John Benjamins, Amsterdam (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Silge</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>D.: Text</given-names>
          </string-name>
          <string-name>
            <surname>Mining with R. O'Reilly Media</surname>
          </string-name>
          ,
          <string-name>
            <surname>Sebastopol</surname>
          </string-name>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sherstinova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitrofanova</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skrebtsova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zamiraylova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kirina</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Topic Modelling with NMF vs. Expert Topic Annotation: the Case Study of Russian Fiction</article-title>
          . In:
          <string-name>
            <surname>Martínez-Villaseñor</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herrera-Alcántara</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponce</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castro-Espinoza</surname>
            ,
            <given-names>F.A</given-names>
          </string-name>
          . (eds)
          <source>MICAI</source>
          <year>2020</year>
          , LNCS, 12469. Springer, Cham (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Sherstinova</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skrebtsova</surname>
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Russian Literature Around the October Revolution: A Quantitative Exploratory Study of Literary Themes and Narrative Structure in Russian Short Stories of 1900-1930</article-title>
          . In: Proc. of the International Workshop “Computational Linguistics”
          <article-title>(CompLing-2020)(in print).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Sherstinova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ushakova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melnik</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Measures of Syntactic Complexity and their Change over Time (the Case of Russian)</article-title>
          . In: Balandin,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Turchet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Tuytina</surname>
          </string-name>
          , T. (eds.)
          <source>Proceedings of the 27th Conference of Open Innovations Association FRUCT</source>
          , Trento, Italy,
          <fpage>221</fpage>
          -
          <lpage>229</lpage>
          . Institute of Electrical and Electronic Engineers, New York (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Stamatatos</surname>
          </string-name>
          , E.:
          <article-title>A survey of modern authorship attribution methods</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          ,
          <volume>60</volume>
          (
          <issue>3</issue>
          ),
          <fpage>538</fpage>
          -
          <lpage>556</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Venecky</surname>
            ,
            <given-names>I.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veneckaya</surname>
            <given-names>V.I.</given-names>
          </string-name>
          :
          <article-title>Osnovniye matematiko-statisticheskie ponyatiya i formuly v ekonomicheskom analize [</article-title>
          <source>Basic Math and Statistics Concepts and Formulas in Economic Analysis]. Statistika</source>
          , Moscow (
          <year>1979</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Yule</surname>
          </string-name>
          , G.:
          <article-title>On sentence-length as a statistical characteristic of style in prose: with application to two cases of disputed authorship</article-title>
          .
          <source>Biometrika</source>
          ,
          <volume>30</volume>
          (
          <issue>3</issue>
          /4),
          <fpage>363</fpage>
          -
          <lpage>390</lpage>
          (
          <year>1939</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>