<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Statistical Description of Russian Texts: Parameters and Factors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anastasia M. Amieva</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viktor V. Filimonov</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey A. Zhivodyorov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anna A. Kramarenko</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Central Scientific Library Ural Branch of the Russian Academy of Sciences</institution>
          ,
          <addr-line>Sofia Kovalevskaya, 22 Ekaterinburg 620137</addr-line>
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ural Federal University, Department of Technical Physics</institution>
          ,
          <addr-line>Mira, 21, Ekaterinburg 620002</addr-line>
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Ural Federal University, Department of printing arts and web-design</institution>
          ,
          <addr-line>Mira, 32, Ekaterinburg 620002</addr-line>
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This work describes the parameters for the attribution of Russianlanguage texts and the methods for obtaining them. The technique of machine attribution of Russian-language texts is presented. The methodology is based on applying factor analysis to study the relationships between parameters. The work was done at the department of printing arts and web-design IRIT-RtF UrFU.</p>
      </abstract>
      <kwd-group>
        <kwd>Modeling</kwd>
        <kwd>Attribution of texts</kwd>
        <kwd>Statistics χ2</kwd>
        <kwd>Law of large numbers</kwd>
        <kwd>Diffusion coefficient</kwd>
        <kwd>Factor analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This work represents a step towards the construction of a theoretical model of written
text. We think that, first of all, it is necessary to understand the laws of text
construction and on their basis to construct a theory for describing the text. The laws of text
construction should correspond the following requirements:
- be expressed numerically;
- be summarized for different texts;
- be subjected to formal mathematical analysis and / or modeling.</p>
      <p>
        With such a formulation of the problem, our research associate with the problem of
interdisciplinary interactions in science (Petrov V.M.) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Modern digital
technologies and natural scientific concepts penetrate deeply into the humanities. Within the
natural science process, areas associated with humanitarian objects are formed
(cognitive studies, neuro-esthetics). In our case, it is a Russian-language texts.
      </p>
      <p>Researches of a text are traditionally conducted in three directions: spatial,
linguistic and statistical.</p>
      <p>
        Spatial direction involves measuring the length of a line, leading, font size, etc.
(Artyomov V.A. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Ushakova M.N. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Tarasov D.A. [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ], Tyagunov A.G.
[
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], Sergeev A.P. [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ], Filimonov V.V. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Weber A. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Cohn H. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and
others).
      </p>
      <p>
        The linguistic direction includes the study of meaning-bearing units (sentences,
paragraphs) and structural features (Matezius V. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], Skalicka V. [10], Halliday
M.A.K. [11], PalmerF.R. [12], Martinet A. [13], Tenyer L. [14], Lotman Yu.M.
[15, 16], Mintz Z.G. [17], and others).
      </p>
      <p>Statistical studies are related to the study of quantitative characteristics of texts.
Studies of quantitative parameters of texts conduct for a long time. The first
theoretical result in the field of statistical studies of the text is the empirical «Zipf law». It can
be formulated as follows: «The product of the frequency of occurrence of a word and
its position in the frequency dictionary is approximately constant value»1.
Quantitative parameter in the Zipf's law is the frequency of occurrence of words in the text.</p>
      <p>In 1991, the American researcher Wentian Li proved that this law is fulfilled for
any random sequence of symbols. Thus, he suggested that the law is a statistical
phenomenon. And this phenomenon is not connected with the semantics of the text [18].</p>
      <p>At this stage of the work, we develop a technique for machine attribution of texts.
The technique can be used to determining authorship and evaluate the usability of the
text.</p>
      <p>We build our research on the following assumptions:
1. there are hidden structural elements in the text. They can be detected by methods
of mathematical modeling and mathematical statistics;</p>
      <p>2. The author and the reader are interpreting the text differ-ently because they have
difference life experiences [19]. So we excluded «meaning» from consideration for
the objectivity of the study.</p>
      <p>Our works [19–23] was dedicated to studies of texts by the methods of frequency
analysis and using the model of random walks. We considered the repeatability of
individual letters and of their triples (the three vowel letters). By triples we mean
three vowel letters that consistently appear in the text. These approaches were chosen
by us, because they:
- are objective;
- allow to get rid of subjective and conventional effects;
- can be associated with digital data processing [19].</p>
      <p>We base on the results of the conducted studies and we can say that the results of
our work can be used to develop algorithms for machine attribution of texts without
preliminary expert evaluation and without taking into account the meaning.</p>
      <p>In the course of research, we have obtained and used a number of parameters.
These parameters can be used as text attributes.</p>
      <p>We use factor analysis to investigate the relationships between parameters in this
work.</p>
      <p>Studies were conducted on Russian-language texts (about 1,500 texts of various
genres and directions) [21].
where   1,   2,   3 are the frequencies of occurrence of the first, second and third
letters in the triples, which are calculated using the following formula:
where   is the number of emergence of the individual vowel in the text.</p>
    </sec>
    <sec id="sec-2">
      <title>Obviously, that  theor cannot be equal to zero.</title>
    </sec>
    <sec id="sec-3">
      <title>The observed amount of triples (</title>
      <p>emp) was obtained by recalculating all triples in
many versions of the triples are missing. Those  emp can be equal to zero.
the text. Recalculation was carried out using the program «Qlines»2. In real texts,</p>
      <p>We used the statistics χ2 to estimate the difference in the calculated distribution of
the triples from their observed distribution:</p>
      <p>theor =  theor ∙ 
 theor =   1 ∙   2 ∙   3
   =
 

(1),
(2),
(3),
(4).</p>
      <sec id="sec-3-1">
        <title>Parameters of text</title>
        <p>This section provides brief descriptions of parameters of the text. The methods for
measuring the parameters are presented in the corresponding articles. References to
the articles are indicated in the text.</p>
        <p>Technique for studying texts with using the χ2 statistics was presented in [20, 21].
We compared the calculated and observed number of triples of the vowels in the texts.</p>
        <p>The calculated quantity was calculated using the formula:
where  theor is the frequency of the appearance of the triples of vowel letters, N is
the number of vowels in the text:</p>
        <p>2 = ∑ =1
( theor− emp)2

 theor
plied that the values of χ2 for different texts could be compared with each other.</p>
        <p>The values of χ</p>
        <p>2 were calculated for approximately 1,500 texts. The values
obtained were distributed as follows:
1. from 0 to 2 000 there are not texts;
2. from 2 000 to 2 400 there are only poetic works;
3. the majority of fiction texts are located in the range from 2 000 to 6 000;
4. the majority of scientific texts are located in the range from 6 000 to 12 000;
5. the religious, socio-political and journalistic texts are located in the range from
4 000 to 12 000;</p>
        <p>6. the administrative texts are located in the range from 10 000 to 30 000.
2 Specially written for research by the staff of the Central Scientific Library Ural Branch of the
Russian Academy of Sciences L.G. Gorbich</p>
        <p>Thus, the texts were grouped in different groups. The groups basically coincided
with the expert classification of texts. The machine did not focus on the meaning of
the text and its name, but only analyzed the sequence of signs.</p>
        <p>The next stage of the study [22] was the search for additional text attribution
parameters. It was suggested that «the differences between the values of the χ2 statistic
are random and can be related to the finiteness of the length of the text» [22]. The
longer the text length, the smaller the standard deviation (SD) of the values χ2. In the
case of «sufficiently large» texts, SD asymptotically tends to zero according to the
law of large numbers:
  = √cN
(5),
where c is the coefficient of proportionality.</p>
        <p>The coefficient c does not depend on the number of vowels N and can be related to
the peculiarities of the text itself. Thus, the coefficient c can be an attribute of the text
or be characteristic of the language as a whole.</p>
        <p>As a result of the study it was found that:
1) 1,5 &lt;  &lt; 5 for the main part of the texts considered;
2) religious (0,3 &lt;  &lt; 3) and administrative texts (5 &lt;  &lt; 38) are distinguished.</p>
        <p>The second approach [23] is based on the use of the mathematical model of
random walks.</p>
        <p>In constructing the model, an analogy was made between the thermal motion of
particles in Euclidean space and the displacement of a phase point in the state space.</p>
        <p>Einstein's law states that the mean square of displacement of a Brownian particle in
the absence of external forces is directly proportional to time. For the
twodimensional case, the law can be written as follows:
 ̅2 = 4
(6),
where R is the displacement, D is the coefficient similar to the diffusion coefficient
for the physical system (hereinafter diffusion coefficient), t is the time (corresponds to
the ordinal number of the letter from the beginning of the text).</p>
        <p>The text is modeled as the displacement of the phase point according to certain
transition rules. Each vowel is associated with a definite vector. Its length is
determined by the inverse frequency of occurrence the letter. Directions of vectors
corresponding to the various letters are not the same and are distributed uniformly over the
circumference with intervals 40.</p>
        <p>The coefficients D were calculated for more than 100 texts. The values were
distributed over two ranges:  1 &lt; 124,  2 &gt; 124. Poetic, fiction, scientific and
journalistic texts are located in the first range. Administrative texts are located in the second
range. Religious texts are located on the border between these ranges. Religious texts
are located on the border between these ranges.</p>
        <p>It turned out that the dependence of the mean square of displacement from time is
ambiguously approximated by a straight line. Therefore, a relative correction to the
Einstein law (RC) was analyzed. RC illustrates the difference between a real process
and strictly random and has unique meaning for each text:
where a2 the coefficient for senior term of the polynomial of the second degree, and a1
is the coefficient for senior term of the straight line.</p>
        <p>The values of RC were distributed over three ranges:  1 &lt; 12%,
12% &lt;  2 &lt; 24%,  3 &gt; 24%. Fiction and journalistic texts are located in the
first range, scientific and religious texts are located in the second range, and
administrative texts are located in the third range. Poetry is located on the border between the
first and second ranges.
3</p>
      </sec>
      <sec id="sec-3-2">
        <title>Factor analysis</title>
        <p>and reduces the number of variables that are used to describe3. 16 parameters were
grouped to five factors as a result of our study. At that the three parameters were
insignificant (did not enter into in any factor), and the others were grouped into five
factors.</p>
        <p>Table 1 shows the meanings of factor loadings. The value of factor loads greater
than 0,7 is standard when deciding whether to include a parameter in the factor. The
values of factor loads for the parameters included in the factor are highlighted in bold
type.</p>
        <p>The diffusion coefficient D, the year the text was created, the frequencies of the
letters «и» and «э» are included in the first factor.</p>
        <p>The number of vowels, the coefficient of proportionality of the law of large
numbers c, frequency of the letter «a» are included in the second factor.</p>
        <p>The frequencies of the letters «e» and «o» are included in the third factor.
The frequencies of the letters «y» and «ю» are included in the fourth factor.
The frequencies of the letters «ы» and «я» are included in the fifth factor.
Thus, we got the opportunity to classify texts with a probability of 88%.</p>
        <p>The distribution of parameters by factors is not interpreted by us at this stage of the
study, since insufficient number of data was used.
3http://studme.org/79025/psihologiya/vidy_protsedury_faktornogo_analiza</p>
      </sec>
      <sec id="sec-3-3">
        <title>Conclusion</title>
        <p>The work describes methods for obtaining parameters for the attribution of
Russianlanguage texts. Also, the use of factor analysis for these purposes is considered.</p>
        <p>It was possible to reduce the number of variables from 16 to 5 with the help of
factor analysis. The method allows you to classify texts with a probability of 88%.</p>
        <p>Studying the interrelationships between the parameters will make it possible to
obtain a more accurate classification of Russian-language texts according to their
direction (poetry, fiction prose, scientific, journalistic, administrative and religious texts).</p>
        <p>Classification can be used for machine processing of text, which will solve a
number of problems:
1) construction of a mathematical model of the text;
2) research tasks (determining authorship, etc.);
3) assessment of readability (usability);
4) accounting for text specificities for editing and layout;
5) automated search for texts of a given direction in arrays of heterogeneous and
poorly structured information.
10. Skalichka, V. Asymmetric dualism of linguistic units / V. Skalichka // History of linguistics
of XIX-XX centuries in essays and extracts. — M.: Flint, 2012. — P. 54–63.
11. Halliday, MAK. Place of the functional perspective of the proposal (FPP) in the system of
linguistic description [Electronic resource]. — Access mode:
http://socialtranslation.ru/article.php?article_id=548 (circulation date is 2.04.2016).
12. Palmer, F.R., Mood and modality [Electronic resource]. — Access mode:
www.academia.edu/3704213/Irreilis__ and_reality (circulation date is 7.04.2016)
13. Martinet, A. A functional view of language [Electronic resource]. — Access mode:
https://www.questia.com/library/3057879/a-functional-view-of-language (circulation date is
19.03.2016).
14. Tenyer, L. Cours elementaire de syntaxe structurale [Electronic resource]. — Access mode:
http://www.classes.ru/grammar/172.Tesniere/ source / worddocuments / _1.htm (circulation
date is 19.03.2016).
15. Lotman, Yu.M. Inside the thinking worlds [Electronic resource]. — Access mode:
http://lms.hse.ru/content/lessons/64192/%D1% 81% D0% B5% D0% BC% D0% B8% D0%
BD% D0% B0% D1% 80% 207 -8 / lotman_yu
_m_izbrannye_stati_v_treh_tomah_tom_1_stati_po_semi (circulation date is 3.05.2016).
16. Lotman, Yu.M. Articles on semiotics of culture and art [Electronic resource]. — Access
mode: http://libatriam.net/read/899058/0/ (circulation date is 10.05.2016).
17. Mintz, Z.G. The structure of the sentence and the typology of artistic texts [Electronic
resource]. — Access mode: http://www.studfiles.ru/view/ 3911949/ (circulation date is
16.05.2016).
18. Wentian, Li. Random Texts Exhibit Zipf's-Law-Like Word Frequency Distribution. — Santa</p>
        <p>Fe Institute, 1991. — P. 1–8.
19. Amieva, A.M. Machine attribution of Russian-language texts: a review of methods /
А.М. Amieva, A.A. Kramarenko, V.V. Filimonov, A.A. Zhivodyorov // Materials of the X
International Scientific and Practical Conference (Ekaterinburg, 27 February–3 March 2017)
New information technologies in education and science. Ekaterinburg: RGPPU — 2017. —
P. 371–375.
20. Filimonov V.V., Zhivodyorov A.A., Gorbich L.G. Expression and order in written speech //
Izvestia UrFU. Series 1 The problems of education, science and culture. — 2012. — №3
(104). — P. 313–319.
21. Filimonov V.V., Amieva A.M., Sergeev A.P. Clustering of Russian-language texts using χ2
statistics. // Proceedings of the International Scientific and Practical Conference
(Ekaterinburg, January 12–13, 2016). Information: transmission, processing, perception. Ekaterinburg:
UrFU named after the first President of Russia B.N. Yeltsin — 2016. — P. 164–174.
22. Filimonov V.V., Amieva A.M., Zhivodyorov A.A., Kramarenko A.A. Attribution of
Russian-language texts using the law of large numbers. Proceedings of the International
Scientific and Practical Conference (Ekaterinburg, January 12–13, 2017). Information:
transmission, processing, perception. Ekaterinburg: UrFU named after the first President of Russia
B.N. Yeltsin — 2017. — P. 10–18.
23. Kramarenko A.A., Filimonov V.V., Zhivodyorov A.A., Amieva A.М. Application of the
random walk model for describing Russian-language texts. Proceedings of the International
Scientific and Practical Conference (Ekaterinburg, January 12–13, 2017). Information:
transmission, processing, perception. Ekaterinburg: UrFU named after the first President of
Russia B.N. Yeltsin — 2017. — P. 138–164.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Petrov</surname>
            ,
            <given-names>V.M.</given-names>
          </string-name>
          <article-title>Art history and exact sciences? // Proceedings of an international scientific and practical conference dedicated to the memory of Herman Alekseevich Golitsyn (September 20-</article-title>
          22,
          <year>2012</year>
          )
          <article-title>Quantitative methods in art history</article-title>
          .
          <source>Ekaterinburg: Artifact</source>
          ,
          <year>2013</year>
          . - S. 6-
          <fpage>7</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Artyomov</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          <article-title>Technographic analysis of the total letters of the new alphabet [Electronic resource]</article-title>
          . - Access mode: http://nowa.cc/printthread.php?
          <source>t=282586 (circulation date is 18.05</source>
          .
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ushakova</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          <article-title>New font for newspapers [Electronic resource]</article-title>
          . - Access mode: http://www.liveinternet.ru/users/1013940/post267598201/ (circulation date is
          <volume>22</volume>
          .05.
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Tarasov</surname>
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akhmetova</surname>
            <given-names>A.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sergeev</surname>
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tyagunov</surname>
            <given-names>A.G.</given-names>
          </string-name>
          <article-title>Substantiation and derivation of the formula for the speed of reading based on the spatial characteristics of textual information //</article-title>
          <source>Proceedings of the International Scientific and Practical Conference (Ekaterinburg</source>
          ,
          <year>2015</year>
          ) Information Technologies, Telecommunications and
          <string-name>
            <given-names>Control</given-names>
            <surname>Systems</surname>
          </string-name>
          . Ekaterinburg:
          <article-title>UrFU named after the first President of Russia BN</article-title>
          . Yeltsin,
          <year>2015</year>
          . - P.
          <fpage>140</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Tarasov</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          <article-title>The account of spatial characteristics of a strip of a typing in the formula</article-title>
          of reading / D.A.
          <string-name>
            <surname>Tarasov</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          <string-name>
            <surname>Sergeev</surname>
          </string-name>
          , A.G. Tyagunov // News of Higher Educational Establishments.
          <article-title>Problems of polygraphy and publishing</article-title>
          .
          <source>- 2014</source>
          . - No. 6. - P.
          <fpage>3</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Tarasov</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          <article-title>Legality of Textbooks: a literature review / D.A</article-title>
          .
          <string-name>
            <surname>Tarasov</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          <string-name>
            <surname>Sergeev</surname>
            , V.V. Filimonov // Procedia - Social and
            <given-names>Behavioral</given-names>
          </string-name>
          <string-name>
            <surname>Sciences</surname>
          </string-name>
          .
          <article-title>-</article-title>
          <year>2015</year>
          . - P.
          <fpage>1300</fpage>
          -
          <lpage>1308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Weber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <article-title>Ueber die Augenuntersuchungen in den hoheren schulen zu Darmstadt [Electronic resource]</article-title>
          . - Access mode: http://www.pearsonified.com/
          <year>2011</year>
          /12/golden-ratiotypography.
          <source>php (circulation date is 30.05</source>
          .
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cohn</surname>
          </string-name>
          , H. Die Hygiene des Auges in den Schulen. [Electronic resource]. - Access mode: http://www.pearsonified.com/
          <year>2011</year>
          /12/golden-ratio-typography.
          <source>php (circulation date is 30.05</source>
          .
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Matezius</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <article-title>On the potentiality of linguistic phenomena / V. Matezius // Selected works on linguistics. Translation from Czech</article-title>
          and
          <string-name>
            <surname>English - M.</surname>
          </string-name>
          ,
          <year>2012</year>
          . - P.
          <fpage>3</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>