<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Decisions of Russian Constitutional Court: Lexical Complexity Analysis in Shallow Diachrony</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>St. Petersburg State University</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Universitetskaya nab.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>St. Petersburg</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>o.blinova</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>s.a.belov</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>m.revazov}@spbu.ru</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University Higher School of Economics</institution>
          ,
          <addr-line>190068, St. Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The paper is aimed at studying the texts of Russian Constitutional Court decisions, issued from 1992 to 2018. We analyzed the corpus, consisting of 584 decisions or 3,426,747 tokens (incl. punctuation marks) and tested the hypothesis about increasing lexical complexity of the documents. Using the R package stylo and MFW statistics, we got a picture that reflects the differences of the texts by years. The results of cluster analysis show that the texts of the 90s and 2000s are combined into the first large cluster. The second large cluster includes the texts of the 2010s. Using the R package quanteda, we obtained the values of 11 lexical diversity measures. We chose the index K (Yule's K) as a basic measure, relatively more reliable and independent of the text length, and then interpreted the values of this measure. In general, the value of K decreases over the years, except for the texts of 2006, in which there is a noticeable increase in the index value, and the texts of 1993, in which the outlier is observed. The calculation hapax proportion shows a picture of a gradual decrease in the share of hapaxes. If we apply the traditional approach to the interpretation of TTR values and derived metrics, we can conclude that, as the lexical diversity decreases and the proportion of hapaxes decreases, the texts become easier to read.</p>
      </abstract>
      <kwd-group>
        <kwd>Legal Linguistics</kwd>
        <kwd>Decisions of the Russian Constitutional Court</kwd>
        <kwd>Stylometric Analysis</kwd>
        <kwd>Most Frequent Words</kwd>
        <kwd>Lexical Complexity</kwd>
        <kwd>Lexical Diversity</kwd>
        <kwd>TTR</kwd>
        <kwd>Yule's K</kwd>
        <kwd>R packages</kwd>
        <kwd>stylo</kwd>
        <kwd>quanteda</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>This paper is aimed at studying the texts of Russian Constitutional Court decisions,
issued from 1992 to 2018. The purpose is to verify the hypothesis, according to which
the texts became more complex during the specified period (i.e. in shallow
diachrony). At the moment, we were primarily interested in the lexical complexity.</p>
      <p>The Constitutional Court is one of the youngest legal institutions in Russia. Its
appearance in 1991 was associated with large-scale changes in the legal and political</p>
      <p>Copyright ©2020 for this paper by its authors.</p>
      <p>Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
system caused by the rejection of Soviet legal and political system and the creation of
a new democratic state in Russia.</p>
      <p>The Constitutional Court, more than any other courts, was perceived as an “alien
body” in the judicial system, since the Court was significantly different from all other
judicial bodies in its objectives and duties. The task of the Constitutional Court was
neither to solve a specific case, nor to draw conclusions about the rights and
obligations of a particular citizen, but to compare the norms of a law challenged by a citizen
and the provisions of the Constitution, ensuring the protection and implementation of
constitutional principles.</p>
      <p>The major function of the Constitutional Court is direct application and appropriate
interpretation of the constitutional text. Citizens and legal entities’ complaints about
the violation of their constitutional rights noticeably prevail among the cases
considered by the Russian Constitutional Court. For example, in 2019, out of more than
3,500 judgments and rulings of the Constitutional Court, only three were made at the
request of state bodies, about 30 were made at the request of the courts, and all the
rest were made on complaints.</p>
      <p>At the same time, it cannot be concluded that decisions of the Constitutional Court,
initiated by citizens, are addressed specifically to these citizens. Of course, an
applicant should understand whether the Constitutional Court supported her arguments as
well as the outcome of the case consideration. However, only a small (operative) part
of а Constitutional Court decision is devoted to this.</p>
      <p>The rest is primarily addresses those bodies that must restore the violated rights,
and not only and not so much the citizen who applied directly to the Constitutional
Court, as those who find themselves in a similar situation. In this regard, the main
addressees of Constitutional Court decisions are the legislative bodies, which should
amend corresponding laws, and law-enforcement bodies (both executive and
judiciary), which should interpret the legal provisions that they apply in accordance with
decisions of the Constitutional Court.</p>
      <p>However, it would be wrong to exclude citizens and legal entities from the
addressees of the decisions. The Constitutional Court very rarely finds itself concluding
the absolute, complete and unconditional contradiction of examined norms of the
Constitution. More often, conclusions about unconstitutionality are made in relation
to a particular interpretation of the impugned norm. Acts of the Constitutional Court
become part of existing law, shall be applied along with statutes, the interpretation of
which they strongly influence. Coming to court and demanding application of a
provision, for example, establishing a social payment, citizens often must refer not only to
the provisions of the law, but also to their constitutional interpretation in the practice
of the Constitutional Court. That is why the decisions of the Constitutional Court
should be clear to all citizens.</p>
      <p>On the one hand, the field of activity for the Constitutional Court is an area of
refined jurisprudence, free from description of factual circumstances, proving their
existence and assessment of such evidence; on the other hand, it is a part of the
existing legal regulation, along with statutes. The decisions of the Constitutional Court
have the most significant difference with the latter, as these decisions are a result
of the work of judges (i. e. professional lawyers). Many of the Constitutional Court
judges are professors of law, others were appointed to the Constitutional Court after
many years career in other courts.</p>
      <p>A draft of each decision is prepared by a judge-rapporteur, while the final text of
any decision becomes a result of the collective creativity of all judges and has no
authorship.</p>
      <p>That is why the assessment of the decisions’ complexity is important and
indicative. Such an assessment demonstrates the ability of professional lawyers to be
clear, to write specialized texts addressing to a wide range of citizens, in an
accessible manner.
1</p>
    </sec>
    <sec id="sec-2">
      <title>This Paper’s Structure and Recent Works</title>
      <p>
        To test the hypothesis of texts becoming more complex with time, we used the
capabilities of two R software packages (stylo and quanteda) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Both packages
allow to analyze non-structured text data.
      </p>
      <p>Using the stylo package, we got a general picture that reflects the differences
between the texts by year. Using the quanteda package, we received more detailed
information about the lexical complexity of texts by different time periods.</p>
      <p>
        The diachronic study of legal documents is the actively developing area. In
particular, there are diachronic corpora of legal texts, for example, Corpus of Historical
English Law Reports (CHELAR) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A study of the texts of Russian legal documents in
dynamics was carried out in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        A research on the readability of texts of Constitutional Court is presented in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
The corpus of decision was analyzed using a simple readability metric, the
FleschKincaid formula, adapted for the Russian by I.V. Oborneva [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The Flesch-Kincaid formula for the Russian looks as follows:</p>
      <p>FRE = 206,836 − 60,1 × ASL − 1,3 × ASW,
(1)
where ASL is the average sentence length in words, and ASW is the average word
length in syllables.</p>
      <p>
        However, two points should be emphasized. Firstly, the coefficients of the
Oborneva’s formula were obtained by calculating the statistical characteristics of about 100
works of famous English-language literary classics (and translating these works into
Russian). Thus, the formula is not quite universal, but is applicable primarily for the
analysis of the complexity of texts of (translated) fiction; about the indicated problem,
see [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Secondly, recent studies show that “sentence and word length measures
likely do not tap directly into linguistic components related to readability … nor
are they the only linguistic features related to readability” [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>In this paper to assess the text complexity (lexical complexity) we also use the
traditional method of assessing, calculating the TTR (type-token ratio), more precisely,
we use a number of derived metrics.</p>
      <p>
        TTR is the ratio of the number of unique tokens (types) to all document tokens. It
is known, however, that the values of the TTR measure are not independent of text
length, that is, documents of equal length should be compared to obtain relevant
results. The solution to this problem can also be the use of derivative measures, see
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Such measures are provided in the quanteda package.
      </p>
      <p>In addition, to assess the lexical complexity, we use information on hapax richness
and hapax proportion (the hapax is a token that appears in a sample once). Hapax
richness is a measure that describes text from the same perspective as the TTR
measure.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Methods for Assessing Text Complexity and Lexical</title>
    </sec>
    <sec id="sec-4">
      <title>Complexity</title>
      <p>
        There is a rather long tradition of applying methods for assessing complexity
(readability) to texts in Russian; for a review, see, for example, [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. There is, among others,
the traditional direction mentioned above, associated with the use of a wide variety of
readability formulas. Only 5 or 7 readability formulas were adapted for Russian [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. So, the following metrics are used on the “LeStCor: Levelled Study Corpus
of Russian” resource: Flesch-Kincaid grade level, Coleman Liau Index score,
(Gunning) Fog, SMOG index, Automated Readability Index, New Dale Chall Adjusted
Grade Level, Powers-Sumner-Kearl Grade Level [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        To assess the lexical complexity, we can use information on:
• lexical density, the proportion of various content words in the texts;
• lexical richness and lexical diversity, measured by calculating the values of TTR or
derived measures;
• number of words with abstract or concrete meaning;
• number of ambiguous words;
• number of function words (particularly prepositions);
• number of abbreviations
etc., see [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and many others. TTR values well predict Russian text
complexity, see [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Material</title>
      <p>We analyzed a collection of judgments, consisting of 584 documents relating to the
period since 1992, when the first judgment of the Court appeared. The distribution of
decisions by year is described in Table 1.</p>
      <p>
        The full texts of the decisions were taken from the database of the ConsultantPlus
information system [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and from the web-portal of the Constitutional Court [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>There are no 1994 decisions in the text collection. This is due to the fact that the
Constitutional Court suspended work at the end of 1993. The reason was the need to
adopt a new law, regulating the Constitutional Court activities. As a result, in 1994
the Federal Constitutional Law “On the Constitutional Court of the Russian
Federation” was adopted. The appointment of judges also took some time. In an updated
form, the Constitutional Court resumed its work in 1995.
We used the corpus, which contains texts combined by years. Accordingly, the corpus
files received names like “1992”, “2003”, “2018”.</p>
      <p>The text collection consists of 3,426,747 tokens (including punctuation), see Table 2.</p>
    </sec>
    <sec id="sec-6">
      <title>MFW Statistics</title>
      <sec id="sec-6-1">
        <title>Analysis Procedure in stylo</title>
        <p>
          The stylo package was created for quantitative studies of writing style and can be used
in authorship verification (including forensic linguistics) and diachronic studies,
see [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          We performed unsupervised multivariate analysis. Using the basic functions does
not imply preprocessing and markup of the text collection (segmentation into
sentences, lemmatization, etc.). We downloaded text data directly from the corpus files. Text
metadata were included in the file names. Using the basic stylo() function, it is
possible to analyze a corpus with the assistance of the following methods.
• Cluster analysis (CA), the results of which are visualized as a dendrogram, or a
graph, showing the clustering of texts.
• Multidimensional scaling (MDS), as a result of which texts are displayed as
ordered on the basis of several variables, so that similar texts are placed next to each
other, and heterogeneous texts are separated, see [
          <xref ref-type="bibr" rid="ref19 ref23">23, 19</xref>
          ].
• Principal Component Analysis (PCA), which operates on the covariance between
features (PCV) [Ibid, 933].
• Principal Component Analysis (PCA), which operates on the correlation
coefficient matrix between features (PCR) [Ibid, 933].
• Building a Bootstrap Consensus Tree (BCT), summarizing various cluster analysis
results based on the most frequent features occurrences and culling parameter
values [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>So, by means of the package one can find out how much the analyzed texts or text
collections differ. As features for the analysis, n-gram sequences of tokens and
characters can be used.
4.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Analysis Results</title>
        <p>
          At the stage of preprocessing, we performed tokenization and removal of stop words.
When tokenizing, we used the built-in features of the package. To remove stop words,
we took a stop word list from [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].1
        </p>
        <p>
          The corpus size after the removal of stop words was 2,103,608 tokens. We formed
a list of 1000 frequent features, then found the features that are used in at least 90% of
the texts. In this way, we got a list of 1684 MWF, and then in the analysis we used
100 or from 100 to 1600 of them.
1 We used a list of stop words to remove units that are not able to characterize the lexical
peculiarity of a text or text collection. The list [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] consists of 159 high-frequency words and
includes primarily function words, as well as some most common nouns, adverbs and verbs
(человек ‘person’, говорил ‘said’ etc.).
        </p>
        <p>
          Then we performed cluster analysis, multidimensional scaling, principal
component analysis. As a measure of distance, where relevant, we used the Eder’s Delta
measure, which is recommended for highly inflected languages [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ].
        </p>
        <p>The results of the PCA, MDS, and CA (see Fig. 1 below) show that the texts of the
90s, 2000s, and 2010s can be described as two separate groups. One can see the
following pattern: the texts of the 90s and 2000s are combined into the first large cluster.
The second large cluster combines primarily the texts of the 2010s.</p>
        <p>Thus, MFW statistics shows, that in general the texts before 2010 and after 2010
are clearly opposed, but the texts of 2005 and 2007 are adjacent to the group of texts
of the 2010s. In addition, the texts of 1992 and 1993 are opposed to the rest of the
texts written before 2010.
On the whole, the results of calculating the Euclidean distance on normalized token
frequency demonstrate a similar, but non identical patterns, see Fig. 2 (the
dendrogram displaying normalized token frequency was obtained using quanteda package).
The texts of 1992 and 1993 are contrasted with the rest of the texts in the corpus; in
addition, the texts of 2012 fell into a large cluster containing the remaining texts of
the 1990s and texts of the 2000s.</p>
        <p>Finally, BCT allowed us to obtain the combined results of a cluster analysis (see
Fig. 3).
5
5.1</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Measures of Lexical Diversity</title>
      <sec id="sec-7-1">
        <title>Analysis Procedure in quanteda</title>
        <p>
          The quanteda package provides tools for a range of natural language processing tasks,
see [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], it allows to perform tokenization, stemming, n-grams forming, selection and
weighing of features [Ibid].
We used the package capacities, related to the lexical diversity assessment. More
specifically, we used TTR calculation and calculation of derived measures such as
Herdan’s C (С), Guiraud’s Root TTR (R), Carroll’s Corrected TTR (CTTR), Dugast’s
Uber Index (U), Summer’s index (S), Yule’s K (K), Herdan’s Vm (Vm), Maas’
indices (Maas, logV0, logeV0). The variables in all formulas are the number of types (V),
the number of tokens (N), as well as fv (i,N), that is, the number of types occurring i
times in a sample of length N [Ibid]. In addition, we calculated the amount and
proportion of hapaxes.
        </p>
        <p>
          When forming the corpus, we performed the removal of stop words, numbers and
punctuation marks (since by default numbers were considered as tokens). The
package uses the “Snowball” list of stop words [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ].
5.2
        </p>
      </sec>
      <sec id="sec-7-2">
        <title>Analysis Results</title>
        <p>
          Using the package, we obtained the values of 11 measures of lexical diversity listed
above.
It is well known that the value of a simple TTR is affected by the text (or the sample)
length, see for example [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ], [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ], and many others. This problem can be solved
in three ways.
        </p>
        <p>
          1. It is possible to use samples of the same length.
2. It is possible to apply formulas with logarithms or with other transformations
of variables N and V (Herdan’s C, Guiraud’s Root TTR, Carroll’s Corrected
TTR, Dugast’s Uber Index, Maas’ indices).
3. It is possible to apply measures that make use of elements of the frequency
spectrum (for example, the “Yule’s K” measure), e.g. measures that take into account
the number of hapax legomena (as in the “Honoré’s R” measure) or hapax
dislegomena, see [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] for more details. We used the second and third
possibilities, see Table 3.
        </p>
        <p>The values of the indices TTR, C, S, K, Vm, Maas demonstrate a general decrease in
time. The values of the indices R, CTTR, logV0 demonstrate a general increase
in time. In general, the interpretation of the data is quite tricky, since the values of
different measures are somewhat contradictory.</p>
        <p>
          Therefore, based on the findings of [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], we chose the K (Yule’s K) index as
a basic measure, relatively more reliable and independent of the text length, and then
interpreted the values of this particular measure.
        </p>
        <p>The value of the index K varies in the range from 23.40 to 83.69 (and the value of
83.69 observed in 1993 should be considered an outlier, see Fig. 4 and 5). In general,
we can say that the value of K decreases over the years (except 2006 and 1993).
Accordingly, it can be argued that the lexical diversity of the texts of in time is
decreasing.</p>
        <p>The calculation of hapax richness (the number of tokens that appear in the sample
only once) and the proportion of hapaxes show that the share of hapaxes gradually
decreases, see Table 4, Fig. 6. However, the texts of 2006, 2008 and, to a lesser
extent, 2007, 2005 and 2016 do not correspond to this general scheme.
Thus, the share of hapaxes decreases over the years, the lexical diversity decreases
over the years (texts have more and more repeating words). These two text evaluation
options are easy to interpret in a consistent manner.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusion</title>
      <p>As a result of analyzing the corpus of Constitutional Court decisions we found out the
following.
• MFW statistics shows that, in general, the texts before 2010 and after 2010
inclusive are clearly opposed.
• The texts of 1992 and 1993 are contrasted with all other texts (see, in particular,
the clustering results after calculating the Euclidean distance on normalized token
frequency). This can be explained by the fact that in 1994 the composition of the
Constitutional Court was updated.
• The hapax proportion decreases over the years, the lexical diversity of texts also
decreases.</p>
      <p>If we apply the traditional approach to the interpretation of TTR values and derivative
metrics, we can make a general conclusion that, since the lexical diversity is reduced
and the proportion of hapaxes decreases, the texts become easier to read. Thus, our
hypothesis of an increase in lexical complexity has not been confirmed.</p>
      <p>
        There is an opposite approach to the interpretation of TTR values, we quote: “a lot
of formal repetitions of the same words denoting legal entities and various legal terms
interfere with the perception of the meaning of the sentence. In this case, we can say
that reducing diversity not only does not simplify the text, but also causes the opposite
effect” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Legal texts use many repetitions. Though the presence of repetitions tires, it also
allows to avoid problems with the interpretation of coreferential expressions. In
addition, the process of text perception is affected by the priming effect (in particular,
lexical priming).</p>
      <p>Apparently, for the successful application of vocabulary-based measures for text
complexity assessment, it is necessary to take into account at least some words’
characteristics, that is, their semantics (first of all, abstractness/concreteness), their
belonging to a certain part-of-speech class and general-language frequency.
Acknowledgement. The research was supported by the Russian Science Foundation, project
#19-18-00525 “Understanding official Russian: the legal and linguistic issues”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R Core</given-names>
            <surname>Team. R:</surname>
          </string-name>
          <article-title>A language and environment for statistical computing</article-title>
          .
          <source>R Foundation for Statistical Computing</source>
          , Vienna, Austria, (
          <year>2018</year>
          ), https://www.R-project.
          <source>org/.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Eder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rybicki</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kestemont</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Stylometry with R: a package for computational text analysis</article-title>
          .
          <source>R Journal</source>
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <fpage>107</fpage>
          -
          <lpage>121</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Benoit</surname>
            ,
            <given-names>K</given-names>
          </string-name>
          , Watanabe,
          <string-name>
            <surname>K</surname>
          </string-name>
          , Wang,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Nulty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Obeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Matsuo</surname>
          </string-name>
          , A.:
          <article-title>quanteda: An R package for the quantitative analysis of textual data</article-title>
          .
          <source>Journal of Open Source Software</source>
          ,
          <volume>3</volume>
          (
          <issue>30</issue>
          ),
          <volume>774</volume>
          (
          <year>2018</year>
          ). DOI:
          <volume>10</volume>
          .21105/joss.00774.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Fanego</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodríguez-Puente</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>José</surname>
            López-Couso,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Méndez-Naya</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NúñezPertejo</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blanco-García</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamaredo</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>The Corpus of Historical English Law Reports 1535-1999 (CHELAR): A resource for analysing the development of English legal discourse</article-title>
          .
          <source>ICAME Journal</source>
          ,
          <volume>41</volume>
          ,
          <fpage>53</fpage>
          -
          <lpage>82</lpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kuchakov</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saveliev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The complexity of legal acts in Russia: Lexical and syntactic quality of texts: analytic note</article-title>
          . European University at Saint Petersburg, St.
          <source>Petersburg</source>
          (
          <year>2018</year>
          ). [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Saveliev</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuchakov</surname>
            ,
            <given-names>R.K.</given-names>
          </string-name>
          :
          <article-title>Decisions of arbitration courts of Russian Federation: lexical and syntactic quality of texts, analytic note</article-title>
          . European University at Saint Petersburg, St.
          <source>Petersburg</source>
          (
          <year>2019</year>
          ). [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dmitrieva</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          <article-title>: “The art of legal writing”: A quantitative analysis of Russian Constitutional Court rulings</article-title>
          .
          <source>Comparative Constitutional Review</source>
          ,
          <volume>118</volume>
          (
          <issue>3</issue>
          ),
          <fpage>125</fpage>
          -
          <lpage>133</lpage>
          (
          <year>2017</year>
          ). [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Oborneva</surname>
            ,
            <given-names>I.V.</given-names>
          </string-name>
          :
          <article-title>Automation of text perception quality assessment</article-title>
          .
          <source>Bulletin of the Moscow City</source>
          Pedagogical University,
          <volume>2</volume>
          ,
          <fpage>221</fpage>
          -
          <lpage>233</lpage>
          (
          <year>2005</year>
          ). [in Russian].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Solnyshkina</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solovyev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Readability Formula for Russian Texts: A Modified Version</article-title>
          . In: Batyrshin,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Martínez-Villaseñor</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Ponce</given-names>
            <surname>Espinosa</surname>
          </string-name>
          , H. (
          <article-title>eds) Advances in Computational Intelligence</article-title>
          .
          <source>MICAI 2018. LNCS</source>
          , vol.
          <volume>11289</volume>
          . Springer, Cham (
          <year>2018</year>
          ). DOI:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -04497-8_
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Crossley</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skalicky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dascalu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Moving beyond classic readability formulas: new methods and new models</article-title>
          .
          <source>Journal of Research</source>
          in Reading,
          <volume>42</volume>
          (
          <issue>3-4</issue>
          ),
          <fpage>541</fpage>
          -
          <lpage>561</lpage>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Tweedie</surname>
            ,
            <given-names>F.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baayen</surname>
            ,
            <given-names>R.H.</given-names>
          </string-name>
          :
          <article-title>How Variable May a Constant Be? Measures of Lexical Richness in Perspective</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>32</volume>
          (
          <issue>5</issue>
          ),
          <fpage>323</fpage>
          -
          <lpage>352</lpage>
          (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Oakes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Corpus Linguistics and stylometry</article-title>
          . In: Lüdeling,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Kytö</surname>
          </string-name>
          , M. (eds.)
          <source>Corpus Linguistics: An International Handbook</source>
          , vol.
          <volume>2</volume>
          , Berlin, Mouton de Gruyter, pp.
          <fpage>1070</fpage>
          -
          <lpage>1091</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Reynolds</surname>
          </string-name>
          , R.J.:
          <article-title>Insights from Russian second language readability classification: complexity-dependent training requirements, and feature evaluation of multiple categories</article-title>
          .
          <source>In: Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          ,
          <fpage>289</fpage>
          -
          <lpage>300</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Begtin</surname>
            ,
            <given-names>I.V.</given-names>
          </string-name>
          :
          <article-title>Readability.io [API for readability assessment of texts in Russian]</article-title>
          , https://github.com/ivbeg/readability.io/wiki/API.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Batinic</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Birzer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zinsmeister</surname>
          </string-name>
          , H.:
          <article-title>Creating an extensible, levelled study corpus of Russian</article-title>
          .
          <source>In: Proceedings of the 13th Conference on Natural Language Processing (KONVENS</source>
          <year>2016</year>
          ), Bochum:
          <string-name>
            <surname>Ruhr-Universität Bochum</surname>
          </string-name>
          , pp.
          <fpage>38</fpage>
          -
          <lpage>43</lpage>
          [Bochumer Linguistische Arbeitsberichte 16] (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>16. LeStCor: Levelled Study Corpus of Russian, http://www.lestcor.com/calculator/.</mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Collins-Thompson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Computational assessment of text readability: a survey of current and future research</article-title>
          . In: François,
          <string-name>
            <given-names>Th.</given-names>
            ,
            <surname>Bernhard</surname>
          </string-name>
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <article-title>Recent Advances in Automatic Readability Assessment</article-title>
          and
          <string-name>
            <given-names>Text</given-names>
            <surname>Simplification</surname>
          </string-name>
          . Special issue of
          <source>International Journal of Applied Linguistics</source>
          ,
          <volume>165</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>97</fpage>
          -
          <lpage>135</lpage>
          (
          <year>2014</year>
          ).
          <source>DOI: 10.1075/itl.165.2</source>
          .01col.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Evert</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wankerl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nöth</surname>
          </string-name>
          , E.:
          <article-title>Reliable measures of syntactic and lexical complexity: The case of Iris Murdoch</article-title>
          .
          <source>In: Proceedings of the Corpus Linguistics 2017 Conference</source>
          , Birmingham,
          <string-name>
            <surname>UK</surname>
          </string-name>
          , (
          <year>2017</year>
          ), http://www.stefan-evert.de/PUB/EvertWankerlNoeth2017.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Alfter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volodina</surname>
          </string-name>
          , E.:
          <article-title>Towards Single Word Lexical Complexity Prediction</article-title>
          .
          <source>In: Proceedings of the Thirteenth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>88</lpage>
          (
          <year>2018</year>
          ). DOI:
          <volume>10</volume>
          .18653/v1/
          <fpage>W18</fpage>
          -0508.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Ivanov</surname>
            ,
            <given-names>V.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solnyshkina</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solovyev</surname>
            <given-names>V.D.</given-names>
          </string-name>
          :
          <article-title>Efficiency of text readability features in Russian academic texts</article-title>
          .
          <source>In: Komp'juternaja Lingvistika i Intellektual'nye Tehnologii</source>
          ,
          <volume>17</volume>
          , pp.
          <fpage>277</fpage>
          -
          <lpage>287</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <article-title>Information-legal system “Consultant Plus”</article-title>
          , http://www.consultant.ru/.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <article-title>The Constitutional Court of the Russian Federation</article-title>
          , http://www.ksrf.ru/.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Miner</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elder</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fast</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hill</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nisbet</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          et al.:
          <article-title>Practical text mining and statistical analysis for non-structured text data applications</article-title>
          . Elsevier/Academic Press, Amsterdam (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Snowball</surname>
          </string-name>
          <article-title>Stemming language and algorithms</article-title>
          . http://snowball.tartarus.org/ algorithms/russian/stop.txt.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Eder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kestemont</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Rybicki</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Stylometry with R: a suite of tools</article-title>
          .
          <source>DH</source>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Mason</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Parameters of Collocation: The Word in the Centre of Gravity</article-title>
          . In: Kirk,
          <string-name>
            <surname>J.M.</surname>
          </string-name>
          (ed.)
          <article-title>English language research on computerised corpora; Corpora galore</article-title>
          ,
          <volume>30</volume>
          , pp.
          <fpage>267</fpage>
          -
          <lpage>280</lpage>
          (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Chipere</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malvern</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Richards, В.:
          <article-title>Using a corpus of children's writing to test a solution to the sample size problem affecting type-token ratios</article-title>
          . In: Aston,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Bernardini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Stewart</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Corpora and Language Learners [Studies in Corpus Linguistics</source>
          <volume>17</volume>
          ], John Benjamins, Amsterdam, pp.
          <fpage>139</fpage>
          -
          <lpage>147</lpage>
          (
          <year>2004</year>
          ). DOI:
          <volume>10</volume>
          .1075/scl.17.
          <year>10chi</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Kettunen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Can</surname>
          </string-name>
          Type-Token Ratio be Used to Show
          <source>Morphological Complexity of Languages? Journal of Quantitative Linguistics</source>
          ,
          <volume>21</volume>
          ,
          <issue>3</issue>
          ,
          <fpage>223</fpage>
          -
          <lpage>245</lpage>
          (
          <year>2014</year>
          ).
          <source>DOI: 10.1080/09296174</source>
          .
          <year>2014</year>
          .
          <volume>911506</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>