<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Computing Syntactic Parameters for Automated Text Complexity Assessment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Valery Solovyev</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marina Solnyshkina</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Ivanov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ivan Rygaev</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Innopolis University</institution>
          ,
          <addr-line>Innopolis</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute for Information Transmission Problems</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Kazan Federal University</institution>
          ,
          <addr-line>Kazan</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>62</fpage>
      <lpage>71</lpage>
      <abstract>
        <p>The article focuses on identifying, extracting and evaluating syntactic parameters in uencing the complexity of Russian academic texts. Our ultimate goal is to select a set of text features e ectively measuring text complexity and build an automatic tool able to rank Russian academic texts according to grade levels. models based on the most promising features by using machine learning methods The innovative algorithm of designing a predictive model of text complexity is based on a training text corpus and a set of previously proposed and new syntactic features (average sentence length, average number of syllables per word, the number of adjectives, average number of participial constructions, average number of coordinating chains, path number, i.e. average number of sub-trees). Our best model achieves an MSE of 1.15. Our experiments indicate that by adding the abovementioned syntactic features, namely the average number of participial constructions, average number of coordinating chains, and the average number of sub-trees, the text complexity model performance will increase substantially.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>E ective reading comprehension implies that reading materials correspond
readers' cognitive and language abilities. The idea behind the existing practice in
education is to ensure that students are exposed to the age-appropriate
materials which are neither too complicated nor too simple for a reader. The
\ageappropriateness" has been traditionally measured by the Grade level which is
viewed as \what all students need to know and be able to do at each grade level"
to progress through their education1.</p>
      <p>Grade level descriptors identify the speci c content (knowledge, skills,
abilities) and the language of a particular course which students at a particular
education stage (or grade) are exposed to2. There are also a number of English
text complexity analyzers available online for any educator selecting a text for
students34. Eliminating the gap between critically the texts and students'
abilities scholars have been developing tools to pro le texts that students would be
able to and want to read56. The existing automatic analyzers use hundreds of
parameters ranging from quantitative, i.e. word length and sentence length only,
to qualitative (levels of meaning or purpose; structure; language conventionality
and clarity; and knowledge demands) to match a particular reader and a text78.</p>
      <p>
        T.E.R.A., for instance, is an engine developed in 2012 by SoLET Lab which
analyzes ve textual components, such as narrativity, syntactic simplicity, word
concreteness, referential cohesion and deep cohesion9. T.E.R.A. also predicts the
grade level of the text using the Flesch-Kincaid Grade Level readability formula
(FKGL) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Knowing the grade level of texts in a corpus, educators select the
most suitable text for the target audience.
      </p>
      <p>The situation in Russia is di erent. Signi cant gaps have been reported
between the complexity levels of texts that students are asked to read in high
school as well as at tertiary levels and students' abilities: the books are either
too simple or too complicated for students10.</p>
      <p>
        Researchers also have evidence of students' lack of wish to read11 who in
many cases are caused by the inappropriate selection of a book by an educator.
Unfortunately, Russian text complexity analyzers so far apply no other variables
but quantitative, i.e. word length and sentence length[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In the paper, we aim
at the following research question: Which syntactic text features better correlate
with the complexity of Russian academic texts?
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>
        Though the terms text complexity', text di culty' and text readability' are still
sometimes used interchangeably, the notions of the concepts began being
separated in Russian academic literature as early as in 1970-s [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. I.Lerner de ned
complexity \as a category that characterizes the range of activities necessary
to solve a cognitive task, regardless of who performs this activity"[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. While
text di culty is viewed by the researcher as \a category characterizing a
person's readiness to overcome obstacles while comprehending a reading text" [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
Accepting I.Lerners' point of view V. Tcetlin speci ed that \complexity of any
3 http://www.readabilityformulas.com/free-readability-formula-tests.php
4 https://www.wikihow.com/Determine-the-Reading-Level-of-a-Book
5 http://readability.io/
6 https://netpeak.net/ru/blog/readability/
7 https://readable.io/text/
8 http://casemed.case.edu/cpcpold/students/module4/Word Readability.pdf
9 http://129.219.222.66/Publish/tera.html
10 https://elibrary.ru/item.asp?id=20134801, http://www.miep.edu.ru/uploaded/zvezdova
oreshkin.pdf, http://www.mdk-arbat.ru/bookcard?book id=934790
11 https://www.science-education.ru/ru/article/view?id=22229,
https://cyberleninka.ru/article/v/chtenie-kak-sotsialnaya-problema,
https://trvscience.ru/2015/11/17/pochemu-ne-chitayut-shkolniki/
educational material is its objective characteristics whereas di culty is a
subjective factor of students' preparedness to overcome complexity" [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The general
approach to text readability once proposed by M. Vogel and K. Washburn was
taken as a basis in all subsequent works. It is based on the objective
characteristics of the text highly correlating with the quantitative results of a test, and
implies designing a regression equation between the success of a text
comprehension, on the one hand, and text parameters { on the other [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. There are also
a number of graphic parameters of a text in uencing its comprehension such as
fonts, indentations, spaces, colors etc., which are beyond the authors interest in
the article. The rst attempts to assess Russian texts complexity were made in
late 1970-s, about 50 years later than the corresponding English studies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In
1985 Yu. Tomina [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] suggested that lexical indicators of the linguistic di culty
of texts are the number of unfamiliar words and abstract words, while
syntactic indicators are the number of participial constructions, the number of similar
parts of a sentence, prepositional-nominal groups. Though the formulas proposed
to predict Russian texts readability in late 1980-s were based on a number of
objective variables borrowed from the similar English texts complexity formulas,
i.e. word length and sentence length. Assessments of texts readability were
initially carried out by hand, and then in 2010s it was done by means of computer
programs. They were all based on I. Oborneva's readability formula derived in
2006 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]:
      </p>
      <p>F RE = 206:836
(1:52 ASL)
(65; 14 ASW )
Here, ASL is the average sentence length, i.e., the number of words divided by the
number of sentences; ASW is the average number of syllables per word, i.e., the
number of syllables divided by the number of words in a text. The constants were
calculated based on the similar English formulas as well as the comparison of 100
parallel English/Russian literary texts and words in two academic dictionaries:
Explanatory Dictionary of the Russian Language by Ozhegov S.I. with 39174
words registered12 and English-Russian Muller's dictionary with 41977 words13.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Datasets</title>
      <p>Two collections of texts were assembled for the research. The rst collection of
7 texts from textbooks on Social Studies by L. N. Bogolubov marked \BOG"
was selected to teach the predictive model and de ne independent variables of
the text variation in the range of 5 { 11 Grade Levels. The second collection
of 7 texts from textbooks on Social Studies by A.F. Nikitin marked \NIK" also
aimed for 5 { 11 Grade Levels. Further we refer to the two collection as a Russian
Readability Corpus (RRC). Both sets of textbooks are from the \Federal List
of Textbooks Recommended by the Ministry of Education and Science of the
Russian Federation to Use in Secondary and High Schools"14.
12 https://slovarozhegova.ru/
13 https://slovar-vocab.com/english-russian/muller-vocab.html
14 http://www.fpu.edu.ru/fpu/</p>
      <p>To ensure reproducibility of results, we uploaded the corpus on a website
thus providing its availability online15. Note, however, that the published texts
contain shu ed order of sentences. The sizes of BOG and NIK collections of
texts are presented in Table 1.</p>
      <p>Tokens Sentences Words per sentence Syllables per word
Grade BOG NIK BOG NIK BOG NIK BOG NIK
5-th { 17,221 { 1,499 { 11.49 { 2,35
6-th 16,467 16,475 1,273 1,197 12.94 13.76 2.56 2,71
7-th 23,069 22,924 1,671 1,675 13.81 13.69 2.84 2,70
8-th 49,796 40,053 3,181 2,889 15.65 13.86 2.96 2,88
9-th 42,305 43,404 2,584 2,792 16.37 15.55 3.04 3,00
10-th 75,182 39,183 4,468 2,468 16.83 15.88 3.07 3,12
10-th* 98,034 { 5,798 { 16.91 { 3.05 {
11-th { 38,869 { 2,270 { 17.12 { 3,11
11-th* 100,800 { 6,004 { 16.79 { 3.19 {</p>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <sec id="sec-4-1">
        <title>Corpus preprocessing</title>
        <p>For the sake of convenience, we have preprocessed all texts from the corpus
in the same way. Common preprocessing included tokenization, splitting text
into sentences and part-of-speech tagging (using the TreeTagger for Russian16).
During the preprocessing step we excluded all extremely long sentences (longer
than 120 words) as well as too short sentences (shorter than 5 words) which we
consider outliers. Clearly, such sentences can be not outliers at all in another
domain, but in case of school textbooks on Social Studies sentences shorter than
5 words are outliers.</p>
        <p>Extremely short sentences mostly appear as names of chapters and sections
of the books or as a result of incorrect sentence splitting. We omit those
sentences, because the average sentence length is a very important feature in text
complexity assessment and hence should not be biased due to splitting errors. At
the same time sentences with ve to seven words in Russian can still be viewed
as short sentences.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Lexical level Features</title>
        <p>We have explored an extended feature set for text complexity modeling:
15 https://kpfu.ru/slozhnost-tekstov-304364.html
16 http://www.cis.uni-muenchen.de/~schmid/tools/TreeTagger/
{ frequency of content words (FREQ),
{ average words per sentence (ASL),
{ average syllables per word (ASW), and
{ features based on POS-tags:
number of nouns per sentence (NOUNS),
number of verbs per sentence (VERBS),
number of adjectives per sentence (ADJ),
number of pronouns per sentence (PRONOUNS),
number of personal pronouns per sentence (PERS. PRONOUNS),
number of negations per sentence (NEG),
number of connectives per sentence (CONN).</p>
        <p>We have tested the features for their importance in linear regression model.
The two features have shown better performance than others: FREQ and ADJ.
To assess the quality of proposed models the mean squared error (MSE) and the
coe cient of determination R2 were used. For the results of tting the parameter
values in the corpus, see Table 2.
The modern tools for processing of Russian texts are able to extract several
syntactic features of texts thus permitting to include a number of syntactic
structures as features in readability formulas. To this end we use a syntactic analyzer
\ETAP-3" that employs a very detailed description of Russian grammar.</p>
        <p>
          The parser of the multipurpose linguistic processor ETAP-3 is a program
that performs parsing. In linguistic terms, parsing results in a dependency tree
structure, the nodes of which are word tokens of the input sentence, and `the
edges' are the established syntactic dependencies. Thus, every tree node
corresponds to a word token in the sentence processed, whilst the directed arcs are
labeled with names of syntactic relations [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. All the syntactic dependencies have
directions, therefore all the dependencies have the original node, its host, and
the nal one - the dependent node. Each word token is represented in the form
of the initial form of a word and a set of its morphological characteristics [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. All
the texts in the collection were processed with the syntactic parser: each
sentence was converted into a dependency tree structure. Then the following which
the following 14 numeric features were extracted:
{ An `average path' is the quotient of the number of nodes and the number
of \leaves" in a sentence
{ An `average sochin length' as the average length of coordinating
constructions is the number of nodes in \branches" starting with coordinating
constructions divided by the number of such branches; all types of nodes are
processed including conjunctions and modi ers
{ The `deeprich rate' as the average number of verbal participles (verbal
adverb phrases) is the quotient of the number of verbal participles and the
number of sentences. The verbal participles are de ned as a verbal adverb
with at least one dependent modi er.
{ The `deeprich v' as the average span of a verbal adverb phrase is the
number of verbal adverb dependent nodes (in all "branches") divided by the
number of verbal adverb phrases.
{ The leaves number' or the average number of "leaves" (terminal nodes,
i.e., words that are not anyone's "hosts") in a sentence. Calculation formula:
the number of all "leaves" in the text is divided by the number of sentences.
{ The longest path as the average average length of the longest branch is
the sum of the lengths of the longest branches of the sentence divided by the
number of sentences.
{ The nouns dep' as the average number of modi ers in a nominal group,
i.e. the sum of the all the nodes that depend on the nouns divided by the
number of nouns; coordinating and explanatory links were ignored.
{ The podchin number', i.e. the ratio of sentences in which there is at least
one syndetic with subordinate conjunctions or relational links calculated as
the number of sentences with at least one link of the type divided by the
number of sentences.
{ The podchin rate', i.e. the average number of subordinate links calculated
as the number of syndetic with subordinate conjunctions and relational links
divided by the number of sentences
{ The prich rate' as the average number of participial construction is
calculated as the number of participial constructions divided by the number of
sentences; participial constructions are de ned as a participle that has at
least one dependent.
{ The prich v' as the average span of a participial construction is the quotient
of the number of nodes that depend on the participle (in all \branches") and
the number of participial constructions.
{ The sentsoch number' as the average number of compound sentences is
the quotient of the number of coordinating constructions and the number of
sentences.
{ The sochin number' is de ned as the average number of coordinating
chains and calculated by dividing the average number of coordinating chains
in the sentence by the number of sentences; a chain is a sequence of nodes
connected by the \coordinating" links, thus conjunctions \break" the chain.
{ The path number' is de ned as the average number of sub-trees (in a
sentence), calculated with an external algorithm.
{ The verbs dep' is de ned as the average number of nite dependent verbs
and is calculated as the sum of nodes directly dependent on the nite verb
divided by the number of nite verbs; coordinating and explanatory links
were ignored.
        </p>
        <p>The list of features extracted by the ETAP syntax parser was preprocessed
to group similar features. We provide the results of correlation analysis in the
Table 3. In general, some syntactic feature are similar to others and correlate
with the target variable (readability, measured as a class number). However, it
is evident that all the syntactic features have lower correlation coe cient with
the target feature (`Grade Level'), than the two `classical' lexical features (ASL
and ASW) do.</p>
        <p>Nevertheless, syntactic features have high correlation with the target variable.
This information could be useful for readability and text complexity prediction
in Russian. Our goal is to evaluate syntactic features in the next subsection with
respect to the capability to serve as predictors in a linear regression model. As
shown in Table 3 few of the syntactic features have low correlation with the
target variable (`Grade Level'). In the rest of the paper we consider only those
syntactic features with correlation coe cient above 0:8 (as depicted in Table 3).
4.4</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluation of syntactic features</title>
        <p>For evaluation of the syntactic features we carried out the two experiments. In
the rst experiment, we clustered the syntactic features with respect to their
similarity to each other. By similarity we treat the correlation of the features'
values derived in RRC. The derived groups of features are the following:
{ Group 1, (G-1): Features related to the structure of the syntax tree
leaves number (LN)
average path (AP)
longest path (LP)
path number (PN)
{ Group 2, (G-2): Features related to noun and participial linear sequences
prich rate (PR)
nouns dep (ND)
{ Group 3, (G-3): Features related to coordinating constructions
average sochin length (SL)
sochin number (SN)</p>
        <p>The selected syntactic features could serve better predictors in a linear
regression model due to their high correlation with the target variable. We measured
the performance of the resulting model in the following way. A linear regression
model was trained on the `BOG' collection and tested on the `NIK' collection
and vice versa. In both cases, we use the MSE as a measure of model's
performance 4. First, we evaluate three features with the highest correlation: SN, PR
and ND. Next, three rows of the Table 4 correspond to the groups of syntactic
we found. Finally, we evaluate three features, one coming from a certain group.
From each group, we pick a feature with the highest correlation: PN from G-1,
PR from G-2 and SN from G-3. It can be seen from the Table 4 that syntactic
features from groups G-2 and G-3 are better than those from G-1. In the second
experiment, we compared syntactic features to the lexical level features (ASW,
ASL, FREQ and ADJ) with respect to their performance. Finally, we have built
linear models using combinations of both syntactic and lexical features (Table
5). Using only syntactic features without any additional features leads to worst
values of MSE. It is shown in the 5th row of Table 5. However, combination of
lexical level features (ASL and ASW) with syntactic features improves
performance of linear model for readability assessment.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion and Future work</title>
      <p>
        Earlier, in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], authors considered text complexity formulas based only on
lexical parameters (listed in Section 4.2). In [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], machine learning methods were
applied to the same corpus of texts and the same set of parameters. The obtained
results are superior to previously known, but using only lexical parameters in
the modeling of text complexity seems unnecessarily restrictive. Intuitively,
syntactically complex sentences should cause di culty in understanding the text.
Therefore, it was natural to include syntactic features into the set of studied
parameters. Also, the choice of a parser for the Russian language is natural
because the parser of the ETAP system is the most powerful of the existing ones.
For the study, based on expert assessment, 15 features were selected that were
supposed to have an impact on the text complexity.
      </p>
      <p>
        The main challenge to cope with is the selection of optimal values for the
constants in the formulas. We show that syntactic features could be useful in
readability assessment. However, the in uence of syntactic features was not as
great as one might expect. This indicates the future direction of research, related
to the expansion of a set of syntactic properties. In future research, we plan to
apply semantic features, such as features based on syntactic n-grams [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ] and
other types of information extracted from text [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>The research indicates that syntactic features have a high correlation with the
target variable. The ndings received could be useful for readability and text
complexity prediction of Russian texts. Since this is a preliminary study, there
is ample space to improve all major stages in the future process of syntactic
feature extraction. The potential working direction is viewed by the authors as
improving the accuracy of sentence and clause boundary detectors and parsers. We
will further experiment with syntactic complexity measures to balance the text
complexity construct multiplicity and model's simplicity. Furthermore, we can
also experiment with additional types of machine learning models and tune
parameters to derive text complexity assessment models with better performance.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>This research was nancially supported by the Russian Science Foundation, grant
# 18-18-00436, the Russian Government Program of Competitive Growth of
Kazan Federal University, state assignment of Ministry of Education and Science,
grant agreement # 34.5517.2017/6.7. The Russian Academic Corpus (section 3
in the paper) was created without support from the Russian Science Foundation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>I.</given-names>
            <surname>Boguslavsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Iomdin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lazursky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mityushin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sizov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kreydlin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Berdichevsky</surname>
          </string-name>
          .
          <article-title>Interactive resolution of intrinsic and translational ambiguity in a machine translation system</article-title>
          . In Alexander Gelbukh, editor,
          <source>Computational Linguistics and Intelligent Text Processing</source>
          , pages
          <volume>388</volume>
          {
          <fpage>399</fpage>
          , Berlin, Heidelberg,
          <year>2005</year>
          . Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>W. H. DuBay.</surname>
          </string-name>
          <article-title>The principles of readability</article-title>
          .
          <source>Online Submission</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>J.</given-names>
            <surname>Kincaid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fishburne</surname>
          </string-name>
          <string-name>
            <surname>Jr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rogers</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Chissom</surname>
          </string-name>
          .
          <article-title>Derivation of new readability formulas (automated readability index, fog count and esch reading ease formula) for navy enlisted personnel</article-title>
          .
          <source>Technical report, Naval Technical Training Command Millington TN Research Branch</source>
          ,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>I</given-names>
            <surname>Ya</surname>
          </string-name>
          <article-title>Lerner</article-title>
          .
          <article-title>Kriterii slozhnosti nekotorykh elementov uchebnika: Problemy shkolnogo uchebnika [the criteria for the complexity of some elements of the textbook: Problems of a school textbook</article-title>
          ],
          <year>1974</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>I. Obobroneva.</surname>
          </string-name>
          <article-title>Avtomatizirovannaya otsenka slozhnosti uchebnykh tekstov na osnove statisticheskikh parametrov [semiautomatic evaluation of the complexity of academic texts on the base of statistic parameters]. moscow: Rae institute of content and methods of teaching. M.: RAS Institut soderzhaniya i metodov obucheniya</article-title>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shpakovskiy</surname>
          </string-name>
          et al.
          <article-title>Otsenka trudnosti vospriyatiya i optimizatsiya slozhnosti uchebnogo teksta [Evaluation of the di culty of perception and optimization of the text complexity]</article-title>
          .
          <source>PhD thesis</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          .
          <article-title>Non-linear construction of n-grams in computational linguistics</article-title>
          . Mexico: Sociedad Mexicana de Inteligencia Arti cial,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>G.</given-names>
            <surname>Sidorov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Velasquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gelbukh</surname>
          </string-name>
          , and L.
          <string-name>
            <surname>Chanona-Hernandez</surname>
          </string-name>
          .
          <article-title>Syntactic dependency-based n-grams as classi cation features</article-title>
          .
          <source>In Mexican International Conference on Arti cial Intelligence</source>
          , pages
          <fpage>1</fpage>
          <lpage>{</lpage>
          11. Springer,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>M</given-names>
            <surname>Solnyshkina</surname>
          </string-name>
          and
          <string-name>
            <given-names>A</given-names>
            <surname>Kiselnikov</surname>
          </string-name>
          .
          <article-title>Slozhnostteksta: Etapy izucheniya v otechestvennom prikladnom yazykoznanii [text complexity: Study phases in russian linguistics]. Vestnik Tomskogo gosudarstvennogo universiteta</article-title>
          . Filologiya [Tomsk State University Journal of Philology],
          <volume>6</volume>
          :
          <fpage>38</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>V.</given-names>
            <surname>Solovyev</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          .
          <article-title>Knowledge-driven event extraction in Russian: corpusbased linguistic resources</article-title>
          .
          <source>Computational intelligence and neuroscience</source>
          ,
          <year>2016</year>
          :
          <volume>16</volume>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Valery</surname>
            <given-names>Solovyev</given-names>
          </string-name>
          , Vladimir Ivanov, and
          <string-name>
            <given-names>Marina</given-names>
            <surname>Solnyshkina</surname>
          </string-name>
          .
          <article-title>Assessment of reading di culty levels in russian academic texts: Approaches and metrics</article-title>
          .
          <source>Journal of Intelligent &amp; Fuzzy Systems</source>
          ,
          <volume>34</volume>
          (
          <issue>5</issue>
          ):
          <volume>3049</volume>
          {
          <fpage>3058</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Valery</surname>
            <given-names>Solovyev</given-names>
          </string-name>
          , Marina Solnyshkina, and
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Ivanov</surname>
          </string-name>
          .
          <article-title>Prediction of reading di culty in russian academic texts</article-title>
          .
          <source>Journal of Intelligent &amp; Fuzzy Systems</source>
          ,
          <volume>36</volume>
          :
          <fpage>4553</fpage>
          {
          <fpage>4563</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Yu</surname>
            <given-names>A</given-names>
          </string-name>
          Tomina.
          <article-title>Obektivnaja otcenka jazykovoj trudnosti tekstov (opisanie, povestvovanie</article-title>
          , rassuzhdenie, dokazatelstvo).
          <source>(in russian)</source>
          .
          <year>1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>VS</given-names>
            <surname>Tsetlin</surname>
          </string-name>
          .
          <article-title>Didactic requirements: the case of academic texts (didakticheskie trebovaniya k kriteriyam slozhnosti uchebnogo materiala)</article-title>
          .
          <article-title>Novye issledovaniya v pedagogicheskih naukah</article-title>
          .{M.: Pedagogika, (
          <volume>1</volume>
          ):
          <volume>30</volume>
          {
          <fpage>33</fpage>
          ,
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>