<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Characterizing Text Complexity with Core Vocabulary Distributional Patterns: Corpus-based Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marina Solnyshkina</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Ivanov</string-name>
          <email>v.ivanov@innopolis.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valery Solovyev</string-name>
          <email>maki.solovyev@mail.ru</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Innopolis University</institution>
          ,
          <addr-line>1, Universitetskaya st., Innopolis</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kazan Federal University</institution>
          ,
          <addr-line>18, Kremlyovskaya st., Kazan</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we report a corpus study aimed at testing the hypothesis that bigram distributional information is related to text complexity. We explored a corpus of Russian textbooks on Social studies for middle and high school to examine how the number of bigrams of the core vocabulary correlates with reading levels of texts within the grade range 5 { 11. The corpus contains 45380 sentences from 14 textbooks, written by two independent groups of authors. Each word in the corpus has a part-of-speech tag derived by TreeTagger. Due to the nature of the domain, we focus our study on a single, but high-frequency pattern: `chelovek' (a man) + verb. The ndings are particularly relevant for text complexity theory as they are consistent with the previous results of corpus investigations on a correlation of text complexity with a number of text features.</p>
      </abstract>
      <kwd-group>
        <kwd>a corpus</kwd>
        <kwd>bigram</kwd>
        <kwd>text complexity</kwd>
        <kwd>distributional patterns</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The research presented in this article is a part of Russian Academic Text
Complexity (RATC) project, which aims at de ning roles of di erent metrics in
academic text complexity analysis and has been carried out at Kazan Federal
University for over a year. The ultimate goal of the project is to provide
cognitive and linguistic pro les of Russian academic texts for middle and high schools
based on the complex linguistic analyses of the latter, describe and determine
correlations between text complexity (or grade levels) and academic text
features [
        <xref ref-type="bibr" rid="ref18 ref9">18,9</xref>
        ]. This current study is aimed at exploring two research questions: 1.
To what extent the number of collocates (and ngrams) of a particular noun may
constitute a valid and reliable index that can be used to objectively discriminate
between texts of di erent grade levels? 2. How does the size of a semantic class
of verbs correlate with text complexity across the grade levels 5 { 11?
      </p>
      <p>
        Related works: Lexical Features in Text Complexity
Studies
In modern text complexity studies, lexico-semantic features of reading texts are
considered valuable and essential metrics in assessing text complexity (see [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]).
The research shows that word frequency as a text feature impacts accuracy of
perception [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], word identi cation ability of readers [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and readers' speed of
performance in language tasks [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        At present, lexical metrics are used in over one hundred readability
formulas including those of Spache [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and Dale [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] (see [
        <xref ref-type="bibr" rid="ref10 ref7">10,7</xref>
        ]). In 1969, W.B. Elley
suggested using the term and feature of \mean noun frequency level" to de ne
readability levels of texts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Another reliable feature discriminating text
readability and its grade pro le is found to be text lexical diversity (LD) or variation
de ned as `the range and variety of vocabulary deployed in a text by either a
speaker or a writer' [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The type-token ratio (TTR), i.e. the number of word
types (or di erent words) divided by the number of tokens.
      </p>
      <p>
        Later on, TTR was acknowledged to be sensitive to the text length, and
numerous revised indices such as Root TTR and Corrected TTR, which take the
logarithm and square root of the text length instead of the direct word count as
denominator were suggested and proved to produce better results [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Experts
in Text complexity also admit that "a phrase (n-gram) gives more information
than just a single word" and better represents a text than just a word [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] as it
provides better document representation than simple \Bag of Words". Semantic
classes of words sharing a number of meaning components are viewed in Natural
Language Processing as classes not only useful for predicting certain correlations
between syntax and semantics Another text feature, i.e. lexical tightness, which
is proved to strongly correlate with grade level in \a collection of expertly rated
reading materials" (p.29) is viewed by M.Flor et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] as a metric representing
the degree to which a text tends to use words that are highly inter-associated,
i.e. semantically connected in the language. All these provide a foundation for
the hypothesis that the number of bigrams (and semantic connections) of a
highfrequency word in a text has a tendency to grow alongside with text complexity
thus representing text complexity. In present study we introduce text corpus of
academic texts and use it to investigate the abovementioned hypothesis.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Corpus Description</title>
      <p>For the purpose of this study, Russian Academic Corpus (RAC) was compiled
of two batteries of textbooks for Russian students: edited by Bogolyubov and
by Nikitin. In the Russian Federation the course on Social Studies is taught for
7 years: it starts in Grade 5 when children are typically aged 11 and nishes in
Grade 11 where the predominant majority of students is 17 years old. The course
nishes with a matriculation exam in the 11th Grade when a certain number of
students select the subject for a high-stake exam to continue their education at
universities.</p>
      <p>To ensure reproducibility of results, we uploaded the corpus on a website
thus providing its availability online3. Note, however, that the published texts
contain shu ed order of sentences. The sizes of BOG and NIK collections of
texts are presented in Table 1.</p>
      <p>Tokens Sentences ASL ASW
Grade BOG NIK BOG NIK BOG NIK BOG NIK
5-th { 17,221 { 1,499 { 11.49 { 2,35
6-th 16,467 16,475 1,273 1,197 12.94 13.76 2.56 2.71
7-th 23,069 22,924 1,671 1,675 13.81 13.69 2.84 2.70
8-th 49,796 40,053 3,181 2,889 15.65 13.86 2.96 2.88
9-th 42,305 43,404 2,584 2,792 16.37 15.55 3.04 3.00
10-th 75,182 39,183 4,468 2,468 16.83 15.88 3.07 3.12
10-th* 98,034 { 5,798 { 16.91 { 3.05 {
11-th { 38,869 { 2,270 { 17.12 { 3.11
11-th* 100,800 { 6,004 { 16.79 { 3.19 {
3.1</p>
      <sec id="sec-2-1">
        <title>Preprocessing of the corpus</title>
        <p>For the convenience, we have preprocessed all texts from the corpus in the same
way. Common preprocessing included tokenization and splitting text into
sentences. During the preprocessing step we excluded all extremely long sentences
(longer than 120 words4) as well as too short sentences (shorter than 5 words)
which we consider outliers. Clearly, such sentences can be not outliers at all
in another domain, but for the case of school textbooks on Social Studies
sentences shorter than 5 words are outliers. Sentence and word-level properties of
the preprocessed dataset are presented in Table 1.</p>
        <p>Extremely short sentences mostly appear as names of chapters and sections
of the books or as a result of incorrect sentence splitting. We omit those
sentences, because the average sentence length is a very important feature in text
complexity assessment and hence should not be biased due to splitting errors. At
the same time sentences with ve to seven words in Russian can still be viewed
as short sentences, because the average sentence length (in our corpus) is higher
than ten.</p>
        <p>The last two columns in Table 1 present well-known features that have been
widely exploited for assessment of readability of English texts. Average sentence
length (or average words per sentence, ASL) and average syllables per word5,
3 http://kpfu.ru/portal/docs/F1554781210/shu ed.zip
4 Indeed, very long sentences appear as long citations from o cial documents which
styles are completely di er from school textbooks
5 Number of syllables in a Russian word can be computed as a number of vowels in
the word.</p>
        <p>
          ASW, are the parameters in Flesch and Flesch-Kincaid formulas [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Table 1
demonstrates that values of ASL and ASW, as it is generally expected, increase
with the grades.
        </p>
        <p>All annotations in the corpus are performed on three levels: text-level,
sentencelevel and word-level. At the text-level meta-annotations refer to a number of
sentences and a set of tokens6, an author and a grade-level of a given text.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Selection of pattern for analysis</title>
        <p>
          At the word-level we have part-of-speech tag for each word. POS-tagging has
been performed with the use of the TreeTagger for Russian7. The tagset is
available from the website of the project. As reported in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] accuracy of 96% on POS
tags and of 92% on the whole tagset was achieved by TreeTagger for Russian.
Kuzmenko [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] reported POS tagging accuracy of TreeTagger between 88% and
95% depending on dataset.
        </p>
        <p>
          The contrastive analysis proved high frequency of content words, mostly
nouns and adjectives thus con rming two texts characteristics: (1) all the texts
studied are quali ed as informative; (2) their narrativity level is much lower
than that of ction texts [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The research also identi ed the word `chelovek'(a
man) to be the most frequent noun in the corpus. The word (in di erent forms)
occurs 7447 times which is around 3.1% of all nouns' mentions. In the list of
most frequent nouns it is followed by nouns such as `law' (`pravo'), `society'
(`obschestvo'), `life' (`zhizn').
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Analysis of collocations in corpus</title>
      <sec id="sec-3-1">
        <title>Collocations as a Feature of Text</title>
        <p>
          Corpus studies reveal patterns of word use in natural languages [Hanks 2004].
These patterns can be analyzed and applied in text complexity research as a
metric which shows semantic distinctions of words in contexts of di erent
complexity. We computed the Corpus with the aim to retrieve the list of verbs with
which the word `chelovek' (a man) collocates and nd out how the semantic range
and variety of these verbs extend across the textbooks of grade levels 5 { 11.
The semantic range of verbs was analysed based on based on "Russian semantic
dictionary" [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] in which the author provides a taxonomy of 35000 verbs and
according to their meanings classi es them into 3 broad and numerous ne-grained
subclasses. Shvedova's taxonomy as well as all traditional semantic classi cations
(see [
          <xref ref-type="bibr" rid="ref1 ref12 ref21">12,21,1</xref>
          ]) goes back to the idea of semantic elds of Trier [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and de nes
three main classes of verbs:
{ functional are the verbs with weakened and/or incomplete content meaning:
link-verbs and semi-notional verbs joining subjects and predicatives, verbs
6 Tokens include words, numbers, punctuation, etc.
7 http://www.cis.uni-muenchen.de/ schmid/tools/TreeTagger/
denoting the Beginning, Middle and End of an Event, modal verbs, verbs
denoting connections, relations and naming, deictic verbs;
{ `existance' (entity) verbs are the verbs denoting self-evident and directly
perceived existence of a referent;
{ `event' verbs are the verbs denoting active actions, activity, states of activity.
        </p>
        <p>The third class of verbs is a widely versatile system, represented in the
following sets: a) verbs denoting mental and emotional activity, as well as the activity
of thought and spirit; b) verbs referring to actions related to indivisible spiritual
and physical sphere: naming actions - behaviors and contacts, information, as
well as actions denoting work, various physical actions and movements. Each of
these sets combines cross-branched sections. This class also includes verbs
denoting inactive procedural states { physical and physiological. In Tables 2 and 3 we
present the classes of verbs which co-occur in all the grade-level texts and, thus,
form high-frequency bigrams with the word `chelovek' (a man) in the Corpus.</p>
        <p>As we see, the core, i.e. words that most frequently collocate with 'chelovek
(a man) is made by the semantic classes of functional verbs (Russian byt' (to
be), stat' (become)), modal verbs (moch' (can, be able to)), verbs of possession
(imet' (to have, possess). The verbs of other semantic classes are less frequent,
though there are two more verbs characterized with above average frequency,
i.e. Russian zhit'(to live) and stremit'sya (to seek). Semantic classes of verbs
that collocate with the word `chelovek' (a man) in texts of Grade 5 textbook are
presented below. The most lexically diversi ed in texts of grade 5 is the class
of verbs denoting di erent types of actions. They are represented by the core
vocabulary of the Russian language, the words which posses a high frequency in
the National language: live, make, develop, sleep, produce, try, turn, socialize,
communicate, answer, etc.</p>
        <p>The lexico-grammatical constructions or bigrams of verbs co-occurring with
the word `chelovek' (a man) functioning as a subject in texts of the 11th grade
demonstrated a wider range of the verbs used. The twenty highest ranking verbs
are the following: can (25), be (15), be (13), have (12), live (5), become (5), strive
(4), understand (4), de ne (4), speak (3), engage in (3), acquire (3), manifest (3),
follow (3), anticipate (3), call (3), see (3), create (3), join (3). The contrastive
analysis of the bigrams retrieved from the subcorpora of texts of di erent classes
(5 { 11) demonstrate that the variety of the semantic groups of the verbs are
the same but the number of the verbs with which the word 'chelovek' (a man)
collocates in each group is much higher thus the groups are 'densely populated'.
E.g. the groups of verbs denoting existence and status increase dramatically
acquiring a wider range of verbs. As it is seen in Fig. 4.1 the number of functional
verbs increase over the grade level line from 0.28 in Grade 5 to 0.37 in Grade 12
(which in the Graph marks textbooks of the 11th Grade of the Advanced level).
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In this paper we report a corpus study aimed at testing the hypothesis that
bigram distributional information could be a function of text complexity. In a
623782-word corpus of Russian textbooks on Social studies for middle and high
school we explore how the number of bigrams of the core vocabulary correlate
with reading levels of texts within the grade range 5 { 11. The applied methods
and techniques are exempli ed with the bigrams (Noun + Verb) of the most
frequent content noun in the corpus, i.e. `chelovek' ( a man). As it is unanimously
accepted that (a) texts with low frequency words are more di cult to read
and (b) frequency of separate words and bigram has proven high discriminative
power among other readability metrics in many languages, in this study we have
put our focus on two features providing rich text representation and readability
prediction:
{ the number of bigrams `chelovek' (a man) + VERB and
{ and the semantic range of the verbs in the bigrams `chelovek' (a man) +</p>
      <p>VERB.</p>
      <p>The ndings reveal that the distributional patterns of the identi ed core bigrams
construct particular semantic classes which tend to increase from grade to grade.
The research results may contribute in the following areas of professional
knowledge:
1. Textbook writers on Social sciences may be supplied with a better
understanding of the prevalence and type of verbs to be used to generate texts of
a certain reading pro le.
2. Researchers may be better able to identify type markers and rank reading
texts.
3. Identifying the ratio of functional, event (action) and entity (existance) verbs
as a marker of text type, genre and complexity can be bene cial for the
development of better reading formulas.</p>
      <p>The ndings are particularly relevant for text complexity theory as they are
consistent with the previous results of corpus investigations on correlation of
text complexity and a number of text features. The results of the research may
also have major implications for natural language processing in text complexity
research. The issue which we decided not to address in the present work is
granularity of verb classes. It is obvious that the `appropriate' level of class and
subclass granularity may vary from one research to another. In the present study
we provided a general purpose classi cation suitable for various purposes, and
in the future we intend to re ne and organize semantic classes of verbs into
taxonomies of higher degrees of granularity.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This research was nancially supported by the Russian Science Foundation,
grant # 18-18-00436, the Russian Government Program of Competitive Growth
of Kazan Federal University, and the subsidy for the state assignment in the
sphere of scienti c activity, grant agreement # 34.5517.2017/6.7. The Russian
Academic Corpus (section 3 up to subsection 3.2 in the paper) was created
without supporting by the Russian Science Foundation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Babenko</surname>
          </string-name>
          .
          <source>Explanatory Ideographical Dictionary of Russian Verbs</source>
          . Moscow, Ast-Press,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A</given-names>
            <surname>Bhakkad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.C.</given-names>
            <surname>Dharamadhikari</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kulkarni</surname>
          </string-name>
          .
          <article-title>E cient approach to nd bigram frequency in text document using e-vsm</article-title>
          .
          <source>International Journal of Computer Applications</source>
          ,
          <volume>68</volume>
          (
          <issue>19</issue>
          ),
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>E.</given-names>
            <surname>Dale</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Chall</surname>
          </string-name>
          .
          <article-title>A formula for predicting readability: Instructions</article-title>
          . Educational research bulletin, pages
          <volume>37</volume>
          {
          <fpage>54</fpage>
          ,
          <year>1948</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>O.V.</given-names>
            <surname>Dereza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.A.</given-names>
            <surname>Kayutenko</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.S.</given-names>
            <surname>Fenogenova</surname>
          </string-name>
          .
          <article-title>Automatic morphological analysis for Russian: A comparative study</article-title>
          .
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Elley</surname>
          </string-name>
          .
          <article-title>The assessment of readability by noun frequency counts</article-title>
          . Reading research quarterly, pages
          <volume>411</volume>
          {
          <fpage>427</fpage>
          ,
          <year>1969</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Flor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.B.</given-names>
            <surname>Klebanov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. M.</given-names>
            <surname>Sheehan</surname>
          </string-name>
          .
          <article-title>Lexical tightness and text complexity</article-title>
          .
          <source>In Proceedings of the Workshop on Natural Language Processing for Improving Textual Accessibility</source>
          , pages
          <volume>29</volume>
          {
          <fpage>38</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>E</given-names>
            <surname>Fry</surname>
          </string-name>
          .
          <article-title>Readability: Insights, sidelights, and hindsight</article-title>
          .
          <source>JV Ho man &amp; YM Goodman (Red.)</source>
          ,
          <article-title>Changing literacies for changing times: an historical perspective on the ture of reading research, public policy, and classroom practices</article-title>
          , pages
          <volume>174</volume>
          {
          <fpage>185</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Gibson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Osser</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Hammond</surname>
          </string-name>
          .
          <article-title>The role of grapheme-phoneme correspondence in the perception of words</article-title>
          .
          <source>The American Journal of Psychology</source>
          ,
          <volume>75</volume>
          (
          <issue>4</issue>
          ):
          <volume>554</volume>
          {
          <fpage>570</fpage>
          ,
          <year>1962</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>V.V.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.I.</given-names>
            <surname>Solnyshkina</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.D.</given-names>
            <surname>Solovyev</surname>
          </string-name>
          .
          <article-title>E ciency of text readability features in Russian academic texts</article-title>
          .
          <source>In Computational Linguistics and Intellectual Technologies</source>
          , volume
          <volume>17</volume>
          , pages
          <fpage>277</fpage>
          {
          <fpage>287</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>G. R.</given-names>
            <surname>Klare</surname>
          </string-name>
          et al.
          <source>Measurement of readability</source>
          .
          <year>1963</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. E. Kuzmenko.
          <article-title>Morphological analysis for Russian: integration and comparison of taggers</article-title>
          .
          <source>In International Conference on Analysis of Images, Social Networks and Texts</source>
          , pages
          <volume>162</volume>
          {
          <fpage>171</fpage>
          . Springer,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>B.</given-names>
            <surname>Levin</surname>
          </string-name>
          .
          <article-title>English verb classes and alternations: A preliminary investigation</article-title>
          . University of Chicago press,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>J. M. Mason</surname>
          </string-name>
          .
          <article-title>The roles of orthographic, phonological, and word frequency variables on word-nonword decisions</article-title>
          .
          <source>American Educational Research Journal</source>
          ,
          <volume>13</volume>
          (
          <issue>3</issue>
          ):
          <volume>199</volume>
          {
          <fpage>206</fpage>
          ,
          <year>1976</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Philip M McCarthy</surname>
            and
            <given-names>Scott</given-names>
          </string-name>
          <string-name>
            <surname>Jarvis</surname>
          </string-name>
          .
          <article-title>vocd: A theoretical and empirical evaluation</article-title>
          .
          <source>Language Testing</source>
          ,
          <volume>24</volume>
          (
          <issue>4</issue>
          ):
          <volume>459</volume>
          {
          <fpage>488</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>P.D. Pearson</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Studt</surname>
          </string-name>
          .
          <article-title>E ects of word frequency and contextual richness on children's word identi cation abilities</article-title>
          .
          <source>Journal of Educational Psychology</source>
          ,
          <volume>67</volume>
          (
          <issue>1</issue>
          ):
          <fpage>89</fpage>
          ,
          <year>1975</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>N.Y.</given-names>
            <surname>Shvedova</surname>
          </string-name>
          .
          <source>Russian Semantic Dictionary</source>
          , volume
          <volume>4</volume>
          . Moscow, Azbukovnik,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>M. Solnyshkina</surname>
          </string-name>
          , E. Harkova,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Kiselnikov</surname>
          </string-name>
          .
          <article-title>Comparative coh-metrix analysis of reading comprehension texts: Uni ed (Russian) state exam in English vs Cambridge rst certi cate in English</article-title>
          .
          <source>English Language Teaching</source>
          ,
          <volume>7</volume>
          (
          <issue>12</issue>
          ):
          <fpage>65</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>V.</given-names>
            <surname>Solovyev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Solnyshkina</surname>
          </string-name>
          .
          <article-title>Assessment of reading di culty levels in Russian academic texts: Approaches and metrics</article-title>
          .
          <source>Journal of Intelligent &amp; Fuzzy Systems</source>
          ,
          <volume>34</volume>
          (
          <issue>5</issue>
          ):
          <volume>3049</volume>
          {
          <fpage>3058</fpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>G.</given-names>
            <surname>Spache</surname>
          </string-name>
          .
          <article-title>A new readability formula for primary-grade reading materials</article-title>
          .
          <source>The Elementary School Journal</source>
          ,
          <volume>53</volume>
          (
          <issue>7</issue>
          ):
          <volume>410</volume>
          {
          <fpage>413</fpage>
          ,
          <year>1953</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>J.</given-names>
            <surname>Trier</surname>
          </string-name>
          .
          <article-title>Der deutsche Wortschatz im Sinnbezirk des Verstandes: von den Anfangen bis zum Beginn des 13</article-title>
          .
          <string-name>
            <surname>Jahrhunderts</surname>
          </string-name>
          , volume
          <volume>31</volume>
          .
          <string-name>
            <surname>C. Winter</surname>
          </string-name>
          ,
          <year>1931</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>A.</given-names>
            <surname>Wierzbicka</surname>
          </string-name>
          .
          <article-title>Lingua mentalis: the semantics of natural language</article-title>
          .
          <year>1980</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>M. Xia</surname>
            , E. Kochmar, and
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Briscoe</surname>
          </string-name>
          .
          <article-title>Text readability assessment for second language learners</article-title>
          .
          <source>In Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , pages
          <volume>12</volume>
          {
          <fpage>22</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>