<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Lexicon and Syntax: Complexity across Genres and Language Varieties</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pietro dell'Oglio</string-name>
          <email>pietrodelloglio@live.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dominique Brunato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Felice Dell'Orletta</string-name>
          <email>felice.dellorlettag@ilc.cnr.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Pisa Istituto di Linguistica Computazionale “Antonio Zampolli” (ILC-CNR) ItaliaNLP Lab -</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This paper presents first results of an ongoing work to investigate the interplay between lexical complexity and syntactic complexity with respect to nominal lexicon and how it is affected by textual genre and level of linguistic complexity within genre. A cross-genre analysis is carried out for the Italian language using multi-leveled linguistic features automatically extracted from dependency parsed corpora.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Linguistic complexity is a multifaceted notion
which has been addressed from different
perspectives. One established dichotomy distinguishes a
“global” vs a “local” perspective, where the
former considers the complexity of the language as a
whole and the latter focuses on complexity within
each sub-domains, i.e. phonology, morphology,
syntax, discourse
        <xref ref-type="bibr" rid="ref16">(Miestamo, 2008)</xref>
        . While
measuring global complexity is a very ambitious and
probably hopeless endeavor, measuring local
complexities is perceived as a more doable task
        <xref ref-type="bibr" rid="ref18">(Kortmann and Szmrecsanyi, 2012)</xref>
        . The level of
complexity within each subdomains indeed has been
formalized in terms of distinct parameters that
capture either internal properties of the language
(in the “absolute” notion of complexity) or
phenomena correlating to processing difficulties from
the language user’s viewpoint (in the “relative”
notion of complexity)
        <xref ref-type="bibr" rid="ref16">(Miestamo, 2008)</xref>
        . For
instance, complexity at lexical level has been
computed in terms of length (measured in characters or
syllables), of frequency either of the whole surface
word
        <xref ref-type="bibr" rid="ref17 ref5">(Randall and Wayne, 1988; Chiari and De
Mauro, 2014)</xref>
        or of its internal components (see
e.g. the root frequency effect
        <xref ref-type="bibr" rid="ref4">(Burani, 2006)</xref>
        ),
ambiguity and familiarity, among others. At syntactic
level, much attention has been paid on canonicity
effects due to word order variation
        <xref ref-type="bibr" rid="ref10 ref14 ref9">(Diessel, 2005;
Hawkins, 1994; Futrell et al., 2015)</xref>
        , as well as on
long-distance dependencies
        <xref ref-type="bibr" rid="ref11 ref12">(Gibson, 1998;
Gibson, 2000)</xref>
        proving their effect on a wide range
of psycholinguistic phenomena, such as the
subject/object relative clauses asymmetry or the
garden path effect in main verb/reduced–relative
ambiguities.
      </p>
      <p>An interesting question addressed by recent
corpus-driven research is how language
complexity is affected by textual genre. At syntactic level,
the study by Liu (2017) on ten genres taken from
the British National Corpus showed that
genrespecific stylistic factors have an influence on the
distribution of dependency distances and
dependency direction. Similarly for Italian, Brunato and
Dell’Orletta (2017) investigated the influence of
genre, and level of complexity within genre, on
a range of factors of syntactic complexity
automatically computed from dependency-parsed
corpora. Inspired by that work, we also intend to
analyze the effect of genre on linguistic
complexity. However, unlike the dominant local approach,
where each subdomain is typically studied in
isolation, our contribution intends to address the
interrelation between different levels, i.e. lexicon
and syntax. Specifically, we investigate the
following questions:
to what extent is lexical complexity
influenced by genre?
to what extent is lexical complexity
influenced by the level of complexity within the
same genre?
is there a correlation between lexical
complexity and syntactic complexity? Does it
vary according to genre and level of
complexity within the same genre?</p>
      <p>To answer these questions, we conducted an
indepth analysis for the Italian language based on
automatically dependency parsed corpora aimed at
assessing i) the distribution of simple and complex
nominal lexicon in different genres and different
language varieties for the same genre ii) the
syntactic role bears by “simple” and “complex” nouns
characterizing each corpus iii) the correlation
between “simple” and “complex” nouns with
features of complexity underlying the syntactic
structure in which they occur.</p>
      <p>In what follows we first describe the corpora
considered in this study. We then illustrate how
lexical and syntactic complexity have been
formalized. In Section 4 we discuss some
preliminary findings obtained from the comparative
investigation across corpora.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Corpora</title>
      <p>Four genres were considered in this study:
Journalism, Scientific prose, Educational writing and
Narrative. For each genre, we chose two corpora,
selected to be representative of a complex and of
a simple language variety for that genre. The level
of complexity was established according to the
expected target audience.</p>
      <p>The Journalistic corpora are Repubblica (Rep)
for the complex variety, and Due Parole (2Par) for
the simple one. Rep is a corpus of 232,908
tokens and it is made of all articles published
between 2000 and 2005 on the newspaper of the
same name; 2Par contains 322 articles taken from
the easy-to-read magazine Due Parole1, for a total
of about 73K tokens.</p>
      <p>The corpora representative of Scientific writing
are Scientific articles (ScientArt) for the complex
language variety, and Wikipedia articles (WikiArt)</p>
      <sec id="sec-2-1">
        <title>1www.dueparole.it</title>
        <p>for the simple one. The former is made of 84
documents (471,969 tokens) covering various topics
on scientific literature. The latter is made of 293
documents (about 205K tokens) extracted from the
Italian web portal “Ecology and Environment” of
Wikipedia.</p>
        <p>For the Educational writing corpora we relied
on two collections of school textbooks: the
‘complex’ one (EduAdu) contains 70 texts (48,103
tokens) targeting high school students, the ‘simple’
one (EduChi) a sample of 127 texts (48,036
tokens) targeting primary school students.</p>
        <p>Finally, the Narrative corpora are composed
by the original versions of Terence and Teacher
(TTorig), for the complex pole, and the
correspondent simplified versions for the simple pole.
Terence, which is named after the EU Terence
Project2, is made of 32 documents, covering short
novels for children. Teacher contains 24
documents extracted from web sites dedicated to
educational resources for teachers. All Terence and
Teacher texts have a simpler version (TTsemp),
which is the result of a manual simplification
process as described by Brunato and Dell’Orletta
(2017).</p>
        <p>
          All corpora were automatically tagged by the
part-of-speech tagger described in
          <xref ref-type="bibr" rid="ref1 ref8">(Dell’Orletta,
2009)</xref>
          and dependency parsed by the DeSR parser
described in
          <xref ref-type="bibr" rid="ref1">(Attardi et al., 2009)</xref>
          .
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Features of Linguistic Complexity</title>
      <p>
        Assessment of Lexical Complexity
For each corpus we extracted all lemmas tagged as
nouns, without considering proper nouns, and we
classified them as ‘simple’ vs ‘complex’ nouns.
Such a distinction was established according to
their frequency, which is one of the most used
parameter to assess the complexity of vocabulary
(see Section 1). Frequency was here computed
with respect to a reference corpus, i.e. ItWac
        <xref ref-type="bibr" rid="ref2">(Baroni et al., 2009)</xref>
        , which was chosen since this is
the biggest corpus available for standard Italian
thus offering a reliable resource to evaluate word
frequency on a large-scale. After ranking all nouns
for frequency, we pruned those with a frequency
value 3 and we kept the first quarter of nouns as
representative of the sample of simple nouns and
the last quarter as representative of the sample of
complex nouns for each corpus.
      </p>
      <sec id="sec-3-1">
        <title>2www.terenceproject.eu</title>
        <p>
          To investigate our main research questions, that
is how lexical complexity affects syntactic
complexity and the possible influence of genre and
language variety on this relationship, we focused
on a set of features automatically extracted from
the sentence parse tree. These features were
chosen since they are acknowledged to be
predictors of phenomena of structural complexity, as
demonstrated by their use in different scenarios,
such as the assessment of learners’ language
development or the level of text readability (e.g.
          <xref ref-type="bibr" rid="ref15 ref6 ref7">(Collins-Thompson, 2014; Cimino et al., 2013;
Dell’Orletta et al., 2014)</xref>
          ).
        </p>
        <p>For each corpus, all the considered features
were computed for all occurring nouns, for the
subset of complex nouns and for the subset of
simple nouns. Specifically, we focused on the
following ones:</p>
        <p>The linear distance (in terms of tokens)
separating the noun from its syntactic head
(HeadDistance in all following Tables)
The hierarchical distance (in terms of
dependency arcs) separating the noun from the root
of the tree (RootDistance)
The average number of children per noun
(AvgChildren)
The average number of siblings per noun
(AvgSibling)
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>To have a first insight into the effect of genre and
language variety on the interplay between lexical
and syntactic complexity, we compared the main
syntactic roles that nouns play in the sentence by
calculating the frequency of all dependency types
linking a noun to its head. This is shown in
Figure 1, which reports the percentage distribution of
typed dependency relationships linking a noun to
its syntactic head across all corpora. For each
corpus there are three columns: the first one
considers data for all nouns of each corpus without any
complexity label, the second one only data for the
simple noun subset and the last one only data for
the complex noun subset.</p>
      <p>It can be noted that the distribution of nouns
used as prepositional complements (prep) is the
higher one across all corpora although with
differences ranging from the lowest percentage (35.5%)
in the ‘easy’ version of the narrative corpus (i.e.
TTsemp) to the highest one (49.9%) in ScientArt
(i.e. the complex language variety for the
scientific writing genre). The syntactic role of
prepositional complement is especially played by
simple nouns compared to complex nouns. This is
particularly evident in ScientArt and Repubblica,
where the difference between simple and complex
nouns occurring as prepositional complements is
equal respectively to 20 and 15 percentage points.
Conversely, complex nouns are more widely used
as modifiers than simple nouns, especially in
Repubblica. The percentage of nouns occurring in
the subject and object position is less than 20% in
all corpora. Interestingly, the higher occurrence
of nominal subjects is attested in DueParole and
ChildEdu (14.1 and 16, respectively). This might
suggest that simpler language varieties,
independently from genre, make more use of explicit
subjects than implicit or pronominal ones. Besides,
the likelihood of a noun to be simple or complex
does not particularly affect the overall presence
of nominal subjects, unless for ScientArt and Rep
which both show a higher percentage of simple
nouns in the subject position.</p>
      <p>
        A deeper understanding of the relationship
between lexical and syntactic complexity was
provided by the investigation of the syntactic
features described in Section 3.2. Table 1 shows the
average value of the monitored features with
respect to all nouns (All), to the subset of complex
nouns (Comp) and to the subset of simple nouns
(Simp) extracted from all corpora. We assessed
whether the variation between these feature
values was statistically significant in a three different
comparative scenarios: i) between the two corpora
of the same genre, ii) between the complex
corpora of each different genre and ii) between the
simple corpora of each different genre. Table 2
shows linguistic features varying significantly for
all the considered comparisons according to the
Wilcoxon rank-sum test, a non parametric
statistical test for two independent samples
        <xref ref-type="bibr" rid="ref20">(Wild, 1997)</xref>
        .
      </p>
      <p>If we compare the two language varieties within
each genre, it can be seen, for instance, that nouns
are hierarchically more distant from the root in
the complex than in the simple version. Such a
variation, which is highly significant for all
genres, affects more the Journalistic genre
(DueParole: 2.969; Rep: 4.197) and, to a lesser extent,
the Educational one (EduChi: 3.408; EduAdu:
4.269). However, for the other monitored
syntactic features, the Wiki corpus appears as slightly
more difficult than its complex counterpart: it
has nouns that are less close to their head (Wiki:
2.531; ArtScient: 2.162) and have a richer
structure in terms of number of children (Wiki: 1.363;
ArtScient: 1.229). With the exception of root
distance, variations concerning other features within
the Narrative genre are not statistically significant.
This can be possibly due to the particular
composition of the two selected corpora: indeed, both
Terence and Teacher texts in their original version
were already conceived for an audience of
children and young students, and they were not greatly
modified in their simplified version.</p>
      <p>We finally assessed whether the variation of
these features was statistically significant
comparing the simple and the complex noun subset of the
same corpus (Table 3). According to this
dimension, we can observe that complex nouns have,
on average, less dependents (AvgChildren feature)
than simple ones, independently from the
internal distinction within genre; on the contrary, they
tend to occur more distant from the root,
especially in the complex variety of Scientific prose
(ArtScient Comp: 5.132; ArtScient Simp: 4.598).
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>While language complexity is a central topic in
linguistic and computational linguistics research,
it is typically addressed from a local perspective,
where each subdomain is investigated in
isola2Par vs Rep
Wiki vs ArtScient
EduChild vs EduAdu
TTsempl vs TTorig
ArtScient vs EduAdu
Rep vs ArtScient
Rep vs EduAdu
Rep vs TTorig
TTorig vs ArtScient
TTorig vs EduAdu
2Par vs EduChild
2Par vs TTsemp
2Par vs Wiki
TTsemp vs EduChild
TTsemp vs Wiki
Wiki vs EduChild
tion. In this preliminary work, we have defined a
method to study the interplay between lexical and
syntactic complexity restricted to the nominal
domain. We modeled the two notions in terms of
frequency, with respect to lexical complexity, and of
a set of parse tree features formalizing
phenomena of syntactic complexity. Our approach was
tested on corpora selected to be representative of
different genres and different levels of complexity
within each genre, in order to investigate whether
noun complexity differently affects syntactic
complexity according to the two dimensions. We
observed e.g. that nouns tend to appear closer to the
root in simple language varieties, independently
from genre, while the effect of genre and linguistic
complexity is less sharp with respect to the other
considered features.</p>
      <p>To have a deeper understanding of the observed
tendencies we are currently carrying out a more
in depth analysis focusing on fine-grained features
of syntactic complexity, such as the depth of the
nominal subtree. Further, we would like to enlarge
this approach to test other constituents of the
sentence, such as the verb.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>The work presented in this paper was partially
supported by the 2–year project (2017-2019)
PERFORMA – Personalizzazione di pERcorsi
FORMativi Avanzati, funded by Regione Toscana
(Progetti Congiunti di Alta Formazione – POR FSE
2014-2020 Asse A – Occupazione) in
collaboration with Meta srl company.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Attardi</surname>
          </string-name>
          , Felice Dell'Orletta, Maria Simi,
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Turian</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Accurate dependency parsing with a stacked multilayer perceptron</article-title>
          .
          <source>In Proceedings of EVALITA 2009 - Evaluation of NLP and Speech Tools for Italian</source>
          <year>2009</year>
          ,
          <string-name>
            <given-names>Reggio</given-names>
            <surname>Emilia</surname>
          </string-name>
          , Italy,
          <year>December 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Marco</given-names>
            <surname>Baroni</surname>
          </string-name>
          , Silvia Bernardini, Adriano Ferraresi, and
          <string-name>
            <given-names>Eros</given-names>
            <surname>Zanchetta</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The WaCky wide web: A collection of very large linguistically processed web-crawled corpora</article-title>
          .
          <source>In Language Resources and Evaluation</source>
          ,
          <volume>43</volume>
          :3, pp.
          <fpage>209</fpage>
          -
          <lpage>226</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Brunato and Felice Dell'Orletta</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>On the order of words in Italian: a study on genre vs complexity</article-title>
          .
          <source>International Conference on Dependency Linguistics (Depling</source>
          <year>2017</year>
          ),
          <fpage>18</fpage>
          -
          <lpage>20</lpage>
          September 2017, Pisa, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Burani</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Morfologia: i processi</article-title>
          . In: A. Laudanna and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Voghera (cur</article-title>
          .)
          <article-title>Il linguaggio</article-title>
          . Strutture.
          <article-title>Strutture linguistiche e processi cognitivi</article-title>
          . Bari, Laterza,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Isabella</given-names>
            <surname>Chiari and Tullio De Mauro</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The New Basic Vocabulary of Italian as a linguistic resource</article-title>
          .
          <source>Proceedings of the First Italian Conference on Computational Linguistics (CLIC-IT)</source>
          ,
          <source>Pisa 15-19</source>
          dicembre
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Cimino</surname>
          </string-name>
          , Felice Dell'Orletta,
          <string-name>
            <given-names>Giulia</given-names>
            <surname>Venturi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Linguistic Profiling based on General-purpose Features and Native Language Identification</article-title>
          .
          <source>Proceedings of Eighth Workshop on Innovative Use of NLP for Building Educational Applications</source>
          , Atlanta, Georgia, June 13, pp.
          <fpage>207</fpage>
          -
          <lpage>215</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Kevyn</given-names>
            <surname>Collins-Thompson</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Computational Assessment of text readability</article-title>
          .
          <source>Recent Advances in Automatic Readability Assessment and Text Simplification</source>
          . Special issue of
          <source>International Journal of Applied Linguistics</source>
          ,
          <volume>165</volume>
          :
          <fpage>2</fpage>
          , John Benjamins Publishing Company,
          <fpage>97</fpage>
          -
          <lpage>135</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Felice</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Ensemble system for partof-speech tagging</article-title>
          .
          <source>In Proceedings of EVALITA 2009 - Evaluation of NLP and Speech Tools for Italian</source>
          <year>2009</year>
          ,
          <string-name>
            <given-names>Reggio</given-names>
            <surname>Emilia</surname>
          </string-name>
          , Italy,
          <year>December 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Holger</given-names>
            <surname>Diessel</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Competing motivations for the ordering of main and adverbial clauses</article-title>
          .
          <source>Linguistics</source>
          ,
          <volume>43</volume>
          (
          <issue>3</issue>
          ):
          <fpage>449</fpage>
          -
          <lpage>470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Richard</given-names>
            <surname>Futrell</surname>
          </string-name>
          , Kyle Mahowald and
          <string-name>
            <given-names>Edward</given-names>
            <surname>Gibson</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Quantifying word order freedom in dependency corpora</article-title>
          .
          <source>In Proceedings of the Third International Conference on Dependency Linguistics (Depling</source>
          <year>2015</year>
          ),
          <fpage>91</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Edward</given-names>
            <surname>Gibson</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Linguistic complexity: Locality of syntactic dependencies</article-title>
          .
          <source>Cognition</source>
          ,
          <volume>68</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Edward</given-names>
            <surname>Gibson</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>The dependency Locality Theory: A distance-based theory of linguistic complexity. Image, Language and Brain</article-title>
          , In W.O.
          <string-name>
            <given-names>A.</given-names>
            <surname>Marants</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Miyashita</surname>
          </string-name>
          (Eds.), Cambridge, MA: MIT Press, pp.
          <fpage>95</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Gildea</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Temperley</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <source>Do Grammars Minimize Dependency Length? Cognitive Science</source>
          ,
          <volume>34</volume>
          (
          <issue>2</issue>
          ):
          <fpage>286310</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>John A. Hawkins</surname>
          </string-name>
          <year>1994</year>
          .
          <article-title>A performance theory of order and constituency</article-title>
          . Cambridge studies in Linguistics, Cambridge University Press,
          <volume>73</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Felice</given-names>
            <surname>Dell'Orletta</surname>
          </string-name>
          , Martjin Wieling, Andrea Cimino, Giulia Venturi, and
          <string-name>
            <given-names>Simonetta</given-names>
            <surname>Montemagni</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Assessing the Readability of Sentences: Which Corpora and Features</article-title>
          .
          <source>Proceedings of the 9th Workshop on Innovative Use of NLP for Building Educational Applications (BEA</source>
          <year>2014</year>
          ), Baltimore, Maryland, USA.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Matti</given-names>
            <surname>Miestamo</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Grammatical complexity in a crosslinguistic perspective</article-title>
          . In:
          <string-name>
            <surname>Miestamo</surname>
            <given-names>M</given-names>
          </string-name>
          , Sinnema¨ki
          <string-name>
            <given-names>K.</given-names>
            and
            <surname>Karlsson</surname>
          </string-name>
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds), Language Complexity: Typology, Contact Change, Amsterdam: Benjamins,
          <fpage>23</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Randall</given-names>
            <surname>James Ryder and Wayne H. Slater</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <article-title>The relationship between word frequency and word knowledge</article-title>
          .
          <source>The Journal of Educational Research</source>
          ,
          <volume>81</volume>
          (
          <issue>5</issue>
          ):
          <fpage>312</fpage>
          -
          <lpage>317</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Kortmann</given-names>
            <surname>Berndt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Szmrecsanyi</given-names>
            <surname>Benedikt</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <string-name>
            <given-names>Linguistic</given-names>
            <surname>Complexity. Second Language</surname>
          </string-name>
          <string-name>
            <surname>Acquisition</surname>
          </string-name>
          , Indigenization, Contact. Berlin, Boston: De Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Yaqin</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Haitao</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The effects of genre on dependency distance and dependency direction</article-title>
          .
          <source>Language Sciences</source>
          ,
          <volume>59</volume>
          ,
          <fpage>135</fpage>
          -
          <lpage>157</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>Chris Wild</source>
          <year>1997</year>
          .
          <article-title>The Wilcoxon Rank-Sum Test</article-title>
          . University of Auckland, Department of Statistics
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>