<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An extended version of the KoKo German L1 Learner corpus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Abel</string-name>
          <email>andrea.abel@eurac.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aivars Glaznieks</string-name>
          <email>aivars.glaznieks@eurac.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lionel Nicolas</string-name>
          <email>lionel.nicolas@eurac.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Egon Stemle</string-name>
          <email>egon.stemle@eurac.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Specialised Communication and Multilingualism EURAC Research Bolzano/Bozen</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>English. This paper describes an extended version of the KoKo corpus (version KoKo4, Dec 2015), a corpus of written German L1 learner texts from three different German-speaking regions in three different countries. The KoKo corpus is richly annotated with learner language features on different linguistic levels such as errors or other linguistic characteristics that are not deficit-oriented, and is enriched with a wide range of metadata. This paper complements a previous publication (Abel et al., 2014a) and reports on new textual metadata and lexical annotations and on the methods adopted for their manual annotation and linguistic analyses. It also briefly introduces some linguistic findings that have been derived from the corpus.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Italiano. Il contributo descrive una
versione estesa del corpus KoKo
(versione KoKo4, Dic 2015), corpus che
raccoglie produzioni scritte di apprendenti di
tedesco L1, provenienti da tre distinte
regioni germanofone, a loro volta situate in
tre diversi paesi. Il corpus KoKo e`
annotato dettagliatamente su differenti livelli
linguistici rilevanti, quali gli errori o
altre caratteristiche linguistiche non
direttamente ricollegabili a deficit individuali, ed
arricchito da un’ampia gamma di
metadati. Questo contributo integra una
precedente pubblicazione
        <xref ref-type="bibr" rid="ref1 ref2">(Abel et al., 2014a)</xref>
        e`
informa sui nuovi metadati testuali e sulle
nuove annotazioni lessicali cosi come sui
metodi adottati per la loro annotazione
manuale e per le loro analisi linguistiche.
Inoltre presenta brevemente alcuni
risultati ricavati dal corpus.
      </p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        The study of linguistically annotated learner
corpora has received a growing interest over the past
20 years
        <xref ref-type="bibr" rid="ref18">(Granger et al., 2013)</xref>
        . In learner
corpus linguistics, such corpora are usually defined as
“systematic computerized collections of texts
produced by language learners”
        <xref ref-type="bibr" rid="ref27">(Nesselhauf, 2005)</xref>
        .
Unlike most learner corpora focusing on L2/FL
learners (i.e. learners learning a foreign language),
the KoKo corpus focuses on advanced L1 speakers
that are still learning their mother tongue, which
typically happens in educational contexts.
      </p>
      <p>
        This paper describes an extended version of the
KoKo corpus
        <xref ref-type="bibr" rid="ref1 ref2">(Abel et al., 2014a)</xref>
        , a corpus
created for the purposes of the KoKo project which
aims at investigating the writing skills of
Germanspeaking secondary school pupils. The creation of
the corpus was guided by two goals: on the one
hand to describe writing skills at the end of
secondary school, on the other hand to consider
external socio-linguistic factors (e.g. gender,
socioeconomic background etc.).
      </p>
      <p>
        The previous description focused on the data
collection, the data processing, the annotation of
orthographic and grammatical features as well as
on aspects regarding annotation quality
        <xref ref-type="bibr" rid="ref1 ref2">(Abel et
al., 2014a)</xref>
        . This paper, however, introduces the
new textual metadata and lexical annotations.
      </p>
      <p>The paper is structured as follows. In section 2,
key facts are briefly reported, including references
to related work. The new textual metadata and
lexical annotations are then described in section 3,
alongside with the methods adopted for their
manual annotation and linguistic analyses and some
examples of linguistic findings. In section 4,
future works are discussed right before concluding
in section 5.</p>
    </sec>
    <sec id="sec-3">
      <title>Key Information about the Corpus</title>
      <p>
        The KoKo corpus is a collection of 1,503
authentic argumentative essays, and the corresponding
survey information about their authors, produced
in classrooms under standardized conditions by
learners of 85 classes of 66 schools from three
different German-speaking areas: South Tyrol in
Italy, North Tyrol in Austria and Thuringia in
Germany.1 Such areas are particularly suitable for
comparative studies because of differences
regarding the German standard varieties, the use of
dialectal vs. standard varieties and the monolingual
vs. plurilingual environments
        <xref ref-type="bibr" rid="ref1 ref2">(Abel et al., 2014a)</xref>
        .
      </p>
      <p>
        The corpus is roughly equally distributed over
the three regions and amounts to 824,757 tokens
(punctuation excluded). All writers were attending
secondary schools one year before their
schoolleaving examinations. 83% of the pupils were
native speakers of German. The corresponding
L1 part of the corpus amounts to 726,247
tokens. Metadata annotations amount to 52,605
annotations whereas manual annotations amount to
117,422 annotations. Furthermore, 366 features
to measure linguistic complexity2
        <xref ref-type="bibr" rid="ref18 ref19 ref20">(Hancke et al.,
2012; Hancke and Meurers, 2013)</xref>
        were
automatically calculated per text (550,098 in total) and
added as metadata.
      </p>
      <p>
        Previous evaluation showed high accuracy of
manual transcriptions (&gt; 99%), and automatic
tokenization (&gt; 99%), sentence splitting (&gt; 96%)
and POS-tagging (&gt; 96%)
        <xref ref-type="bibr" rid="ref1 ref16">(Glaznieks et al.,
2014)</xref>
        .
      </p>
      <p>
        As it is among the first accessible richly
linguistically annotated German L1 learner corpora, the
KoKo corpus is particularly relevant to L1 learner
language researchers, and for the field of didactics
of German as L1. Other comparable language
resources are either not accessible
        <xref ref-type="bibr" rid="ref10 ref29 ref6">(Berg et al., 2010;
DESI-Konsortium, 2006; Nussbaumer and Sieber,
1994)</xref>
        , or although accessible, have not been
enriched with linguistic information
        <xref ref-type="bibr" rid="ref14 ref5">(Augst et al.,
2007; Fix and Melenk, 2002)</xref>
        or are only partly
1We followed the privacy policy for such surveys and
requested a signed consent from all adult participants and
parents of minors. In addition, all students participated
anonymously, no names of the students were collected, names of
schools were codified and made anonymous.
      </p>
      <p>
        2e.g. syntactic features such as the average length of NPs,
VPs and PPs as well as their number per sentence,
morphological features such as the number of modal verbs per total
number of verbs or the average compound depth of nouns,
and lexical features such as lexical diversity described by
means of different measures
annotated
        <xref ref-type="bibr" rid="ref38">(Thelen, 2010)</xref>
        . Some other corpora
include L1 data, but as reference for L2/FL learner
corpus research
        <xref ref-type="bibr" rid="ref20 ref32 ref41">(Reznicek et al., 2010;
Zinsmeister and Breckle, 2012)</xref>
        .
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>New Metadata and Annotations</title>
      <p>This section describes the main features of the
latest corpus version KoKo4 (Dec. 2015) that have
been added to the version KoKo3 (Dec. 2014). It
thus focuses on a new set of textual metadata and
a new layer of lexical annotations which is, due to
the selected features and the degree of
granularity, a novelty in (corpus-based) modeling of
L1writing competences for German .
3.1</p>
      <sec id="sec-4-1">
        <title>Textual Metadata</title>
        <p>In the KoKo corpus, two kinds of Metadata
information are available: (1) non-linguistic, i.e.
person-related information provided by each
participant via a questionnaire survey in class that is
available for the whole sample and (2)
linguistic, i.e. text-related information provided for a
subsample of the corpus (569 texts, equally
distributed over the three regions involved) through
an online evaluation form by three different
specially trained raters originating from the different
participating regions.</p>
        <p>
          While type (1) metadata allow for
sociolinguistic analyses in order to detect relations
between linguistic features (e.g. text length, sentence
length, orthographic errors, grammatical errors,
etc.) and non-linguistic person-related
information, type (2) metadata constitute a further
expansion of our analysis by including textual features
as well. Text analysis was done holistically
using an evaluation form and detailed guidelines that
were elaborated on the basis of recent findings in
writing research and text analyses
          <xref ref-type="bibr" rid="ref11 ref13 ref21 ref4 ref5 ref7 ref8">(Brinker, 2010;
Feilke, 2010; Augst et al., 2007; Bo¨ttcher and
Becker-Mrotzek, 2006; Jechle, 1992; Augst and
Faigel, 1986)</xref>
          and the curricula in the participating
regions. The text evaluation form distinguishes
four categories : (A) formal completeness, (B)
content, (C) formal and linguistic means of text
arrangement and (D) overall impression.
        </p>
        <p>For category A, 10 questions of the online
evaluation form focused on the presence of
obligatory text parts (introduction, main part, closing
part) and explicitly requested constituents of
argumentative essays (opinion of the author,
conclusion). The 25 questions of category B belong to
two subcategories: (B1) the topics of the essay (9
questions), (B2) patterns of topic development (16
questions). B1 comprises evaluations on e.g. the
topics of each text part, gaps, and the overall
coherence of the text. B2 refers to the main pattern
of topic development (argumentative, etc.), the
argumentation strategies (point of view, concessive
or not), and the motivation of arguments
(objective vs. subjective stance, quality of arguments).
Formal and linguistic means of text arrangement
(category C, 7 questions) focus on the use of
paragraphs, the explicit announcement of and
commitment to the function of the essay, and the use of
linguistic means to structure the text with regards
to content. Finally, category D (20 questions)
aims for an overall impression and therefore
focuses on the completion of the task (successful or
not), the overall quality of the text and the
overall consistency of both the quality and coherence.
Of all 62 questions of the entire online evaluation
form, we used 57 for each document of the
subcorpus (alltogether 33,972 annotations).</p>
        <p>The analyses revealed, among other things, that
the text quality is classified as quite satisfactory
on a 5 point Likert-scale3. More specifically, there
are significant correlations between text quality
assessment and other linguistic variables: thus, a
lower number of e.g. lexical errors is connected to
a higher text quality score4, and, finally, a variety
of group differences could be detected (e.g.
concerning school type: lower text quality scores
within vocational schools compared to general
high schools5).
3.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Lexical Annotations</title>
        <p>
          As for the manual annotations of orthographic
and grammatical features added to previous
corpus versions
          <xref ref-type="bibr" rid="ref1 ref2">(Abel et al., 2014a)</xref>
          , a specifically
crafted tag set and annotation manual were used
for the annotation of lexical features. 61,728
lexical annotations were manually performed by
trained annotators on a subcorpus of 980 texts,
almost equally distributed over the three regions.
        </p>
        <p>The analyses of lexical features focuses on
lexical knowledge as a central part of lexical
competence which includes the dimensions of
lexical breadth (quantitative aspect) and lexical depth
3percentages: 1 (scarse): 6.2 - 2: 22.9 - 3: 39.0 - 4: 26.3
5 (excellent): 5.7)</p>
        <p>4Kruskal Wallis H Test: FS errors X2(1) = 10.417, p =
.036, single word errors: ANOVA F(4, 338) = 2.805, p = .026
5Kruskal Wallis H Test: X2(1) = 49.147, p = .000</p>
      </sec>
      <sec id="sec-4-3">
        <title>Category</title>
        <p>Single
words</p>
        <sec id="sec-4-3-1">
          <title>Phrasemes</title>
        </sec>
        <sec id="sec-4-3-2">
          <title>Particularities</title>
        </sec>
        <sec id="sec-4-3-3">
          <title>Target hyp.</title>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>Sub-category Total</title>
        <p>
          Neol. &amp; occas. 4,670
Arg. adv. &amp; conj. 14,345
Referential 18,708
Communicative 4,824
Structural 2,704
Semantic 8,397
Stylistic 236
Form 1,923
Metalinguistic 1,412
4,509
(qualitative aspect)
          <xref ref-type="bibr" rid="ref11 ref25 ref26 ref30 ref31 ref37 ref7">(Steinhoff, 2009; Bo¨ttcher and
Becker-Mrotzek, 2006; Mukherjee, 2005; Read
and Nation, 2004; Read, 2000; Nation, 2001)</xref>
          .
Whereas the analyses of quantitative aspects of
lexical knowledge were performed automatically
by using different measures (e.g. lexical diversity
measures such as MTDL and Yule’s K, or lexical
frequency scores based on dlexDB
          <xref ref-type="bibr" rid="ref18 ref19">(Hancke and
Meurers, 2013)</xref>
          ), the analyses of qualitative
aspects were done by means of manual annotations.
We focus hereafter exclusively on the manual
annotations allowing us to model qualitative aspects
of lexical knowledge.
        </p>
        <p>
          For annotating lexical features, we developed
a new hierarchically-structured linguistic
classification scheme inspired by previous work that
focused on L2 learner languages
          <xref ref-type="bibr" rid="ref1 ref2 ref22">(Abel et al., 2014b;
Konecny et al., 2016)</xref>
          . The classification scheme
takes both into account occurrences of selected
lexical phenomena and defective as well as
nondefective particularities of learner languages
considering two dimensions: (1) the linguistic
subcategory, e.g. collocations and idioms, and (2) a
target modification classification, e.g. omission,
addition
          <xref ref-type="bibr" rid="ref1 ref11 ref2 ref7">(D´ıaz-Negrillo and Dom´ınguez, 2006; Abel
et al., 2014b)</xref>
          . Furthermore, we formulated target
hypotheses for those categories that we annotated
as defective in order to make the error
interpretation transparent
          <xref ref-type="bibr" rid="ref24">(Lu¨deling et al., 2005)</xref>
          . The
corresponding annotation scheme contains 77 different
tags including a set of further attributes.
        </p>
        <p>
          In a multi-stage annotation procedure, all
occurrences of phenomena on both single words and
formulaic sequences (FS) were annotated
          <xref ref-type="bibr" rid="ref39">(Wray,
2005)</xref>
          . Annotations for particularities were
subsequently added in order to distinguish between
errors concerning correctness, errors concerning
appropriateness of usage
          <xref ref-type="bibr" rid="ref12 ref34">(Eisenberg, 2007;
Schneider, 2013)</xref>
          , non-defective modifications (to
capture, for example, creative use of language), and
diasystematic markedness. At the single word
level, we considered all out-of-vocabulary tokens
of the part-of-speech tagger
          <xref ref-type="bibr" rid="ref33">(Schmid, 1994)</xref>
          as
candidates of neologisms or occasionalisms. In
addition, we captured a variety of tokens
relevant for the text genre of an argumentative
essay (i.e. argumentative adverbs and conjunctions).
At the level of FS, we applied a function-based
approach distinguishing between three main
categories of phrasemes
          <xref ref-type="bibr" rid="ref9">(Burger, 2007)</xref>
          , each of them
with further subcategories
          <xref ref-type="bibr" rid="ref1 ref17 ref2 ref22 ref3 ref35 ref36 ref9">(Abel et al., 2014b;
Konecny et al., 2016; Granger and Paquot, 2008;
Burger, 2007; Stein, 2007; Steinhoff, 2007)</xref>
          , as
well as a “mixed classification”
          <xref ref-type="bibr" rid="ref9">(Burger, 2007)</xref>
          :
        </p>
        <p>Referential phrasemes include collocations6
and idioms7, distinguished among other things
with respect to their degree of idiomaticity.
Communicative phrasemes are subdivided into those
bound to specific situations8, and those not
bound to specific situations9. Finally, structural
phrasemes comprise complex conjunctions and
prepositions10 and concessive constructions11.</p>
        <p>For particularities, we considered four main
categories, each with further subcategories:</p>
        <p>On a semantic dimension a distinction is
made between denotative errors concerning
correctness or appropriatness of use12, and
connotative markedness or appropriatness of use13. The
stytlistic dimension considers repetition, and
redundancy. The form dimension focusses on
6further divided into restricted and loose collocations,
light verb constructions (called ”Funktionsverbgefu¨ge” in
German) as well as special classes such as irreversible
biand trinominals, similes etc.</p>
        <p>7further divided into nominative idioms and fixed phrases,
and special classes such as irreversible bi- and trinominals
etc.</p>
        <p>8further divided into general routine or speech act
formulas, special classes such as commonplaces, slogans, proverbs
etc., and empty formulas</p>
        <p>9further divided into text organising formulas, and
interaction organising formulas</p>
        <p>10further divided into phraseological connectors and
syntactically complex connectors, and secondary prepositions
11further divided into constructions with ”although” and a
correlate of the ”but”-class, and constructions with a modal
word and a correlate of the ”but”-class</p>
        <p>12further divided into reference/function, contextual
fitness, semantic compatibility, and precision</p>
        <p>13further divided into speaker’s attitude, and
diasystematic markedness concerning language usage, i.e. diaphasic
markedness, diachronic markedness, diatopic markedness
word formation errors (concerning single word
units only14), and on omission, choice, position
and addition errors as well as creative
modifications (concerning FS). Concerning metalinguistic
markers the appropriateness of the use of
quotation marks for highlighting units is considered.</p>
        <p>An overview of the number of annotations is
provided in Table 1.</p>
        <p>Results of the analyses showed, among
others, that pupils use different types of FS quite
frequently, on average 5.12 constructions per
100 words: with 62%, non idiomatic
referential phrasemes constitute the major part,
followed by idiomatic referential phrasemes (19%),
and, finally, structural (10%) and communicative
phrasemes (9%). However, lexical errors in
general affect more often FS than single word units
(10% of the FS vs. 1.04% of the single words).
The latter are most frequently form errors (5.50%
of FS affected, especially choice errors: 4.17%).
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Future Work</title>
      <p>
        The KoKo project was completed and presented
to the public in December 2015. We will start
releasing the data via the corpus exploration
interface ANNIS3
        <xref ref-type="bibr" rid="ref22 ref23">(Krause and Zeldes, 2016)</xref>
        and
for download on request, after signing a license
agreement.15 Aside from the aforementioned data,
future versions will also include additional
metadata information about the authors integrated for
the purposes of future socio-linguistic analyses.
      </p>
      <p>
        Consensus in the annotations among annotators,
and as such an indication of its reliability, will
be evaluated on sub-sets of texts that were
annotated for this purpose by more than one annotator.
Three annotators independently annotated the text
level metadata annotations on 27 texts, and six
annotators independently annotated the lexical level
annotations on the same 27 texts. Inter-annotator
agreement will be calculated for annotations and
segmentation, i.e. the agreement on the decision
which word sequence needs to be tagged vs. what
annotation needs to be assigned to it, and will be
evaluated and reported in the form of Fleiss
Multik and boundary similarity
        <xref ref-type="bibr" rid="ref15 ref17 ref3">(Artstein and Poesio,
2008; Fournier, 2013)</xref>
        .
      </p>
      <p>Finally, thanks to its relatively large size and its
richly annotated nature, potential additional uses
14distinguishing between errors with respect to derivation
and to composition</p>
      <p>
        15We have been trying to make the data available for direct
download – but have to take more legal hurdles.
of the KoKo corpus in Natural Language
Processing and Corpus Linguistics are being considered.
Regarding Natural Language Processing, the error
annotations paired with target hypothesis
annotations allow for creating an aligned corpus. Such
corpora can be used to improve machine
translation for automatically correcting learner texts
        <xref ref-type="bibr" rid="ref28">(Ng
et al., 2014)</xref>
        . Regarding Corpus Linguistics,
machine learning methods can be used (e.g. as being
done in WebAnno
        <xref ref-type="bibr" rid="ref40">(Yimam et al., 2014)</xref>
        ) to drive
linguistic intuitions when performing annotations
or analyses. Because of the richness of its
annotation schemes, the KoKo corpus constitutes a
challenging but at the same time promising dataset to
test if the developed methods are able to uncover
relevant correlations that have already been
investigated, or to uncover even new ones that are worth
considering for future linguistic analyses.
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>This paper described the most recent version of the
KoKo corpus, a collection of richly annotated
German L1 learner texts, and focused on the new
textual metadata and lexical annotations.</p>
      <p>Because other comparable language resources
are either not accessible, or have not been enriched
with linguistic information or are only partly
annotated, the corpus is a valuable resource for
research on L1 learner language, in particular for
the research on writing skills, and for teachers of
German as L1, in particular for the teaching of L1
German writing skills.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Abel</surname>
          </string-name>
          , Aivars Glaznieks, Lionel Nicolas, and
          <string-name>
            <given-names>Egon</given-names>
            <surname>Stemle</surname>
          </string-name>
          . 2014a.
          <article-title>Koko: An L1 learner corpus for german</article-title>
          .
          <source>In Proceedings of LREC 2014</source>
          , pages
          <fpage>2414</fpage>
          -
          <lpage>2421</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Abel</surname>
          </string-name>
          , Katrin Wisniewski, Lionel Nicolas, and
          <string-name>
            <given-names>Detmar</given-names>
            <surname>Meurers</surname>
          </string-name>
          .
          <year>2014b</year>
          .
          <article-title>A trilingual learner corpus illustrating european reference levels</article-title>
          .
          <source>RICOGNIZIONI- Rivista di Lingue</source>
          , Letterature e Culture Moderne,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>111</fpage>
          -
          <lpage>126</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Ron</given-names>
            <surname>Artstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Poesio</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Inter-coder agreement for computational linguistics</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>34</volume>
          (
          <issue>4</issue>
          ):
          <fpage>555</fpage>
          -
          <lpage>596</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Augst</surname>
          </string-name>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Faigel</surname>
          </string-name>
          .
          <year>1986</year>
          .
          <article-title>Von der Reihung zur Gestaltung: Untersuchungen zur Ontogenese der schriftsprachlichen</article-title>
          <source>Fa¨higkeiten von 13-23 Jahren</source>
          , volume
          <volume>5</volume>
          of Theorie und Vermittlung der Sprache. Peter Lang, Frankfurt.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Augst</surname>
          </string-name>
          , Katrin Disselhoff, Alexandra Henrich, Thorsten Pohl, and Paul Vo¨lzing.
          <year>2007</year>
          .
          <article-title>TextSorten-Kompetenz</article-title>
          .
          <article-title>Eine echte Longitudinalstudie zur Entwicklung der Textkompetenz im Grundschulalter</article-title>
          .
          <source>Peter Lang</source>
          , Frankfurt.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Margit</given-names>
            <surname>Berg</surname>
          </string-name>
          , Anne Berkemeier, Reinold Funke, Christian Glu¨ck, Christiane Hofbauer, and Jordana Schneider, editors.
          <year>2010</year>
          . Sprachliche Heterogenita¨t in der Sprachheil- und der Regelschule. Abschlussbericht im Programm ,,Bildungsforschung” der Landesstiftung Baden-Wrttemberg, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Ingrid</given-names>
            <surname>Bo</surname>
          </string-name>
          <article-title>¨ttcher and</article-title>
          <string-name>
            <surname>Michael</surname>
          </string-name>
          Becker-Mrotzek.
          <year>2006</year>
          .
          <article-title>Schreibkompetenz entwickeln und beurteilen</article-title>
          .
          <source>Cornelsen</source>
          , Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Klaus</given-names>
            <surname>Brinker</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Linguistische Textanalyse. Eine Einfu¨hrung in Grundbegriffe und Methoden</article-title>
          .
          <source>Bearbeitet von Sandra Ausborn-Brinker</source>
          ,
          <volume>7</volume>
          ., durchgesehene Auflage, volume
          <volume>29</volume>
          of Grundlagen der Germanistik. Erich Schmidt Verlag, Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Harald</given-names>
            <surname>Burger</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Phraseologie: Eine Einfu¨hrung am Beispiel des Deutschen</article-title>
          , volume
          <volume>36</volume>
          of Grundlagen der Germanistik. Erich Schmidt Verlag, Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>DESI-Konsortium</surname>
          </string-name>
          , editor.
          <year>2006</year>
          .
          <article-title>Unterricht und Kompetenzerwerb in Deutsch und Englisch</article-title>
          . Beltz Verlag, Weinheim - Bern.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Ana D´</surname>
          </string-name>
          ıaz-Negrillo and
          <article-title>Jesu´s Ferna´ndez Dom´ınguez</article-title>
          .
          <year>2006</year>
          .
          <article-title>Error tagging systems for learner corpora</article-title>
          . Revista espan˜ola de lingu¨´ıstica aplicada, (
          <volume>19</volume>
          ):
          <fpage>83</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Peter</given-names>
            <surname>Eisenberg</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Sprachliches Wissen im Wo¨rterbuch der Zweifelsfa¨lle. U¨ ber die Rekonstruktion einer Gebrauchsnorm</article-title>
          .
          <source>Aptum. Zeitschrift fu¨r Sprachkritik und Sprachkultur</source>
          ,
          <volume>3</volume>
          (
          <year>2007</year>
          ):
          <fpage>209</fpage>
          -
          <lpage>228</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Helmuth</given-names>
            <surname>Feilke</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Schriftliches Argumentieren zwischen Na¨he und Distanz am Beispiel wissenschaftlichen Schreibens</article-title>
          .
          <source>Na¨he und Distanz im Kontext variationslinguistischer Forschung</source>
          , pages
          <fpage>209</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Fix</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hartmut</given-names>
            <surname>Melenk</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Schreiben zu Texten-Schreiben zu Bildimpulsen: das Ludwigsburger Aufsatzkorpus; mit 2300 Schu¨lertexten, Befragungsdaten und Bewertungen auf CD-ROM</article-title>
          . Schneider-Verlag,
          <year>Hohengehren</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Chris</given-names>
            <surname>Fournier</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Evaluating Text Segmentation using Boundary Edit Distance</article-title>
          .
          <source>In Proceedings of 51st Annual Meeting of the ACL</source>
          , pages
          <fpage>1702</fpage>
          -
          <lpage>1712</lpage>
          . ACL.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Aivars</given-names>
            <surname>Glaznieks</surname>
          </string-name>
          , Lionel Nicolas, Egon Stemle, Andrea Abel, and
          <string-name>
            <given-names>Verena</given-names>
            <surname>Lyding</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Establishing a standardised procedure for building learner corpora</article-title>
          .
          <source>Apples - Journal of Applied Language Studies</source>
          ,
          <volume>8</volume>
          (
          <issue>3</issue>
          ):
          <fpage>5</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Sylviane</given-names>
            <surname>Granger</surname>
          </string-name>
          and
          <string-name>
            <given-names>Magali</given-names>
            <surname>Paquot</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Disentangling the phraseological web</article-title>
          .
          <source>Phraseology. An interdisciplinary perspective</source>
          , pages
          <fpage>27</fpage>
          -
          <lpage>50</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Sylviane</given-names>
            <surname>Granger</surname>
          </string-name>
          , Gae¨tanelle Gilquin, and
          <string-name>
            <given-names>Fanny</given-names>
            <surname>Meunier</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Twenty Years of Learner Corpus Research</article-title>
          . Looking Back,
          <source>Moving Ahead: Proceedings of the First Learner Corpus Research Conference (LCR</source>
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Julia</given-names>
            <surname>Hancke</surname>
          </string-name>
          and
          <string-name>
            <given-names>Detmar</given-names>
            <surname>Meurers</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Exploring CEFR classification for German based on rich linguistic modeling</article-title>
          .
          <source>In Proceedings of the Learner Corpus Research Conference (LCR</source>
          <year>2013</year>
          ), pages
          <fpage>54</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Julia</given-names>
            <surname>Hancke</surname>
          </string-name>
          , Sowmya Vajjala, and
          <string-name>
            <given-names>Detmar</given-names>
            <surname>Meurers</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Readability classification for german using lexical, syntactic, and morphological features</article-title>
          .
          <source>In Martin Kay and Christian Boitet</source>
          , editors,
          <source>Proceedings of COLING 2012</source>
          , pages
          <fpage>1063</fpage>
          -
          <lpage>1080</lpage>
          , Mumbai.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Jechle</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Kommunikatives Schreiben: Prozess und Entwicklung aus der Sicht kognitiver Schreibforschung</article-title>
          , volume
          <volume>41</volume>
          of ScriptOralia. Gunter Narr Verlag, Tu¨bingen.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Christine</given-names>
            <surname>Konecny</surname>
          </string-name>
          , Andrea Abel, Erica Autelli, and
          <string-name>
            <given-names>Lorenzo</given-names>
            <surname>Zanasi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Identification and Classification of Phrasemes in an L2 Learner Corpus of Italian</article-title>
          . In Gloria Corpas Pastor, editor,
          <source>Computerised and Corpus-based Approaches to Phraseology</source>
          , pages
          <fpage>533</fpage>
          -
          <lpage>542</lpage>
          . Editions Tradulex, Geneva.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Krause</surname>
          </string-name>
          and
          <string-name>
            <given-names>Amir</given-names>
            <surname>Zeldes</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>ANNIS3: A New Architecture for Generic Corpus Query and Visualization</article-title>
          .
          <source>Digital Scholarship in the Humanities</source>
          ,
          <volume>31</volume>
          (
          <issue>1</issue>
          ):
          <fpage>118</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Anke</given-names>
            <surname>Lu</surname>
          </string-name>
          ¨deling, Maik Walter, Emil Kroymann, and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Adolphs</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Multi-level error annotation in learner corpora</article-title>
          .
          <source>In Proceedings of Corpus Linguistics</source>
          <year>2005</year>
          , pages
          <fpage>15</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Joybrato</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>The native speaker is alive and kicking: Linguistic and language-pedagogical perspectives</article-title>
          .
          <source>Anglistik</source>
          ,
          <volume>16</volume>
          (
          <issue>2</issue>
          ):
          <fpage>7</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>I.S.P.</given-names>
            <surname>Nation</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Learning Vocabulary in Another Language</article-title>
          .
          <source>Foreign Language Study</source>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Nadja</given-names>
            <surname>Nesselhauf</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Collocations in a Learner Corpus</article-title>
          , volume
          <volume>14</volume>
          of Studies in Corpus Linguistics. John Benjamins Publishing, Amsterdam.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Hwee</given-names>
            <surname>Tou</surname>
          </string-name>
          <string-name>
            <surname>Ng</surname>
          </string-name>
          , Siew Mei Wu, Ted Briscoe, Christian Hadiwinoto, Raymond Hendy Susanto, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Bryant</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The CoNLL-2014 Shared Task on Grammatical Error Correction</article-title>
          .
          <source>In Proceedings of the Eighteenth Conference on Computational Natural Language Learning: Shared Task</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          , Baltimore, Maryland. ACL.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Markus</given-names>
            <surname>Nussbaumer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Sieber</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Texte analysieren mit dem Zu¨rcher Textanalyseraster</article-title>
          . In Peter Sieber, editor, Sprachfa¨
          <article-title>higkeiten-Besser als ihr Ruf und no¨tiger den je!</article-title>
          , pages
          <fpage>141</fpage>
          -
          <lpage>186</lpage>
          . Verlag Sauerla¨nder, Aarau.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Read</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Nation</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Measurement of formulaic sequences</article-title>
          . In Norbert Schmitt, editor,
          <source>Formulaic sequences: Acquisition, processing and use, Language Learning &amp; Language Teaching</source>
          , pages
          <fpage>23</fpage>
          -
          <lpage>35</lpage>
          . John Benjamins Publishing, Amsterdam.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <given-names>John</given-names>
            <surname>Read</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Assessing vocabulary</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>Marc</given-names>
            <surname>Reznicek</surname>
          </string-name>
          , Maik Walter, Karin Schmidt, Anke Lu¨deling, Hagen Hirschmann, Cedric Krummes, and
          <string-name>
            <given-names>Torsten</given-names>
            <surname>Andreas</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <string-name>
            <given-names>Das</given-names>
            <surname>Falko-Handbuch</surname>
          </string-name>
          .
          <article-title>Korpusaufbau und Annotationen</article-title>
          .
          <source>Technical report, Institut fu¨r deutsche Sprache und Linguistik</source>
          , Humboldt-Universita¨t zu Berlin.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>Helmut</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Probabilistic part-of-speech tagging using decision trees</article-title>
          .
          <source>In International Conference on New Methods in Language Processing</source>
          , pages
          <fpage>44</fpage>
          -
          <lpage>49</lpage>
          , Manchester, UK.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Jan Georg Schneider</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Sprachliche ,Fehler' aus sprachwissenschaftlicher Sicht</article-title>
          . In Sprachreport, volume
          <volume>1</volume>
          -2/
          <year>2013</year>
          , pages
          <fpage>30</fpage>
          -
          <lpage>37</lpage>
          . Institut fr Deutsche Sprache, Mannheim.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <given-names>Stephan</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Mu¨ndlichkeit und Schriftlichkeit aus phraseologischer Perspektive</article-title>
          . In Harald Burger, Dmitrij Dobrovolskij, Peter Ku¨hn, and Neal R. Norrick, editors,
          <source>Phraseologie. Ein internationales Handbuch zeitgeno¨ssischer Forschung</source>
          , volume
          <volume>1</volume>
          , pages
          <fpage>220</fpage>
          -
          <lpage>236</lpage>
          . de Gruyter, Berlin - New York.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <given-names>Torsten</given-names>
            <surname>Steinhoff</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Wissenschaftliche Textkompetenz: Sprachgebrauch und Schreibentwicklung in wissenschaftlichen Texten von Studenten und Experten</article-title>
          , volume
          <volume>280</volume>
          of Reihe Germanistische Linguistik. de Gruyter, Berlin - New York.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <given-names>Torsten</given-names>
            <surname>Steinhoff</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Wortschatz-eine Schaltstelle fu¨r den schulischen Spracherwerb?</article-title>
          , volume
          <volume>17</volume>
          /2009 of SPASS. Universita¨t Siegen,
          <article-title>FB 3</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <given-names>Tobias</given-names>
            <surname>Thelen</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Automatische Analyse orthographischer Leistungen von Schreibanfa¨ngern</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of Osnabru¨ck.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <given-names>Alison</given-names>
            <surname>Wray</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Formulaic language and the lexicon</article-title>
          . Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <given-names>Seid</given-names>
            <surname>Muhie</surname>
          </string-name>
          <string-name>
            <surname>Yimam</surname>
          </string-name>
          , Richard Eckart de Castilho, Iryna Gurevych, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Biemann</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Automatic Annotation Suggestions and Custom Annotation Layers in WebAnno</article-title>
          . In Kalina Bontcheva and Zhu Jingbo, editors,
          <source>Proceedings of the 52nd Annual Meeting of the ACL. System Demonstrations</source>
          , pages
          <fpage>91</fpage>
          -
          <lpage>96</lpage>
          . ACL, jun.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <given-names>Heike</given-names>
            <surname>Zinsmeister</surname>
          </string-name>
          and
          <string-name>
            <given-names>Margit</given-names>
            <surname>Breckle</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>The alesko learner corpus: design-annotationquantitative analyses</article-title>
          .
          <source>Multilingual Corpora and Multilingual Corpus Analysis. Amsterdam: John Benjamins</source>
          , pages
          <fpage>71</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>