<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Research of Morphemic Vocabulary Volume Based on the Material of German Short Stories</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bence Ny´eki</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>nyeki.bence</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@gmail.com</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Saint-Petersburg State University</institution>
          ,
          <addr-line>Saint-Petersburg, Russian Federation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Slovak Academy of Sciences</institution>
          ,
          <addr-line>L</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vladim ́ır Benko</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The dependence of the vocabulary volume on text sample size has been studied on the material of literary texts [Grebennikov and Assel, 2019] as well as everyday spoken language [Kosareva and Martynenko, 2015]. The present research concerns the study of the morphemic type-token ratio in samples from Franz Kafka's and Thomas Mann's short stories in German. The morphemic annotation of the samples from these texts was carried out manually and was aimed at finding the asymptote of the function “morpheme token-morpheme type.” This helps to conclude whether the expansion of the list of lexemes in these authors' texts is due to the occurrence of new stems or to word formation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Our work is based on the idea that the analysis of the morphemic structure of
wordsdeservesserious attention as compounding and affixation (i.e. the concatenation of
morphemes) serve as productive tools for the creation of morphologically complex words
with semantically transparentstructure in many languages. Consequently, the total number
of morphemes in such a language is smaller than that of lexemes. As a result, it might be
easier to obtain a representative sample of morphemes from a corpus than in case of words.
Of course, it should be notedthat not every lexical unit is semantically compositional at the
morphemic level; this feature corresponds to the problem of syntagmatic idiomaticity. Despite
this fact, knowledge about the units of the morphological system of languages with rich word
formation such as German and Russian can still be considered useful. Another important
aspect is that word formation is considered to play a significant role in text building. For
example, Zemskaya[1992: 164] defines six ways of word formation manifestation as an activity
in speech acts:
derivation from a pretext word or syntagma during speech act production
use of set of derivatives of the same type within a text
formation of different derivatives from the same base
use of words with identical derivational meaning
juxtaposition of derivatives from homonymous words
contrastive use of words with the same root.</p>
      <p>These observations suggest that the study of morphemes in text can be a source of
information for the description of an individual style.</p>
      <p>This research is aimed at the modelling of the morphemic vocabulary growth (in number
of morphemes) as a function of sample size on the material of German literary texts. In the
next section, the research of the type-token ratio in Russian texts is briefly reviewed. In the
present work, the attempt was made to adopt this basic idea to the quantitative study of texts
at the morphemic level. In the third section, the data used for the present research and their
annotation are discussed. Results of fitting a distribution to the datareceived are presented in
the fourthsection. The last section is a summary of our conclusions and further plans.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>In [Kosareva and Martynenko, 2015] research was done on the ORD corpus1 (“One
Day of Speech” corpus) to estimate the asymptote of the function modelling the type-token
dependence in spoken Russian texts. The Weibull and Haustein functions were used for
approximation, the latter of which was found to fit the data better, and the asymptotic level
of about 45,000 lexemes was estimated.With the same methods, approximations by the two
functions on the material of short stories of Russian writers were compared in [Grebennikov
and Assel, 2019] but the authors concluded that the Haustein function is not always preferable
depending on the growth stabilization. It is also worth noting that a significant difference was
found between the authors. The growthof thequantity of types in Chekhov’s short stories
considerably slows down at a sample size of around 150,000 tokens (ca 16,000 lexemes). In
case of Averchenko, however, no clear upper limit of the growth could be set.</p>
      <p>The linguistic annotation of the data used in the present work is based on linguistic
principles. [Mel’ˇcuk, 2006: 390] should be mentioned as a theoretical framework and [Fleischer
and Barz, 2018] as a description of contemporary German word formation. More details will
be given in the next section.</p>
      <p>1http://www.ord-corpus.spbu.ru/SocialStudies/ORD.html</p>
    </sec>
    <sec id="sec-3">
      <title>The data</title>
      <sec id="sec-3-1">
        <title>Preparing the data</title>
        <p>In order to perform a quantitative experiment at a subword level, it is essential to have
a large amount of morphemically annotated text samples. Algorithms of automatic word
segmentation into smaller meaningful units are available. For example, MorfessorFlatCat
is based on hidden Markov models with hidden states “stem”, “prefix”, “suffix” and
“nonmorpheme” to enable unsupervised machine learning [Gr¨onroos et al. 2014]. However, it is
obvious that such an algorithm could not be applied for data annotation in our work.</p>
        <p>Firstly, the segmentations should be maximally precise and based on linguistic principles.
Secondly, even the correct identification of morphs, i.e. minimal meaningful substrings of
the words is unsuitable for the planned experiment. Morphs are tokens of morphemes, each
of which should be represented by a single form. For instance, the morpheme KEIT ‘-ness’
appears in forms -keit and -heit in the German words Wichtigkeit (importance) and Dunkelheit
(darkness), respectively. Thirdly, the elaboration of a fine-grained system of supplementary
morpheme tags is necessary in order to disambiguate homonymous morphemes. Due to these
conditions, the data was annotated manually.</p>
        <p>For the given experiment, the texts of two German authors were chosen. The short stories
of Thomas Mann in a collection available in Project Gutenberg2 as well as all literary works
(but not diary entries and private letters) of Franz Kafka in Project Gutenberg-DE3 were
copied and saved as plain texts. The short stories only may not have contained enough tokens,
so Kafka’s novels were also included into the material. Thus the document containing Kafka’s
texts (ca 290,000 words) is much larger than the collection of Mann’s short stories (ca 39,000
words) but still big enough for sampling.</p>
        <p>The texts were first annotated by TreeTagger4, which is a part-of-speech (POS) tagger
that applies a Markov model and a decision tree [see Schmid 1994, 1995]. Apart from a
relatively large number of POS-tags 5, lemmatization is also provided by the software. This
information simplifies the process of morphemic annotation as a tuple consisting of a POS-tag
and a lemma usually unambiguously determines the correct morphemic analysis.</p>
        <p>Sampling was implemented as follows. From each of the two collections, 60
nonoverlapping text fragments of 250 tokens each were randomly selected.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Morphemic tags</title>
        <p>The words in the obtained samples were analyzed manually; punctuation marks were
ignored. As mentioned above, contextual information was not usually needed to find the
right annotation, so each POS-tag–lemma tuple was processed only once. In the case when
disambiguation was necessary, the given words were analyzed in context.</p>
        <p>The annotation includes the assignment of a string that represents a morpheme and a tag
which is analogous to POS-tags to each identified word segment. The latter (for simplicity,
it may be referred to as a morphemic POS-tag or MPOS-tag) is a two-level tag the first
component of which takes a value from the set “ST”, “PF”, “SF”; the elements of this set stand
2https://www.gutenberg.org/files/36766/36766-0.txt
3https://gutenberg.spiegel.de/autor/franz-kafka-309
4https://www.cis.uni-muenchen.de/ schmid/tools/TreeTagger/
5Documentation for German available: https://www.cis.uni-muenchen.de/ schmid/tools/TreeTagger/
data/stts_guide.pdf
for “stem”, “prefix” and “suffix”, respectively (so they correspond to three of the hidden states
of model applied by MorfessorFlatCat). The second component is determined by the part of
speech of the whole word form when the given morpheme is added to the stem. It takes a value
from the set “NN”, “VV”, “ADJ”, “ADV”, “PP” “DET”, “PPER”, “PREL”, “PD”, “PUF”, “PAV”,
“PWAV”, “KOUS”, “KOUI”, “KON”, “KOKOM”, “APPR”, “PTKVZ”, “PTKZU”, “PTKNEG”,
“PTKA”, “PTKANT”, “ITJ”. These values are connected to the POS-tags assigned to the words
of the samples by TreeTagger. However, it should be noted that this set is smaller than the set
of all TreeTagger POS-tags,so each one of the morphemic annotation units may be associated
with several original TreeTagger tags. For example, attributive and substitutive/predicative
roles are not reflected at the morphemic level. A one-morpheme adjective gets the same
morphemic tag (ST-ADJ) in both attributive (ADJA) and predicative positions (ADJD). Of
course, an attributively used adjectival word form normally consists of at least two morphemes,
the last of which is an inflectional ending. Inflectional morphemes, however, were simply
ignored as their main function is marking syntactic relations within a sentence and can
hardly be considered informative from the aspect of textuality or individual style. Furthermore,
considering inflectional morphemes in the segmentation would make it impossible to determine
the correct morphemic analysis without information about the concrete word form which is
to be annotated. So-called “Fugenelemente” (meaningless segments of words on the borders of
constituent morphemes) were ignored as well. Examples are given in
Table 1-2.</p>
        <p>It is worth noting that word-level POS-tags are also used in modules of word sense
disambiguation systems [Wilks and Stevenson, 1997]. This means that the method of identification
of morphemes by normal form and an MPOS-tag makes use of an analogy between morpheme
and word level.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Annotation principles</title>
        <p>Although it might seem trivial to define a morpheme as the smallest segmental meaningful
part of a word, in practice it can be difficult to find a theoretically supportable morphemic
analysis of a particular word. For example,it needs explanation whether the following words
are to be split into two morphemes or not:</p>
        <sec id="sec-3-3-1">
          <title>a) Mädchen ‘girl’</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>b) bekommen ‘get, receive’</title>
        </sec>
        <sec id="sec-3-3-3">
          <title>c) entdecken ‘discover’</title>
        </sec>
        <sec id="sec-3-3-4">
          <title>d) Augenglas ‘glasses’</title>
          <p>All examples above were marked as single morphemes in the samples. Word a) consist
of the diminutive suffix -chen and a pseudostem which cannot be considered as a sign of
standard German. Following the concepts in [Mel’čuk, 2006: 384], any linguistic sign should
representable as a triplet of a signified, signifier and syntactics, which condition is not met by
the given pseudostem due to the lack of a signified at a synchronic level. In b) both be- and
komm(en) are valid linguistic signs. However, the first is an abstract verbal prefix while the
second is a verb meaning ‘come’. Obviously, the meaning of the derivative cannot be composed
of these semantic elements. The possible constituents of c) are ent-, which is a verbal prefix</p>
          <p>Meaning
‘war years’</p>
          <p>Word form
Kriegsjahren</p>
        </sec>
        <sec id="sec-3-3-5">
          <title>Wirkungen Ha¨nde Table 1: Examples of morphological analysis. Part 1 POS</title>
          <p>Meaning
‘is’
‘one,
anyone’</p>
          <p>STPUF_man</p>
          <p>STDET_dies</p>
          <p>STPRELS_die
ST-PD_die</p>
        </sec>
        <sec id="sec-3-3-6">
          <title>Comment</title>
          <p>In case of suppletion,
the lemma is</p>
          <p>analyzed.</p>
          <p>ST-PUF is a special
tag without an</p>
          <p>equivalent in
TreeTaggertagset. It
is assigned to a small
set of uninflectable</p>
          <p>words
ST-DET (determiner)
is a frequent tag that
is assigned to articles
and demonstrative
pronouns without</p>
          <p>homonyms
The next examples</p>
          <p>demonstrate the
disambiguating effect
of MPOS-tags (in this</p>
          <p>case they
disambiguate whole
words)
that implies the cancellation of an action or a state as one of its senses, and deck(en) (cover,
protect, hide). Although not hiding and discovering something can be regarded as cognitively
associated senses, such a weak relation did not seem to be sufficient to split the word into two
parts. The compound Auge (eye) + Glas (glass) is presented in d). In fact, it is close to be
semantically compositional but still the meaning of Glas is too general.</p>
          <p>In linguistics, degrees of semantic compositionality, which can also be referred to as
transparency, are sometimes distinguished. In [Ransmayr et al., 2016: 267–268] eleven degrees
of transparency of German derived words with the diminutive suffix -chen were defined. The
opaquest derivatives are words without a synchronically identifiable stem, like a) above, or
with a non-noun stem like Frühchen (premature infant).The meanings of the most transparent
derivatives are simply constructed of those of their constituent morphemes. Sometimes they
can be affected by pragmatic restrictions, for example, diminutives denoting clothes such as
Jäckchen (small jacket) are usually used to refer to women’s or children’s clothes. Between
these extremes there are instances of weakly (e.g. metaphorically) motivated compositions
such as Eichhörnchen (squirrel) and Hörnchen (croissant).</p>
          <p>All derived words which are not maximally transparent can be considered, using the
terminology in [Mel’čuk, 2006: 390], quasimorphs, which should be stored in dictionary as
separate entries. However, a typical lexicographic problem often makes this principle more
difficult to apply in practice: it is not always clear how to define the meaning of constituent
morphemes. Taking once again the stem of c as an example, it is not self-evident without
thorough corpus research whether the sense ‘protect, hide’ should be considered a simple metaphor
or a word sense which needs lexicographic description.This is a usual lexicographic problem,
which inevitably occurs and must be solved more or less subjectively in each case. Despite
this fact, some principles of segmentation formulated in advance can serve as considerable
theoretic support. In view of the aspects discussed above, they can be summarized as follows:
1. All constituents of the analyzed word should be meaningful linguistic unitsat a synchronic
level.
2. A bound morpheme is to be added to the morphemic vocabulary only if it occurs in
several non-synonymous words in which it has the same sense (which although might be
highly general or opaque).
3. A complex unit (quasimorph) is preferable only if its meaning is not fully transparent.</p>
          <p>Under the assumption that word meaning can be represented as a finite set of senses,
which are considered to be conventionalized, this means that we expect that all senses
of a morphologically complex word w can be represented as an element of the Cartesian
product of the constituents senses. If some senses of w are elements of the Cartesian
product, while others are not, the correspondent quasimorph should be added to
vocabulary in case of the occurrence of w with a non-compositional sense.</p>
          <p>These principles are simple, but they can help make consequent decisions. For example, it
follows from 2) that such morphemes as the pseudostem of Mädchen (girl) cannot be treated
as separate meaningful segments of words even if several synonymous lexemes exist in the
language with the same pseudostem, e.g. Mädchen and Mädel.This idea can be generalized
to any unique morpheme which occurs only as a constituent of a certain lexeme. Note that
a different viewpoint is also presented in [Fleischer and Barz, 2012: 65] that is based on
structural and not strictly semantic analysis.</p>
          <p>The analysis of the noun Aufzug (the act of lifting, hoisting/elevator/act in theater)
shows the consequences of principle 3). If it occurs as a noun derived straightly from the verb
aufziehen (auf ‘up’ + ziehen ‘pull’), then it is correct to segment it into two constituents.
However, in a sample from Thomas Mann’s texts this is not the case. It occurs in the sense
‘act in theater’ and is therefore added to the morphemic vocabulary as a whole unit.</p>
          <p>It is obvious that the last principle suggests that knowledge about word senses is essential
for morphemic analysis. It makes it necessary to rely on a lexicographic resource. DWDS6
was chosen as such a resource as it ensures quick access not only to lexicographic but also
corpus data if needed. For details about DWDS see [Geyken, 2007].</p>
          <p>However, these ideas were not extended to verbs with separable verb prefixes (and words
compositionally derived from them) although they often form lexical units with semantically
opaque structure. Separable prefixes can take a position very far from the verbal stem in the
sentence, which makes it hard to suggest an appropriate annotation method. This remained
a problem to resolve.</p>
          <p>Since morphemes are not the only elementary linguistic signs [Mel’čuk, 2006: 295–297],
it is necessary to mention how non-segmental signs were handled. In German such signs are
frequently applied by means of conversion and modification. Take the nouns Schritt (step),
Eintreten (the act of coming in) and the verb beenden (finish) as examples. Schritt is derived
from schreiten (to step) by vowel change, Eintreten is obtained by applying conversion to
the verb eintreten (come in) and in beenden the noun Ende (i.e. its allomorph end) can be
observed and as be- is a verbal prefix, it must be concluded that the noun is converted to a
verb (there is also another verb enden with nearly the same meaning). As our annotation
is morphemic, these non-segmental signs are simply ignored. This means that the analyses
of these words are ST-VV_schreit, ST-PTKVZ_ein + ST-VV_tret and PF-VV_be +
STNN_ende, respectively. Of course, this is a simplification: these signs are frequent in German
and the morphological process of conversion is highly productive.</p>
          <p>Now that all major problems of the data annotation are discussed, results can be
presented.
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>Having annotated the text fragments, it was necessary to find a distribution which can
model the growth of morphemic vocabulary as a function of sample size. As mentioned above,
the Weibull and Haustein functions were applied for similar goals [Kosareva and Martynenko,
2015; Grebennikov and Assel, 2019]. Our paper is not aimed at comparing how different
functions fit the data. Only cumulative Weibull distribution was chosen for modelling, which
is usually defined as follows:</p>
      <p>Apart from the exponent, the difference is that the right side of the last equation is
6Digitales Wörterbuch der deutschen Sprache (Digital Dictionary of the German Language):
https://www.dwds.de/
multiplied by Nmax which is the asymptote of the function, i.e. the theoretical maximal
volume of the morphemic vocabulary.</p>
      <p>
        To simplify computation, data was manipulated to enable the use of the former (more
standard) formula. Firstly, empirical values (the registered number of lexemes at a given
sample size) were divided by a hypothetical value of maximal volume. Then the equation was
linearized
        <xref ref-type="bibr" rid="ref2">(analogously to [Kosareva and Martynenko, 2015])</xref>
        to estimate the parameters of
Weibull distribution using the slope and intersect of the linear regression. Now theoretical
values of y could be calculated, which were multiplied by Nmax. The most appropriate value
of Nmax was estimated by the least-squares method: it was determined as the highest observed
vocabulary volume plus 50 (it was clear from the data that the asymptote was significantly
higher than any observed value). Then 50 was added to Nmax again in each following step
until the least sum of quadratic deviations between the theoretical and observed values of y
was reached (some steps could be omitted adding immediately more than 50 to the upper
limit of vocabulary volume). Of course, obtained values are approximate and they can be
defined more precisely; still the results clearly show the difference between the texts of the
two German authors.
      </p>
      <p>Tables 3-4 show some hypothetical values of Nmax and the corresponding sum of quadratic
deviations. The data necessary for calculating theoretical values of y, given the Nmax which
has eventually proved best, are presented in Tables 5-6. These tables serve as illustrations
and they contain only every third observed value of the morphemic type-token function (of
course, the whole sets of observations were used to find distribution parameters). The curves
of cumulative Weibull distributions determined by the calculated parameters and Nmax are
depicted in Figure 1.</p>
      <p>Weibull Theoretic Quadratic
value value yj
devia</p>
      <p>tions
0,12 327,12 146,98
0,19 507,97 143,23
0,25 656,35 347,91
0,29 775,03 782,11
0,33 879,94 787,22
0,37 971,58 339,14
0,4 1054,26 59,95
0,43 1126,25 7,54
0,45 1195,62 0,14
0,48 1259,3 176,94
0,5 1320,37 337,4
0,52 1377,86 355,72
0,54 1431,51 210,51
0,56 1480,44 180,52
0,58 1526,8 249,71
0,59 1570,93 397,28
0,61 1611,8 7,85
0,62 1648,22 116,27
0,64 1683,86 260,59
0,65 1719,99 529,59</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>A logical interpretation of the results is that there are more foreign roots in Thomas
Mann’s short stories than in Franz Kafka’s texts. It can be noticed that Mann uses more
foreign proper names (e.g. Florentinum, Fontana). However, it seems that the type-token ratio
should be considered at word level as well in order to justify this hypothesis. For example, an
author whose morphemic vocabulary grows slowly but uses many different lexemes could rely
on word formation to support expressivity.</p>
      <p>In the present paper, it has been showed that quantitative aspects of derivational
morphology can be regarded as a feature of individual style. A significant difference has been
found between the growth of morphemic vocabulary in two German authors’ texts. In future
research, larger samples need to be taken in order to compare the dependence of the number
of lexeme and morpheme types on sample size. As word formation is a productive linguistic
and cognitive process, this may considerably contribute to the quantitative research of style.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Zemskaya</source>
          , 1992]
          <string-name>
            <surname>Zemskaya E.</surname>
          </string-name>
          (
          <year>1992</year>
          )
          <article-title>Word formation as an activity. (In Rus.) = Slovoobrazovaniekakdeyatel'nost'</article-title>
          .
          <source>Nauka</source>
          , Moscow, Russia. - 221p.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Kosareva and Martynenko</source>
          , 2015]
          <string-name>
            <given-names>Kosareva E. O.</given-names>
            ,
            <surname>Martynenko</surname>
          </string-name>
          <string-name>
            <surname>G. Ya.</surname>
          </string-name>
          (
          <year>2015</year>
          )
          <article-title>The Type-TokenRatio in Everyday Spoken Russian</article-title>
          .
          <source>Structural and Applied Linguistics</source>
          , Vol.
          <volume>11</volume>
          . (In Rus.) = Otnoshenietekst - slovarv
          <source>povsednevnoyustnoyrechi. Strukturnaya I prikladnayalingvistika</source>
          , Vol
          <volume>11</volume>
          .
          <string-name>
            <surname>Saint</surname>
            <given-names>Petersburg</given-names>
          </string-name>
          , Russia. Pp.
          <volume>220</volume>
          -
          <fpage>228</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[Grebennikov and Assel</source>
          , 2019]
          <string-name>
            <surname>Grebennikov</surname>
            <given-names>A. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Assel</surname>
            <given-names>A. N.</given-names>
          </string-name>
          (
          <year>2019</year>
          )
          <article-title>XIX-XX Centuries' Russian Short Stories Corpus</article-title>
          . ApproximationModels. Proceeding of the International Conference «Corpus Linguistics-
          <volume>2019</volume>
          ». Saint Petersburg University Press. (In Rus.) =
          <article-title>Bazarusskogorasskaza XIX-XX vekov</article-title>
          . Modeliapproksimatsii. Trudy mezhdunarodnoykonferentsii«Korpusnayalingvistika - 2019».
          <article-title>Izdatel'stvo Sankt-PeterburgskogoUniversiteta, Saint Petersburg</article-title>
          , Russia.Pp.
          <volume>379</volume>
          -
          <fpage>386</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Mel'čuk, 2006] Mel'
          <string-name>
            <surname>čuk I. A.</surname>
          </string-name>
          (
          <year>2006</year>
          )
          <article-title>: Aspects of the Theory of Morphology. Trends in Linguistics</article-title>
          .
          <source>Studies and Monographs</source>
          <volume>146</volume>
          . Mouton de Gruyter, Berlin, Germany. - 616p. Available at https://anekawarnapendidikan.files.wordpress.com/
          <year>2014</year>
          /04/aspects
          <article-title>-of-the-theoryof-morphology-by-igor-melcuk</article-title>
          .pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Fleischer and Barz</source>
          , 2012]
          <string-name>
            <given-names>Fleischer W.</given-names>
            ,
            <surname>Barz</surname>
          </string-name>
          <string-name>
            <surname>I.</surname>
          </string-name>
          (
          <year>2012</year>
          )
          <article-title>Word Formation of Contemporary German (In Ger</article-title>
          .) =
          <article-title>Wortbildung der deutschen Gegenwartssprache</article-title>
          . Fourth, revisededition. De Gruyter, Berlin, Germany - 481p.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Grönroos et al.,
          <year>2014</year>
          ] Grönroos S.-A.,
          <string-name>
            <surname>Virpioja</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smit</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurimo</surname>
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2014</year>
          )
          <article-title>MorfessorFlatCat: An HMM-based method for unsupervised and semi-supervised learning of morphology</article-title>
          .
          <source>Proceedings of the 25th International Conference on Computational Linguistics. Association for Computational Linguistics</source>
          ,
          <year>August 2014</year>
          , Dublin, Ireland. Pp.
          <volume>1177</volume>
          -
          <fpage>1185</fpage>
          . Available at https://www.aclweb.org/anthology/C14-1111.pdf
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Schmid</source>
          , 1994] Schmid,
          <string-name>
            <surname>H</surname>
          </string-name>
          (
          <year>1994</year>
          )
          <article-title>: Probabilistic Part-of-Speech Tagging Using Decision Trees</article-title>
          .
          <source>Proceedings of International Conference on New Methods in Language Processing</source>
          , Manchester, UK. Pp.
          <volume>44</volume>
          -
          <fpage>49</fpage>
          . Available at https://www.cis.unimuenchen.de/ schmid/tools/TreeTagger/data/tree-tagger1.pdf
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>[Schmid</source>
          , 1995] Schmid,
          <string-name>
            <surname>H.</surname>
          </string-name>
          (
          <year>1995</year>
          )
          <article-title>Improvements in Part-of-Speech Tagging with an Application to German</article-title>
          .
          <source>Proceedings of the ACL SIGDAT-Workshop</source>
          . Dublin, Ireland. Pp.
          <volume>1</volume>
          -
          <fpage>9</fpage>
          . Available at https://www.cis.uni-muenchen.de/ schmid/tools/TreeTagger/data/treetagger2.pdf
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>[Wilks and Stevenson</source>
          , 1997] Wilks Y. and
          <string-name>
            <surname>Stevenson</surname>
            <given-names>M.</given-names>
          </string-name>
          (
          <year>1997</year>
          )
          <article-title>Sense Tagging: Semantic Tagging with a Lexicon</article-title>
          .
          <source>ACL SIGLEX workshop</source>
          , Washington, DC. Pp.
          <volume>74</volume>
          -
          <fpage>78</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Ransmayr et al.,
          <year>2016</year>
          ]
          <string-name>
            <given-names>Ransmayr J.</given-names>
            ,
            <surname>Schwaiger</surname>
          </string-name>
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Durco</surname>
          </string-name>
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Pirker</surname>
          </string-name>
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Dressler</surname>
          </string-name>
          <string-name>
            <surname>W. U.</surname>
          </string-name>
          (
          <year>2016</year>
          )
          <article-title>Grading of the Transparency of Diminutives Ending in -chen: a corpus linguistic research</article-title>
          .
          <source>German Language. Journal for Theory, Practice, Documentation</source>
          <volume>44</volume>
          (
          <issue>3</issue>
          ). (In Ger.) =
          <article-title>Gradierung der Transparenz vonDiminutiven auf -chen: Eine korpuslinguistische Untersuchung</article-title>
          . Deutsche Sprache. ZeitschriftfürTheorie, Praxis, Dokumentation
          <volume>44</volume>
          (
          <issue>3</issue>
          ). Pp.
          <volume>261</volume>
          -
          <fpage>286</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Geyken</source>
          , 2007] GeykenA. (
          <year>2007</year>
          )
          <article-title>The DWDS corpus: A reference corpus for the German language of the 20th century</article-title>
          . In: Collocations and Idioms: Linguistic, lexicographic, and computational aspects. Ed. by Fellbaum C. London, UK. Pp.
          <volume>23</volume>
          -41 Draft available at https://www.dwds.de/dwds_static/publications/text/DWDS-Corpus_
          <article-title>Desc4_draft</article-title>
          .pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>