<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Developing morphologically annotated corpora for minority languages of Russia</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Timofey Arkhangelskiy</string-name>
          <email>tarkhangelskiy@hse.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University Higher School of Economics</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Groningen</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Despite recent progress in developing annotated corpora for minority languages of Russia, still only about a dozen out of about 100 have comprehensive corpora, and even less have computational tools such as machine translation systems or speech recognition modules. However, given that many of them have resources such as dictionaries and grammars, the situation can be improved at relatively low cost. In the paper we demonstrate the pipeline that can be used for developing such corpora, featuring the development of Udmurt and Adyghe corpora. The methods we describe are in principle applicable to any language for which certain kind of linguistic resources are available.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Language corpora are one of the primary
instruments of research in contemporary linguistics.
Corpora allow researchers from all over the world
to analyze raw language data rather than its
interpretations by other linguists in grammars and
articles. Compiling publicly available corpora is
particularly important for more ‘remote’ languages
most researchers have restricted physical access to.</p>
      <p>But, precisely because of their remoteness and
poorer accessibility, there are no corpora for most
such languages, while there are a multitude of
quality corpora for most European languages.
In this paper we speak about developing corpora for
minority languages of Russia. Despite their genetic
diversity, these languages are similar in several
respects, which makes certain approaches
applicable to all or most of them. We will focus
primarily on the cases of Udmurt and Adyghe and
show that solutions we used in their development
can be employed for creating corpora of other
languages of Russia at a reasonably low cost. The
Udmurt corpus was first released in 2014 and is
available at http://web-corpora.net/
UdmurtCorpus/search . Adyghe corpus is currently
under development. The pilot version of the corpus
currently has restricted access, but it is expected to
be released later in 2016 at the same portal.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Languages of Russia and their corpora</title>
      <p>
        There are 93 living indigenous minority languages
spoken in Russia, according to Ethnologue
        <xref ref-type="bibr" rid="ref12">(Lewis
et al., 2016; the number should not be seen as
precise because of the language vs. dialect
uncertainty)</xref>
        . All or almost all of them share several
features important for corpus linguistics.
      </p>
      <p>
        First, vast majority of them are written and have
official orthography, which, with the exception of a
handful of Finnic languages, is based on Cyrillic
alphabet. As virtually all of these orthographies
were developed in the 1930s or later, they represent
the phonology in a pretty straightforward fashion,
unlike in English, Russian or other languages with
long written tradition. Having been developed by
professional linguists, these orthographies faithfully
reflect all phonological distinctions. On the level of
lexicon, these languages share numerous loanwords
from Russian. On the level of grammar, all these
languages are morphologically rich, having on
average more morphologically expressed
grammatical categories than Standard European
languages. What this implies is that in order to be
useful for a wide range of linguistic research, their
corpora should have full morphological tagging
including all morphological categories, rather than
mere POS-tagging. Fine-grained morphological
annotation is also essential for developing ulterior
levels of annotation, such as syntactic parsing or
anaphora resolution, in morphologically rich
languages
        <xref ref-type="bibr" rid="ref15 ref8">(see e.g. Goldberg and Elhadad, 2013 on
syntactic parsing of Hebrew)</xref>
        .
      </p>
      <p>However, what is more important, is that quality
linguistic resources have been created for these
languages. Virtually all of them have grammars and
many have extensive bilingual (usually
X-toRussian) dictionaries. These resources, as we will
show, can be transformed into taggers relatively
easily, and thus are crucial for low-cost corpus
development.</p>
      <p>
        Existing corpora of minority languages of Russia
can be split into two groups: relatively small (almost
always under 100,000 tokens) manually annotated
collections, mainly containing spoken texts, and
larger ones (at least several hundred thousand
tokens, and usually more than one million) with
automatic annotation. Numerous corpora of the
former kind have been collected for various
languages and dialects in linguistic expeditions
since the 1960s. However, their size, which is
naturally constrained by the amount of time and
money required for their collection, is too small for
1 http://web-corpora.net/AvarCorpus/search/
2 http://dag-languages.org/DargwaCorpus/search/
3 http://dag-languages.org/LezgianCorpus/search/
4 http://corpus.ossetic-studies.org/search/ (Iron dialect),
http://corpus-digor.ossetic-studies.org/search/ (Digor
dialect)
5 http://web-corpora.net/RomaniCorpus/search/
many kinds of research, especially if the research
involves statistics. In this paper, we focus on larger,
automatically annotated (and mostly written)
corpora, which are more suitable for low-cost
development. To the best of our knowledge, such
large corpora have been released for the following
13 minority languages of Russia belonging to five
language families:
• East Caucasian: Avar1, Dargwa2, Lezgian3;
• Indo-European: Ossetic4, Romani5;
• Mongolic: Buryat 6
        <xref ref-type="bibr" rid="ref5">(Badmaeva 2015)</xref>
        ,
      </p>
      <p>
        Kalmyk7;
• Turkic: Tatar 8
        <xref ref-type="bibr" rid="ref18">(Suleymanov et al. 2011)</xref>
        ,
Bashkir
        <xref ref-type="bibr" rid="ref7">(Buskunbaeva, Sirazitdinov 2011)</xref>
        ,
Khakas9
        <xref ref-type="bibr" rid="ref17">(Sheymovich 2011)</xref>
        ;
• Uralic: Udmurt, Mari 10
        <xref ref-type="bibr" rid="ref6">(Bradley 2015)</xref>
        ,
      </p>
      <p>Komi11.</p>
      <p>
        Apart from those, there are several ongoing corpus
development projects that we know of, including the
Adyghe corpus project. There are also reports on
developing corpora for Chuvash (Zheltov 2015),
Tuva
        <xref ref-type="bibr" rid="ref15">(Salchak, Bayirool 2013)</xref>
        and Yakut
        <xref ref-type="bibr" rid="ref11">(Leontyev 2014)</xref>
        , but the status of these projects is
unclear.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The pipeline of corpus development</title>
    </sec>
    <sec id="sec-4">
      <title>3.1. Collecting the texts</title>
      <p>
        Books and other printed materials exist for most of
the languages in question, but the cost of scanning,
OCR and proofreading sufficient amount of texts is
prohibitive for a low-budget corpus project. The
only way to obtain a sufficiently large text
collection at low cost is therefore the Internet.
Unfortunately, this constraint makes it impossible to
build corpora for the small and critically endangered
languages that have very low digital vitality, in
terms of
        <xref ref-type="bibr" rid="ref10">Kornai (2013)</xref>
        . However, it seems that
more than one third of the languages in question are
to some extent represented on the web. According
to the estimates of Zaydelman et al. (2016), 30 to 40
languages of Russia have visible amount of texts on
the Internet. The overall size of available texts
6 http://web-corpora.net/BuryatCorpus/search/
7 http://web-corpora.net/KalmykCorpus/search/
8 http://web-corpora.net/TatarCorpus/search/
9 http://khakas.altaica.ru/texts/
10 http://corpus.mari-language.com/
11 http://komicorpora.ru/
varies between a couple of thousand and several
dozen million tokens. Our corpus of Udmurt, 13th
largest minority language and probably the most
digitally well-represented Uralic minority language,
currently contains 7.3 million tokens, which covers
the vast majority of all digitally available texts for
this language. The volumes of the available data are
slowly, but steadily growing: according to the year
distribution of our texts, the growth rate is on
average 0.7 million tokens per year in 2011-2015.
The texts available on the Internet fall mainly into
one of the following groups: digital
newspapers/mass media, blogs/social media and
Wikipedia articles. For the genre composition of the
Udmurt corpus see Table 1. It can be seen that the
corpus is severely unbalanced, as the genre
distribution is skewed in favor of press, followed by
blogs with less than 6%. Our survey of texts in other
minority languages available online suggests that
the distribution is roughly the same for all these
languages (again, with possible exceptions of Tatar
and Bashkir). Lack of balance, which is inevitable
in the proposed method of corpus development, is
one of its largest downsides.
      </p>
      <sec id="sec-4-1">
        <title>Genre</title>
        <p>press
blogs
poetry
fiction</p>
      </sec>
      <sec id="sec-4-2">
        <title>Total</title>
      </sec>
      <sec id="sec-4-3">
        <title>New Testament</title>
      </sec>
      <sec id="sec-4-4">
        <title>Wikipedia articles</title>
        <p>non-fiction</p>
        <p>Tokens
(millions)
6.64
0.42
0.13
0.06
0.03
0.03
0.02
7.33
%
90.56%
5.71%
1.73%
0.84%
0.40%
0.40%
0.36%
Now, the corpus does not include Udmurt posts
from vkontakte, the most popular social network in
Russia, which are estimated to contain more than
0.5 million tokens.</p>
        <p>
          The resulting text collection resembles the corpora
developed within ‘Web as corpus’ approach
          <xref ref-type="bibr" rid="ref9">(Kilgarriff and Grefenstette, 2003)</xref>
          . There is,
however, an important difference between ‘web as
corpus’ and the approach presented here. While the
former aims at gathering vast amounts of data for
NLP purposes, the objective of the latter is to collect
all available texts in a given language, as the size of
the collection is limited for minority languages.
According to our estimate, there are less than 100
web domains containing texts in Udmurt. This order
of magnitude allows for manual inspection of all
relevant web domains (probably with the exception
of Tatar and Bashkir, the most digitally viable of all
these languages) and do not require extraordinary
computational resources to process them.
Another potential pitfall in this process, besides
poor balance, is low quality of texts on Wikipedia.
While for larger languages Wikipedia is often used
as a convenient and reliable source of linguistic
data, Wikipedias in minority languages of Russia
often contain a substantial number of automatically
generated and thus linguistically useless content,
which can be easily seen in their distorted frequency
lists
          <xref ref-type="bibr" rid="ref14">(Orekhov and Reshetnikov, 2014)</xref>
          . If
Wikipedia articles are to be used at all, they should
be filtered (e.g. by length), after which normally
only a small number of articles make it to the
corpus. The corresponding figure in Table 1 shows
the size of the Wikipedia subcorpus after filtering.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>3.2. Morphological tagging</title>
      <p>
        Given that the texts can be collected from the
Internet and that tokenization is not much of a
problem for minority languages of Russia,
development of a morphological tagger is the most
difficult step in corpus development. Both statistical
and dictionary-based taggers require substantial
amount of manual labor if built from scratch. The
former have to be trained on sufficiently large
manually annotated collections, while the latter
require that a grammatical dictionary is compiled
manually. However, the bilingual dictionaries
available for the languages of Russia make
compilation of a grammatical dictionary a much
easier task. This fact, as well as the tradition of
grammatical description of Russian that was started
by
        <xref ref-type="bibr" rid="ref20">Zaliznyak (1977)</xref>
        , is the reason why all corpora
listed in section 2 use the dictionary-based
approach.
      </p>
      <p>
        The idea is to manually write a formalized
description of the morphology based on the
grammars, and then transform a bilingual dictionary
into a grammatical dictionary. In the Udmurt and
Adyghe projects we used the UniParser format and
software for formalized description and tagging,
which were also used for most other aforementioned
corpora
        <xref ref-type="bibr" rid="ref3">(Arkhangelskiy et al., 2012)</xref>
        . There are also
plenty of alternatives, including PC-KIMMO
        <xref ref-type="bibr" rid="ref1">(Antworth, 1992)</xref>
        , used in the Tatar corpus tagger,
or giellatekno infrastructure
        <xref ref-type="bibr" rid="ref13">(Moshagen et al.,
2013)</xref>
        .
      </p>
      <p>The central problem in this step is the fact that
bilingual dictionaries normally do not contain
necessary grammatical information such as part of
speech or declension / vowel harmony type; they
have to be automatically restored. We combined
three approaches to address this issue.</p>
      <p>First, the form of the lemma in some cases clearly
indicates its part of speech. In Udmurt, we tagged as
verbs all lemmata ending in -ɨnɨ or -anɨ (markers of
the infinitive). Manual check found that only one
word, ǯɨnɨ ‘half’, was tagged incorrectly during this
step.</p>
      <p>
        All other parts of speech, however, did not have any
markers that could be used as clues for
part-ofspeech tagging. In Adyghe, a polysynthetic
language where bare stems are used as citation
forms and parts of speech in general are not well
differentiated, this was impossible altogether. The
approach we used for these cases was using the tag
given by a Russian tagger (specifically, mystem
        <xref ref-type="bibr" rid="ref16">(Segalovich, 2003)</xref>
        ) to the first non-abbreviated
word of the translation. This worked surprisingly
well: in Udmurt, around 85% of these tags proved
to be correct. The wrong tags came primarily from
two sources. First, some of the translation
equivalents in both Udmurt and Adyghe
dictionaries had several possible analyses, e.g. in
adjectives which are commonly used as
(substantivized) nouns. Second, Udmurt has an
extensive (hundreds of items) inventory of
ideophones, or imitative words that do not have
Russian translation equivalents and are translated
periphrastically, e.g. čʼɨš-čʼaš ‘about burning of wet
wood’.
      </p>
      <p>
        As the final approach we wrote some simple scripts.
In Udmurt, the only additional field needed for
tagging beyond part of speech is the conjugation
type, which is determined by the last vowel of the
stem (Winkler 2000: 45). In Adyghe, there is a
regular e/a alternation in stems of a certain kind
        <xref ref-type="bibr" rid="ref2">(Arkadyev and Testelets, 2009)</xref>
        . Whenever the
script sees an alternated stem in the lemma, it
generates the base form and adds it to the list of stem
allomorphs in the grammatical dictionary.
Apart from the challenge posed by part of speech
tags, the excessiveness of the information in the
dictionary can be an obstacle. One of its
manifestations is abundance of synonymous
translation equivalents, usage examples and phrases
in dictionary entries, which are usually not needed
in the corpus and thus have to be cut out. In Adyghe,
for which several dictionaries were used as an input,
this lead to especially long translations, since
different dictionaries used different synonyms for
translating the same word. This issue was addressed
by passing the translation equivalent through a
number of transformations. All secondary meanings
and comments were removed by cutting out
segments in parentheses, after semicolons and after
colons if certain threshold length has been reached.
In the case of Adyghe, the synonyms in translation
equivalents were rearranged in a decreasing
frequency order (according to the data from Russian
National Corpus), so that the most frequent
synonym appeared first and all the rest could be
easily deleted during the manual proofreading.
Another manifestation of this problem lies in the list
of words included in the dictionary. Apart from too
many (potential) Russian loanwords that will
probably never appear in a corpus, such as
aerosyomka ‘aerial photography’, dictionaries for
minority languages of Russia tend to include
absolutely compositional and productive
derivatives or word forms as separate entries. For
both Udmurt and Adyghe, this involves, first and
foremost, verbal derivation. In Udmurt, causative,
detransitive and iterative forms of most verbs were
included in the corpus, however only a handful of
them have somewhat non-compositional meaning.
If left as is, the tagger based on such a dictionary
would give seemingly ambiguous results for words
containing these affixes. For example, the word
vera-lʼlʼa-z speak-ITER-PST.3SG will be
ambiguously tagged as ITER.PST.3SG form of the
verb veranɨ ‘speak’ and as PST.3SG form of the verb
veralʼlʼanɨ ‘speak (repeatedly)’. These words were
removed from the dictionary with a script that
searched for a marker of one of these categories and
checked if the remaining part was listed in the
dictionary as a separate verbal stem. The situation is
somewhat more difficult in Adyghe. Adyghe is a
polysynthetic language, which means that the stem
can attach numerous derivational affixes. While
most of these combinations are perfectly
compositional, some are not, therefore manual
check of all such complex stems is required. Words
involving non-compositional combinations of stems
and derivational affixes in Adyghe corpus get two
levels of annotation, one for the original stem, the
other for the combination of the stem and the affix
        <xref ref-type="bibr" rid="ref4">(Arkhangelskiy and Lander, 2016)</xref>
        . The interfaces
enables users to search either for all occurrences of
a given stem, or only those occurrences where it is
not part of a non-compositional combination.
Finally, the dictionaries have to be extended
manually by adding irregular words (mostly
pronouns) and frequent regular words that were
absent due to scarcity of the source dictionary or
conversion errors. The Udmurt tagger, to which all
pronouns and no more than a hundred other frequent
lexemes were added manually, currently covers
about 88% of the tokens in the corpus. Here is an
example of an entry from the resulting Udmurt
dictionary:
-lexeme
lex: кизьыны
stem: киз.
gramm: V,I
paradigm: connect_verbs-I-soft
trans_ru: сеять, посеять, засеять
The entry contains fields indicating its lemma, stem,
grammar tags, set of inflectional affixes and Russian
translation.
      </p>
    </sec>
    <sec id="sec-6">
      <title>4. Conclusion</title>
      <p>The presented pipeline, which we used for
developing the Udmurt corpus and which is
currently used in the Adyghe corpus project, allows
for relatively inexpensive construction of digital
corpora. The proposed approach is applicable to
digitally represented languages which have
grammars and dictionaries. According to our
estimates, there are still 15 to 20 minority languages
of Russia that lack comprehensive written corpora
but have enough resources so that this approach can
be applied to them.</p>
      <p>The resulting corpora will only have morphological
annotation and will probably be severely
unbalanced. However, development of such corpora
constitutes a necessary step for introducing higher
levels of annotation and for achieving better
balance. Our ongoing experiments with OCRed
Udmurt books suggest that adding a simple
ngrambased postprocessor trained on corpus data may
significantly improve its quality, reduce the cost of
proofreading and thus eventually lead to adding
books to the corpus. Finally, language models
trained on such corpora enable other NLP tools for
these other under-resourced languages (as an
example, Yandex launched Udmurt-Russian
machine translation service in 2016, which uses
language model trained on the Udmurt corpus).
This, in turn, can lead to preserving and
revitalization of the minority languages.
descriptive statistics. Computational Linguistics and
Intellectual Technologies: papers from the Annual
conference “Dialogue”.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Antworth</surname>
            ,
            <given-names>E. L.</given-names>
          </string-name>
          <year>1992</year>
          .
          <article-title>Glossing text with the PCKIMMO morphological parser</article-title>
          .
          <source>Computers and the Humanities</source>
          ,
          <volume>26</volume>
          (
          <issue>5-6</issue>
          ):
          <fpage>389</fpage>
          -
          <lpage>398</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Arkadyev</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Testelets</surname>
          </string-name>
          , Ya.
          <year>2009</year>
          .
          <article-title>O trekh cheredovaniyakh v adygejskom yazyke [On three alternations in the Adyghe language]</article-title>
          .
          <source>Ya</source>
          . Testelets (ed.),
          <source>Aspekty polisintetizma [Aspects of polysynthesis]</source>
          .
          <volume>121</volume>
          -
          <fpage>145</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Arkhangelskiy</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Belyaev</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Vydrin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>The creation of large-scaled annotated corpora of minority languages using UniParser and the EANC platform</article-title>
          .
          <source>Proceedings of COLING 2012: Posters, Ch</source>
          .
          <volume>9</volume>
          :
          <fpage>83</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Arkhangelskiy</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lander</surname>
          </string-name>
          , Yu.
          <year>2016</year>
          .
          <article-title>Developing a polysynthetic language corpus: problems and solutions</article-title>
          .
          <source>Computational Linguistics and Intellectual Technologies: papers from the Annual conference “Dialogue”.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Badmaeva</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Natsionalnyj korpus buryatskogo yazyka: predposylki i perspektivnye puti razrabotki [Buryat National Corpus: prerequisites and future development trajectories]</article-title>
          .
          <source>Vestnik Buryatskogo gosudarstvennogo universiteta</source>
          ,
          <volume>72</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Bradley</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>corpus.mari-language.com: A Rudimentary Corpus Searchable by Syntactic and Morphological Patterns</article-title>
          . Septentrio Conference Series,
          <volume>2</volume>
          :
          <fpage>57</fpage>
          -
          <lpage>68</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Buskunbaeva</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Sirazitdinov</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Sistema razmetok v natsionalnom korpuse bashkirskogo yazyka [Annotation system in Bashkir National Corpus]</article-title>
          . Proceedings of “
          <article-title>Yazyki menshinstv v kompyuternykh tekhnologiyakh: opyt, zadachi i perspektivy”</article-title>
          , Yoshkar-Ola:
          <fpage>46</fpage>
          -
          <lpage>51</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Goldberg</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Elhadad</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Word Segmentation, Unknown-word Resolution, and Morphological Agreement in a Hebrew Parsing System</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>39</volume>
          /1:
          <fpage>121</fpage>
          -
          <lpage>160</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Kilgarriff</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Grefenstette</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <year>2003</year>
          .
          <article-title>Introduction to the special issue on the web as corpus</article-title>
          .
          <source>Computational linguistics</source>
          ,
          <volume>29</volume>
          (
          <issue>3</issue>
          ):
          <fpage>333</fpage>
          -
          <lpage>347</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Kornai</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Digital Language Death</article-title>
          .
          <source>PLoS ONE</source>
          <volume>8</volume>
          (
          <issue>10</issue>
          ): e77056. doi:
          <volume>10</volume>
          .1371/journal.pone.0077056
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Leontyev</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Natsionalnyj korpus Internet-sajtov gazet na yakutskom yazyke [National corpus of newspaper web sites in Yakut]</article-title>
          .
          <source>Zhurnal nauchnyx i prikladnykh issledovaniy "Infinity"</source>
          ,
          <volume>4</volume>
          :
          <fpage>35</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Paul</surname>
          </string-name>
          , Gary F. Simons, and Charles D.
          <source>Fennig (eds.)</source>
          .
          <year>2016</year>
          .
          <article-title>Ethnologue: Languages of the World, Nineteenth edition</article-title>
          . SIL International, Dallas, Texas. Online version: http://www.ethnologue.com.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Moshagen</surname>
            ,
            <given-names>S. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pirinen</surname>
            ,
            <given-names>T. A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Trosterud</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Building an open-source development infrastructure for language technology projects</article-title>
          .
          <source>Proceedings of the 19th Nordic Conference of Computational Linguistics (NODALIDA</source>
          <year>2013</year>
          ),
          <fpage>343</fpage>
          -
          <lpage>352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Orekhov</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Reshetnikov</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>K otsenke vikipedii kak lingvisticheskogo istochnika [Assessing Wikipedia as a linguistic source]. Sovremennyj russkiy yazyk v internete</article-title>
          , Moscow:
          <fpage>310</fpage>
          -
          <lpage>321</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Salchak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Bayirool</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Elektronnyj korpus tuvinskogo yazyka: sostoyanie, problemy [Electronic corpus of Tuva: current state, challenges]</article-title>
          . Mir nauki, kultury, obrazovaniya,
          <volume>6</volume>
          (
          <issue>43</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Segalovich</surname>
            ,
            <given-names>I. 2003</given-names>
          </string-name>
          <article-title>A fast morphological algorithm with unknown word guessing induced by a dictionary for a web search engine</article-title>
          .
          <source>Proceedings of the International Conference on Machine Learning; Models, Technologies and Applications</source>
          . MLMTA'
          <volume>03</volume>
          . - Las Vegas:
          <fpage>273</fpage>
          -
          <lpage>280</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Sheymovich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Morfologicheskaya razmetka korpusa khakasskogo yazyka [Morphological annotation of the Khakas corpus]</article-title>
          .
          <source>Rossiyskaya tyurkologiya</source>
          ,
          <volume>2</volume>
          (
          <issue>5</issue>
          ):
          <fpage>48</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Suleymanov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khakimov</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Gilmullin</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>Korpus tatarskogo yazyka: kontseptualnye i lingvisticheskie aspekty [Tatar corpus: conceptual and linguistic aspects]</article-title>
          .
          <source>Filologiya i kultura</source>
          ,
          <volume>26</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Winkler</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <year>2001</year>
          . Udmurt. Lincom Europa, München.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Zaliznyak</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>1977</year>
          .
          <article-title>Grammaticheskiy slovar russkogo yazyka: slovoizmenenie [Grammatical dictionary of the Russian language: inflection]</article-title>
          .
          <source>Russkiy yazyk</source>
          , Moscow.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Zaydelman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krylova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orekhov</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Popov</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Stepanova</surname>
            , E. Russian minority languages: Zheltov,
            <given-names>P.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>Sozdanie natsionalnogo korpusa chuvashskogo yazyka: problemy i perspektivy [Development of Chuvash National Corpus: Challenges and perspectives]</article-title>
          .
          <source>Sovremennye problemy nauki i obrazovaniya</source>
          ,
          <volume>1</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>