<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Experimental Study of Neural Morpheme Segmentation Models for Russian Word Forms</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lomonosov Moscow State University</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Research University Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Morphemic structure of words is useful for various NLP problems, in particular, for deriving a meaning of unknown words in languages with rich morphology, such as Russian. For Russian, several machine learning models for automatic morpheme segmentation of words were built, but only for parsing their lemmas. Meanwhile, signi cantly varying word forms are present in texts, among them unknown words are often encountered, and their lemmas are unknown. The paper reports on experiments for comparing two ways to automatically segment Russian word forms, both ways involve splitting into morphs and classi cation of resulted morphs. The former is based on a neural model trained on a data set automatically augmented with segmented word forms, the latter produces segmentation through predicted lemma and a pre-trained neural morpheme segmentation model for lemmas. It was shown that the models have comparable quality in morpheme segmentation with classi cation, and the model based on the augmented dataset slightly outperforms in word-level classi cation accuracy.</p>
      </abstract>
      <kwd-group>
        <kwd>morphological segmentation</kwd>
        <kwd>morpheme analysis of Russian word forms</kwd>
        <kwd>neural network models for morphology</kwd>
        <kwd>morpheme segmentation with classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Morpheme segmentation as a kind of morphological analysis implies splitting
words into constituent morphs, which are the surface forms of morphemes (roots
and a xes), for example: без-вкус-н-ый, taste-less. Though the task of
automatic morpheme segmentation was studied in early years of natural language
processing (NLP), signi cant progress in its solution has appeared in recent
years, when various machine learning techniques began to be applied.</p>
      <p>
        Since morphemes are the smallest meaningful language units, information
about morphemic structure of words is already in use in various NLP
applications and auxiliary tasks, including machine translation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], recognition of
semantically related words (cognates, paronyms, etc.), creating derivational trees
of words [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], constructing word embeddings [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for handling rare and
out-ofvocabulary words (by deriving their meaning based on distributional word vector
representations) and so on.
      </p>
      <p>Morpheme segmentation is especially topical and at the same time more
difcult for languages with rich morphologies (such as Russian or Finnish). For
morphologically rich languages with many a xes of various types and
meanings, a more complicated task is relevant, which involves besides segmentation
classi cation of segmented morphs. The main types of morphemes are Pre x,
Root, Su x, Ending, for example: без:PREF/вкус:ROOT/н:SUFF/ый:END,
taste:ROOT/less:SUFF.</p>
      <p>
        The rst works on morpheme segmentation were pure statistical and
dictionarybased [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Since during a long time only a small amount of words with labeled
segmented morphemes was available for training, only unsupervised and
semisupervised machine learning techniques were applied, the most known solutions
are implemented in Morfessor system [
        <xref ref-type="bibr" rid="ref12 ref8">8, 12</xref>
        ].
      </p>
      <p>
        The task of morpheme segmentation with classi cation of segmented morphs
remained almost unexplored until recent works [
        <xref ref-type="bibr" rid="ref13 ref4 ref5">4, 5, 13</xref>
        ] undertaken for
Russian, due to powerful supervised machine learning techniques applied to relevant
labeled data, rst of all, the dataset from Tikhonov's derivation dictionary [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
These works presented various supervised models with open-source code:
{ Convolutional neural network (CNN) model3 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ];
{ Gradient boosted decision trees (GBDT) model4 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ];
{ Bidirectional long short-term memory (Bi-LSTM) neural model5 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        The implemented methods consider the task of morpheme segmentation with
classi cation as sequence labeling [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and classify letters of words according to
main types of morphs. As showed by comparative evaluation of the models,
which was undertaken in the works [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], they all achieve F-measure about 98{
99% for detecting morpheme boundaries and they also show high accuracy of
morpheme classi cation: up to 96{98% for letters, and about to 87{89% for whole
words (depending on training datasets and model hyper parameters). Therefore,
these models present state-of-the-art (SOTA) methods for the task of morpheme
segmentation with classi cation.
      </p>
      <p>However, these SOTA models for Russian were developed only for morpheme
segmentation of lemmas (normalized forms of words), so far as only lemmas are
present in the existing labeled datasets. Meanwhile, for morphologically rich and
highly-in ecting Russian language, signi cantly varying word forms are present
in texts, in particular, for verb успеть (to be in time) more than 15 its forms
may be used: успеют, успел, успели and so on. Among various word forms,
unknown ones are often encountered, and their lemmas are unknown. Since it
turned out that the developed SOTA models work poorly for word forms, giving</p>
    </sec>
    <sec id="sec-2">
      <title>3 https://github.com/AlexeySorokin/NeuralMorphemeSegmentation</title>
      <p>4 https://github.com/alesapin/GBDTMorphParsing
5 https://github.com/alesapin/RussianMorphParsing
only about 30% for classi cation accuracy, we aimed to research segmentation
methods applicable for word forms.</p>
      <p>
        In the paper we describe and experimentally compare two ways to
automatically segment Russian word forms, both ways involve splitting into morphs and
classi cation of resulted morphs. The former is based on a neural model trained
on a dataset automatically augmented with segmented and labeled word forms,
the latter produces segmentation through predicted lemma and a pre-trained
segmentation model for lemmas. It is unclear a priory, which of the ways is
preferable, and to evaluate them, we have chosen CNN model as a core for both
ways and have exploited an available dataset containing about 90,000 segmented
words (lemmas) from Tikhonov's dictionary [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. To train the model on word
forms, we have extended this dataset by segmented word forms generated by an
augmentation procedure we have developed.
      </p>
      <p>Experimental evaluation has shown that the model trained on the augmented
dataset (hereafter, model on word forms) and the model trained on lemmas and
supplemented by the rules for segmenting the word form based on its segmented
lemma (hereafter, hybrid model) have comparable quality in morpheme
segmentation with classi cation (as well as comparable with quality of SOTA methods),
while the model on word forms slightly wins in word-level classi cation accuracy
with score 88%.</p>
      <p>The paper starts with an overview of the main works on the morpheme
segmentation, followed by explanation of our augmentation procedure and the
resulted augmented dataset. Then our CNN model architecture and key issues of
training the model on word forms are described, and the results of experiments
with the compared models are reported and discussed. Finally, we present some
conclusion.
2</p>
      <sec id="sec-2-1">
        <title>Related Work</title>
        <p>
          The earliest method of morpheme segmentation was proposed by Z. Harris in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ],
it detects morpheme boundaries by letter variety statistics (LVS) [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Despite
that the method showed only 61% of precision (tested on a small English
dictionary), the statistics was useful in many subsequent researches of the task, in
particular [
          <xref ref-type="bibr" rid="ref11 ref4">4, 11</xref>
          ].
        </p>
        <p>
          In the next years, the most known solutions for morpheme segmentation were
implemented in Morfessor system [
          <xref ref-type="bibr" rid="ref12 ref8">8, 12</xref>
          ], which exploits unsupervised machine
learning methods to be trained on a large unlabelled text. The pure unsupervised
method and its semi-supervised version that uses some labeled data in addition
to the text collection give about 70{80% of F-measure for detected morpheme
boundaries (tested on English, Finnish, and Turkish words).
        </p>
        <p>
          Another kind of semi-supervised machine learning for morpheme
segmentation [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] was based on conditional random elds (CRF), the task was considered
as sequential classifying and labeling letters of a given word. Besides LVS values
and features of letters, the developed CRF classi er exploits some data obtained
by Morfessor, thus increasing F-measure on morpheme boundaries to 84{91%.
        </p>
        <p>
          A pure supervised method with signi cantly better quality for the twofold
task of morpheme segmentation with classi cation was proposed in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], it was
e ective due to applying convolutional neural network (CNN) and training on
the representative labeled data of Tikhonov's dictionary [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. The task is
considered as sequence labeling by classifying letters with 22 classes based on BMES
labeling scheme: the classes account for beginning (B), middle (M), and ending
(E) positions of a letter in the corresponding a x (pre x, root, su x, post x),
as well as single (S) letter variants of a xes, and also hyphen and linking letter
in multi-root and hyphenated words. The trained CNN model is supplemented
with post editing of predicted classes by an auxiliary correcting procedure, which
xes some wrong sequences of classes, according to their probabilities. The model
outperforms all previous morpheme segmentation models, giving F-measure up
to 98% on morpheme boundaries and also achieving classi cation accuracy of
96% for letters and 88% for whole words.
        </p>
        <p>
          Two more supervised machine learning models for morpheme segmentation
with classi cation were developed for Russian words in recent works [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ]: the
rst is based on decision trees with gradient boosting (GBDT), while the second
applies Bi-LSTM neural network. In both models, unlike the CNN model, the
number of letter classes was reduced to 10, since the set of BMES labels is
redundant even for recognizing successive a xes and roots. The GBDT classi er
takes into account features of the letter (in particular, its position in the word
and LVS values), features of its word (some morphological tags), and also
window of 5 previous and 5 subsequent letters. The Bi-LSTM model [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] has three
LSTM layers, the input includes one-hot encoded letters and also some
morphological tags of the word being segmented. Both GBDT and Bi-LSTM morpheme
segmentation models were trained and evaluated on two di erent datasets of
Russian words segmented into labeled morphs, including Tikhonov's dataset.
        </p>
        <p>
          Evaluation of these CNN, GBDT and Bi-LSTM models trained on the same
Russian datasets has showed their comparable quality, about 98{99% of
Fmeasure on morpheme boundaries and 96{98% of classi cation accuracy for
letters and about 87{89% on words [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ]. For now, they are SOTA methods
outperforming the previously developed ones, both for morpheme segmentation
and for segmentation with classi cation. However, they were developed for
segmenting lemmas (normalized word forms), not for various word forms
encountered in texts. Therefore, it seems reasonable to study possible ways to build a
more broad supervised model, and for this purpose, a dataset with word forms
splitted into morphs is needed.
3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Data Augmentation</title>
        <p>
          In order to build a dataset augmented with segmented word forms and thus
suitable for training, we have developed a procedure that produces necessary
segmentation of word forms based on known segmentation of the corresponding
lemmas along with grammatical information about Russian word formation
sufxes and about speci c features of Russian in ection for words of various part
of speech [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>
          The dataset6 based on Tikhonov's dictionary was the source of segmented and
labeled lemmas, and various word forms for a particular lemma were taken from
Open Corpora dictionary7 [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. The dataset encompasses 96,046 words (lemmas)
of main part of speech: nouns, adjectives, verbs, adverbs. Segmented morphs
of words are classi ed according main morpheme types of Russian language
(pre x, root, su x, ending, post x), and successive pre xes and su xes (if any)
are labeled, for example, the verb смазываться (to lubricate) is segmented and
labeled as с:PREF/маз:ROOT/ыва:SUFF/ть:SUFF/ся:POSTFIX.
        </p>
        <p>While applying our augmentation procedure, all lemmas from Tikhonov's
dataset were considered and their corresponding word forms from OpenCorpora
were processed, but those dataset elements that are absent in Open Corpora
dictionary were discarded (approximately, 5 thous. words, the most of them are
very rare, such as гофмейстерский, яспис, спассеровать).</p>
        <p>For a given word form to be segmented and its segmented lemma, the
procedure applies segmenting rules depending on the part of speech and its subclass.
For most nouns, adjective, and participles the rules are quite simple: in the
general case, the given word form and lemma have some common beginning, and if
the rest part of the lemma is labeled as ending, the rest part of the word form
is also annotated as ending, whereas its common part copies segmentation and
labels of the lemma. The following word pair illustrates the rule:
Lemma: разрумяненный раз:PREF/румян:ROOT/енн:SUFF/ый:END
Word form: разрумяненному раз:PREF/румян:ROOT/енн:SUFF/ому:END</p>
        <p>However, for some subclasses of nouns and adjectives (words with a nal
yota: ковбой { ковбоя, соболий { собольего), short adjectives (послушный {
послушен) nouns (words with fugitive vowels: день { дня, замочек { замочка),
as well as for verbs, more di cult segmenting rules were elaborated.</p>
        <p>Speci cally, to segment personal verbal forms and gerund (e.g., увидевши {
у:PREF/вид :ROOT/е:SUFF/вши:SUFF), after detection of the common part
with in nitive form, the segmenting rules sequentially try to recognize and to
label word-formative su xes (ова, ева, ыва, ива, вши, ев, ен, в, л, and so on)
and post x (ся, сь) in the mismatching part of the given word form, and its rest
part (if any) is classi ed as ending. Here is an example:</p>
        <sec id="sec-2-2-1">
          <title>Lemma выходить вы:PREF/ход :ROOT/и:SUFF/ть:SUFF</title>
          <p>Word form выходила вы:PREF/ход :ROOT/и:SUFF/л:SUFF/а:END
In such a way, our augmentation procedure has processed about 92% of
wordforms. Some rare di cult cases were discarded, in particular, consonant
alternation, but such discarding does not impact on the result of comparing.
The resulted dataset augmented with segmented and classi ed word forms has
a total size of 1,130,359 elements: 34% nouns, 32.35% adjectives and participles,
33.56% verbal forms, and 0.07% words of other POS.
6 https://github.com/AlexeySorokin/NeuralMorphemeSegmentation/tree/master/data
7 http://opencorpora.org</p>
          <p>The augmented dataset consists of in ectional paradigms for the processed
lemmas (hereafter, in ectional groups), each group encompasses word forms for a
particular lemma. Groups for nouns and adjectives are relatively small, while for
verbs, a group includes all forms of present, future, and past tense, gerund forms,
up to 31 elements. Here is a fragment of in ectional group for verb обсыпать
(to strew ):
обсыпать об:PREF/сып:ROOT/а:SUFF/ть:SUFF
обсыпал об:PREF/сып:ROOT/а:SUFF/л:SUFF
обсыпала об:PREF/сып:ROOT/а:SUFF/л:SUFF/а:END
обсыпало об:PREF/сып:ROOT/а:SUFF/л:SUFF/о:END
обсыпали об:PREF/сып:ROOT/а:SUFF/л:SUFF/и:END
обсыплю об:PREF/сып:ROOT/л:SUFF/ю:END
обсыпем об:PREF/сып:ROOT/ем:END
обсыплем об:PREF/сып:ROOT/л:SUFF/ем:END
обсыпешь об:PREF/сып:ROOT/ешь:END
обсыплешь об:PREF/сып:ROOT/л:SUFF/ешь:END
обсыпете об:PREF/сып:ROOT/ете:END
4</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Model Architecture</title>
        <p>
          For our study of segmenting word forms and building morpheme segmentation
models, among three SOTA models for morpheme segmentation, namely CNN,
GBDT, and Bi-LSTM we have chosen convolutional neural network (CNN),
because CNN is training much faster than others, and at the same time does not
lose in quality. For simpli cation of experiments, in all our segmentation models
we did not use the auxiliary correction procedure proposed for the original CNN
model, as well as ensembles of several models [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Though such techniques
improve quality of segmentation, but not signi cantly (1{2%), moreover, their
application is not necessary for correct comparison of our model on word forms
and hybrid model, as they use the same neural architecture.
        </p>
        <p>
          All our trained CNN models for segmenting words (word forms) were
implemented with Keras library [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] (based on Tensor ow). As model input we use
letters represented in one-hot encoding format, complementing them with
information about is a particular letter vowel or not, and also with POS tag of
the word, which are taken from morphological analyzer, one-hot encoded and
concatenated with letter vectors. To align all words to the same xed length
(20 letters), we evidently exploit padding, but with masking residual letters
&lt;PAD&gt;(by excluding them while calculating errors), in order to avoid their
in uence on gradient descent. Thereby, one word is represented as an 1120
dimensional vector.
        </p>
        <p>
          The model has several layers, the last layer is fully connected and completed
with a softmax activation function, which outputs a probability distribution
over all possible letter classes. The resulted classes of letters are obtained from
probability distribution with argmax function. Similar to works [
          <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
          ] we apply
simpli ed (i.e., BE) labeling scheme of letters, with 11 classes.
        </p>
        <p>Various hyperparameters of our CNN model were experimentally tested in
preliminary experiments. The resulted model has four layers with 512 lters
in each layer, dropout of 40%, ReLU activation function and kernel size of 5.
More lters in a layer slightly improve the quality (less than 0.5%), but the
model became too heavy both for training and for evaluation. As for additional
layers, they also do not signi cantly improve quality: the model with three layers
gives su cient results, losing to four-layer network only about 1{2%. Among the
gradient descent algorithms (Adam, RMSprop, SGD), the better results were
shown by Adam.
5</p>
      </sec>
      <sec id="sec-2-4">
        <title>Models on the Augmented Dataset</title>
        <p>For all our experiments, the data sets (original Tikhonov's dataset and the
augmented one) were randomly divided in proportion 70:10:20 for training,
validation, and testing, respectively; the training subset of the augmented dataset
includes 791 thous. word forms. After tuning the model with random splits,
for correct evaluation of the models, we have xed our training, testing and
validation sets for reproducibility. All trained and evaluated models are freely
available8.</p>
        <p>In experiments with training our CNN model on the augmented dataset, two
di erent variants of random dividing the dataset were studied:
{ Random mixing of labeled word forms and then splitting them to training
and testing subsets;
{ Random mixing of in ectional groups (each group consists of all word forms
corresponding to the same lemma); and after that splitting to training and
testing subsets is performed (thus, splitting does not divide the groups).</p>
        <p>Thereby we have obtained two trained models, namely, the model on word
forms with simple mixing and the model on word forms with group mixing, the
results of their evaluation are presented in Tables 1, 2. Table 1 shows quality of
only segmentation measured in precision, recall, and F-measure (computed as
mean harmonic of the recall and precision).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>8 https://github.com/alesapin/XMorphy</title>
      <p>recognized classes of letters to the number of all letters, the latter estimates
the ratio of completely correctly segmented words with true classes of all their
letters.</p>
      <p>For comparison, in the last lines of the Tables we have added scores of the
CNN model trained only on lemmas taken from the augmented dataset (more
precise, from its training subset). The scores show that this model signi cantly
loses when applied to word forms: much worse F-measure on morpheme
boundaries (89.52%) and even worse classi cation accuracy (81.19% and 34.30% for
letters and word, respectively). At the same time, almost similar scores for
lemmas con rm consistency of experimental settings.
As for our models on word forms, the model with simple mixing outperforms its
counterpart in all the scores (slightly on morphs boundaries and signi cantly in
classi cation accuracy for words). The explanation is simple: since in ectional
groups may be divided while mixing and splitting to training and testing subsets
for the model with simple mixing, the testing subset can contain some word
forms of the groups, whose elements are present in the training subset, and this
improves evaluation results. At the same time, the quality of the model with
group mixing is comparable with SOTA morpheme segmentation models built
on lemmas. Therefore, it is not quite correct to compare the model with simple
mixing with our hybrid model for segmenting word forms, and we have compared
only the model with group mixing.
6</p>
      <sec id="sec-3-1">
        <title>Comparison with the Hybrid Model</title>
        <p>The hybrid model implements another way to segment word forms, which implies
the following steps:</p>
        <sec id="sec-3-1-1">
          <title>1. converting a given word form into its lemma;</title>
          <p>2. segmenting the latter by the model trained on lemmas (in our experiments,
by the model already learned and indicated in the last lines of Tables 1, 2);
3. transforming the resulted segmented lemma into a segmented word with the
aid of the procedure and segmenting rules described in section 3.</p>
          <p>Using the model already trained on lemmas, we have evaluated the proposed
hybrid model, with precision, recall and F-measure on morph boundaries (for
segmentation, the scores are given in Table 3), and also accuracy both for letters
and whole words (see Table 4). For comparison, in these Tables we repeat scores
of the model on word forms (trained with group mixing). It is important, that
CNN network of the hybrid model was trained on the lemmas taken from the
training dataset for the model on word forms, and it was evaluated on the same
testing set.
One can notice that two our evaluated models for segmenting Russian word
forms have highly close scores for morpheme segmentation, while for classi
cation (Table 4), the model on word forms (group mixing) slightly wins both for
letters and words.
Additionally, we have evaluated ratio of various errors in morpheme
segmentation, depending on wrong boundaries between morphemes of various types,
the results are presented in Table 5. In both models under comparison, the
most frequent errors are related with wrong boundaries between roots and
sufxes, almost half of the errors (column ROOT-SUFF in Table 5). Another types
of frequent errors are wrong recognition of boundary between pre x and root
(PREF-ROOT) and erroneous segmentation of successive roots (ROOT-ROOT)
or su xes (SUFF-SUFF). Below we present some examples of these types. In
general, the presented statistics of errors are about the same, with rare errors of
segmenting word endings.</p>
          <p>{ Root and su x(ROOT-SUFF) { for verb перетлевать, the incorrectly
segmented word form пере:PREF/тл:ROOT/е:SUFF/ва:SUFF/ешь:END
instead of correct пере:PREF/тле:ROOT/ва:SUFF/ешь:END;
{ Pre x and root (PREF-ROOT) { for adjective подоблачный, erroneous
под :ROOT/о:PREF/блач:ROOT/н:SUFF/ою:END instead of the correct
segmentation под :PREF/облач:ROOT/н:SUFF/ою:END;
{ Successive roots and su xes (ROOT-ROOT, SUFF-SUFF) { for adjective
трегубный, instead of correct тр:ROOT/е:LINK/губ :ROOT/н:SUFF/ого:END,
wrong segmentation variant: тре:ROOT/губ :ROOT/н:SUFF/ого:END.
We have developed and evaluated two models of morpheme segmentation with
classi cation, which were proposed speci cally for word forms and are
important for morphologically rich and highly-in ective languages, such as Russian.
The rst model is purely supervised and built on the augmented dataset with
labeling of constituent morphs, the second is the hybrid one combining both the
supervised model based on lemmas and rules for segmenting word forms. For
augmentation of existing dataset with labeled Russian lemmas we have created
the rule-based procedure generating segmented word forms.</p>
          <p>The quality of the developed models turned out to be comparable, and the
model based on the augmented dataset is slightly better in word-level accuracy.
This means, that both models can be used in various NLP experiments with
Russian text. At the same time, the choice of the model may depend on its
computational complexity important in particular applications. For some applied
tasks, a three-layer CNN model instead of our four-layers CNN (as a core of the
hybrid model) is more preferred, as it is faster to train and takes less memory.</p>
          <p>Our future work implies:
{ To resolve some inconsistencies and errors in the original Tikhonov's dataset,
which have been observed while experimenting with it, in order to increase
the quality of the built models for word forms;
{ To elaborate additional segmenting rules for some unconsidered cases of word
forms, such improvement of our augmentation procedure may be useful not
only for improving the morpheme segmentation models, but also for other
tasks.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bocharov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bichineva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Granovsky</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ostapuk</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stepanova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Quality assurance tools in the OpenCorpora project</article-title>
          .
          <source>In.: Computational Linguistics and Intelligent Technologies: Papers from the Annual Int. Conference "Dialogue 2011", Issue</source>
          <volume>10</volume>
          ,
          <issue>101</issue>
          {
          <fpage>109</fpage>
          ,
          <string-name>
            <surname>Moscow</surname>
          </string-name>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Botha</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blunsom</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Compositional morphology for word representations and language modelling</article-title>
          .
          <source>In: Proceedings of the 31th International Conference on Machine Learning (ICML)</source>
          ,
          <year>1899</year>
          {
          <year>1907</year>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bojanowski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grave</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>5</volume>
          , 135{
          <fpage>146</fpage>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bolshakova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sapin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Comparing models of morpheme analysis for Russian words based on machine learning:</article-title>
          <source>In.: Computational Linguistics and Intellectual Technologies: Proc. of the Int. Conference "Dialogue</source>
          <year>2019</year>
          ", Moscow, RGGU (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bolshakova</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sapin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Bi-LSTM Model for Morpheme Segmentation of Russian Words</article-title>
          . In: Ustalov,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Filchenkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Pivovarova</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . (eds)
          <article-title>Arti cial Intelligence and Natural Language</article-title>
          .
          <source>Proceedings of the Int. Conference AINL</source>
          <year>2019</year>
          , CCIS,
          <volume>1119</volume>
          , 151{
          <fpage>160</fpage>
          . Springer, Cham (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chollet</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Keras: Deep learning library for theano and tensor ow</article-title>
          , https://keras.io/,
          <source>last accessed 2020/12/9</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Coltekin, C.:
          <article-title>Improving Successor Variety for Morphological Segmentation</article-title>
          .
          <source>Lot Occasional Series</source>
          ,
          <volume>16</volume>
          , 13{
          <fpage>28</fpage>
          . University of Groningen (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Creutz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagus</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Unsupervised models for morpheme segmentation and morphology learning</article-title>
          ,
          <source>ACM Transactions on Speech and Language Processing</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ),
          <source>Article</source>
          <volume>3</volume>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <article-title>Zellig: Morpheme boundaries within words: Report on a computer test</article-title>
          .
          <source>Transformations and Discourse Analysis Papers</source>
          ,
          <volume>73</volume>
          , 68{
          <fpage>77</fpage>
          (
          <year>1967</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lango</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zabokrtsky</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sevc</surname>
            <given-names>kova</given-names>
          </string-name>
          , M.:
          <article-title>Semi-automatic construction of word-formation networks</article-title>
          .
          <source>Language Resources &amp; Evaluation</source>
          (
          <year>2020</year>
          ). https://doi.org/10.1007/s10579-019-09484-2
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ruokolainen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          et al. :
          <article-title>Painless semi-supervised morphological segmentation using conditional random elds</article-title>
          .
          <source>In: Proceedings of the 14th Conference of the European Chapter of the ACL, Short Papers</source>
          ,
          <volume>84</volume>
          {
          <fpage>89</fpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Smit</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Virpioja</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gronroos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kurimo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Morfessor 2.0: Toolkit for statistical morphological segmentation</article-title>
          .
          <source>In: Proceedings of the Demonstrations at the 14th Conference of the European Chapter of the ACL, Gothenburg</source>
          ,
          <volume>21</volume>
          {
          <fpage>24</fpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Sorokin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kravtsova</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Deep Convolution Networks for Supervised Morpheme Segmentation of Russian Language</article-title>
          . In: Ustalov,
          <string-name>
            <surname>D.</surname>
          </string-name>
          et al.(
          <article-title>eds) Arti cial Intelligence and Natural Language</article-title>
          .
          <source>Proc. of the Int. Conference AINL</source>
          <year>2018</year>
          , CCIS,
          <volume>930</volume>
          , 3{
          <fpage>10</fpage>
          . Springer, Cham (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q. V.</given-names>
          </string-name>
          :
          <article-title>Sequence to sequence learning with neural networks</article-title>
          .
          <source>In: Proceedings of the 27th Int. Conference on Neural Information Processing Systems</source>
          ,
          <volume>2</volume>
          , 3104{
          <fpage>3112</fpage>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Tikhonov</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          :
          <article-title>Word Formation Dictionary of Russian language</article-title>
          . Moscow, Russkij Yazyk Publ. (
          <year>1990</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Zaliznjak</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>Grammatical dictionary of Russian: In ection</article-title>
          . Moscow, Russkij Yazyk Publ. (
          <year>1977</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>