<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>About Turkic Morpheme Portal</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Applied Semiotics of Tatarstan Academy of Sciences</institution>
          ,
          <addr-line>Kazan</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kazan Federal University</institution>
          ,
          <addr-line>Kazan</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Doing scientific research in fields of turkology and agglutinative languages typology requires software that takes into account the structural and functional features of languages in question. This paper presents a description of the Turkic Morpheme Portal, in the development of which an integrated approach is used for the development of computer linguistic models and technologies for Turkic languages processing. This portal was created on the basis of the structural-parametric functional model of the Turkic morpheme and contains special linguistic databases that describe the categories of Turkic languages at different levels: morphological, syntactic, and semantic. The problems of developing complex multilingual linguistic models for low-resource languages and their software implementation are considered. The prospects of using the created portal as a base for the development of linguistic software, as well as an information and reference system, including a multilingual thesaurus, and as a platform for communication of specialists are given.</p>
      </abstract>
      <kwd-group>
        <kwd>Linguistic resource</kwd>
        <kwd>Ontology</kwd>
        <kwd>Thesaurus</kwd>
        <kwd>Turkology</kwd>
        <kwd>Multilingual model</kwd>
        <kwd>Morphology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The importance of Turkic Morpheme Portal development is determined by the
situation that has established around development of software for Turkic languages
processing, which is formed with a number of factors.</p>
      <p>Firstly, historically, development of linguistic software for English is leading in
field of computational linguistics in the world, and in Russia, in turn, it is software for
Russian, therefore, researchers and developers of scientific software for other
languages study technologies for English and Russian, basing their research on these
methods. However, the Turkic languages, unlike Russian and English, fully belong to
agglutinative languages family and are structurally quite different. This means that
computational linguistic models specialized on the agglutinative languages and
technologies for processing of this type of languages are needed.</p>
      <p>Secondly, all Turkic languages, except for Turkish, are low-resource languages and
their lag behind the resource-rich languages continues to accumulate. One of the
reaCopyright © 2020 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
sons for this is lack of specialists working on development of linguistic resources and
computer processing software for Turkic languages.</p>
      <p>Thirdly, despite the increase in research on Turkic languages processing, there is
practically no real integration of the results. According to resolution of TurkLang
2017 Conference, joining efforts of groups working on different Turkic languages can
provide solution to this problem.</p>
      <p>Fourthly, there is a duplication of linguistic models and resources, as well as
language processing software, which are basically 70-80 percent common to all Turkic
languages. Therefore, overcoming such duplication and combining the efforts in
collaborative development and interchange of linguistic software is urgent.</p>
      <p>
        Fifthly, the idea of creating a machine fund of Turkic languages was proposed back
in 1988 in the "Soviet Turkology" journal by V.G. Guzev, R.G. Piotrovsky,
A.M. Shcherbak in their article "On the Creation of the Machine Fund of Turkic
Languages" [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this article, a number of ideas about the principles of machine fund
organization were presented, which remain relevant at the present time. However, for
a number of reasons, these ideas were never implemented. We'll look at them in detail
in the next chapter.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <sec id="sec-2-1">
        <title>Structure of NLP domain</title>
        <p>
          Current state of NLP research domain is characterized by rapid development of
machine learning and neural networks technologies. A problem with machine learning
systems is the need for large amounts of data with different types of annotation, such
as morphological, semantic, and syntactic. The development of machine learning
methods has led to the fact that a number of problems previously solved using
ontologies, frames and semantic networks, began to be solved without these resources. One
of these methods consists in constructing of vector representations for words in a
lowdimensional space, namely word2vec [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], which allows to reduce the problem of
words semantic proximity evaluation to calculating the cosine of the angle between
the corresponding word vectors. Today this method is widely used for automatic
construction and expansion of semantic resources, as well as in classification and
clustering tasks.
        </p>
        <p>
          Despite the advances in machine learning, high quality ontology models, frames,
and semantic networks are still important linguistic resources. They are indispensable
when high accuracy is required, even if it is achieved by narrowing the lexical
coverage. One of the tasks in which linguistic resources built by experts like WordNet [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ],
are still out of competition is word sense disambiguation. Also, the ontological
technologies and semantic networks are effective in evaluation of methods and systems
for natural language processing. For these tasks such resources are used as a gold
standard against which a comparison in some metrics is made.
        </p>
        <p>
          Machine learning methods and ontological models represent two main directions
for artificial intelligence development. Here, neural networks imitate empirical
thinking and perception, while ontological models express logical and abstract thinking.
Machine learning methods require large corpora, and recently a number of corpora for
Turkic languages were developed including but not limited to [
          <xref ref-type="bibr" rid="ref4 ref5 ref6">4-6</xref>
          ]. Other examples
can be found at [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. However, most of them cannot go beyond morphological markup,
precisely because of the lack of ontological resources and software for syntactic and
semantic analysis.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Linguistic resources</title>
        <p>
          Multilingual databases and software constitute an effective typology tool for different
languages, language units, properties, and phenomena classification. An example of
such tool is “Global Lexicostatistical Database” [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. This database presents the basic
vocabulary of the world's languages in the comparative manner. It is intended for the
formation of unified system of basic vocabulary lists to research the degree of world's
languages relations based on the percentage of common words. Then the software
system makes a classification genealogical tree of world languages.
        </p>
        <p>
          Another function set that increases the efficiency and presentability of research
software can be the combination of a linguistic service functions with geographic
information systems. As an example of such service, one can propose a new resource
for Turkic languages called “Maps for Turkic languages” [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which is being
developed by a group of Russian turkologists.
        </p>
        <p>
          There were various attempts to combine different types of linguistic resources into
a common database. One of these projects is BabelNet [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], a unified linguistic
resource that combines 47 resources, including Wikipedia, WordNet, Wikidata,
FrameNet, VerbNet, ImageNet and others. The integration of these resources in BabelNet is
done automatically using a linking and lexical gap filling algorithm. It contains Babel
synsets, presented in many languages and connected by a huge number of semantic
relations: in version 4.0, out of 832 million meanings, more than 6 million concepts
and 9.5 million named entities, linked by more than 1 billion semantic relationships
are extracted.
        </p>
        <p>
          Another example of linguistic resources unification is the SemLink project [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ],
which was developed at the University of Colorado. The authors of this project
propose an approach to unification of the following resources: PropBank, VerbNet.
FrameNet, WordNet.
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <sec id="sec-3-1">
        <title>Problem definition</title>
        <p>The analysis shows that among all the Turkic languages in all global international
projects for linguistic resources, only one actively participated which is Turkish
language. As a result, almost all Turkic languages, except Turkish, belong to
lowresource languages. Therefore, an urgent task is formulated: to combine various kinds
of linguistic resources for Turkic languages in one resource. When solving this
problem, a hypothesis is put forward, that linguistic resources for the Turkic languages can
be combined using the Turkic morpheme as a unifying element.</p>
        <p>
          The choice of this element is based on structural features of Turkic languages.
V.A. Plungyan [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] divides all languages according to morphological models into
three types:
1. Elemental-combinatorial (Item and Arrangement) morphological model. The main
structural tool of this model is linear segmentation.
2. Elemental-procedural (Item and Process) morphological model. In languages with
this model some allomorphs are considered as initial, and others as derivatives,
which can be obtained from the former by applying operations of "phonological
processes".
3. Verbal-paradigmatic (Word and Paradigm) morphological model. In this model,
there is a complete rejection of morphemic division when describing inflection. It
is the word form that is chosen to be the minimum unit of grammatical description
here.
        </p>
        <p>According to this classification, the Turkic languages belong to
elementalcombinatorial type. This allows us to consider the Turkic morphemes as integral
elements in the Turkic language system, which are in different types of relationships
with each other and with other elements of the language and semantics.</p>
        <p>The next argument to the choice of Turkic morpheme as a basic connecting
element is the following fact. The phonological and morphological levels of language
contain a finite number of units, and it is possible to compose the finite alphabet of
language in these units. At the same time, the syntactic and semantic levels of
language operate with an infinite number of language units, which complicates the task
of constructing a linguistic model.</p>
        <p>In this regard, it is possible to extract the meanings of reproducible language units,
i.e., units stored in memory in a complete form. For this, only two types of values can
be distinguished:
 Meanings of morphemes;
 Meanings of words (including meanings of phraseological units).</p>
        <p>
          B.Y. Gorodetsky proposes the folowing levels of analysis of two-sided language
units [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]:
 Morpho-semantic level, represented by meanings of all morphemes distinguished
in a given language;
 Lexico-semantic level, represented by meanings of all lexical units included in the
lexicon of a given language.
        </p>
        <p>Units on each level connect using morpho-semantic and lexical-semantic relations.
The feature of Turkic languages is that a lexeme can coincide with a root morpheme.
As a result, in Turkic languages both the unit of morpho-semantic level and the unit
of lexico-semantic level are essentially the same, it is a morpheme. This determines
the choice of the Turkic morpheme as the basic unit used to connect linguistic models
of different levels.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Turkic Morpheme Model</title>
        <p>
          On the basis of the stated hypotheses, a complex linguistic Turkic Morpheme Model
was developed (Fig. 1). This model is a further development and generalization of the
structural and functional model of Tatar affixal morpheme [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. The original version
was expanded onto all Turkic languages with inclusion of root morphemes. It is
assumed that the resulting complex linguistic model can improve the efficiency of
multilingual word processing software development. It also serves as a basis for solving
other fundamental and applied problems, which require conceptual and formal
linguistic models, common databases, as well as software based on these models. This is
facilitated by the pragmatically oriented approach proposed by D.S. Suleymanov [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
On the basis of the proposed model, the Turkic Morpheme Portal was developed.
Web portal is a web site that provides access to various services in a particular area.
The same way, the Turkic Morpheme Portal is a set of services for the Turkic
languages processing. The portal has a whole range of functions:
1. An information and reference system on the Turkic languages grammar;
2. A place for communication of specialists in computer processing of Turkic
languages for joint turkology research;
3. A set of linguistic resources for Turkic languages, including multilingual thesaurus
and frame ontologies;
4. A database for linguistic services, such as text processors on different levels of
language structure, presented as a pipeline.
5. A service for unification and connection of linguistic resources for Turkic
languages, including corpora.
6. A source of linguistic data on Turkic languages for training the machine learning
models.
        </p>
        <p>The main purpose of the portal is supporting the research and development in the field
of Turkic studies. This feature determines the requirements for the portal as a
multipurpose scientific resource. The main properties of a scientific resource are as
follows:
1. Multifunctionality. The resource is applicable to various tasks in which Turkic
languages texts processing is required (machine translation, multilingual search,
question-answer systems, information extraction).
2. Multilingualism. Resource software and algorithms are separated from linguistic
database, while being focused on processing of Turkic languages, and are equally
applicable to any language in this group.
3. Pragmatic orientation. The software is precisely oriented towards the processing of
languages from Turkic family, it is not universal.
4. Stratification. The sentence analysis algorithms are based on representations of the
sentence at several linguistic levels from morphological to syntactic and semantic.
5. Interactivity. Human-machine dialog interaction is required to resolve complex
cases of ambiguity.</p>
        <p>These principles were taken as a basis for Turkic Morpheme Portal development.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Portal description</title>
      <sec id="sec-4-1">
        <title>Portal infrastructure</title>
        <p>The Turkic Morpheme Portal has wide functionality available in two modes of portal
operation: the reader mode and the expert mode, which actually represent two
subsystems of the portal. The reader mode represents the reference subsystem, and the
expert mode represents the research subsystem.</p>
        <p>The reader can view the materials from system database, but they do not have the
permission to edit it (Fig. 2). The portal provides the reader with descriptions of
Turkic language elements collected by turkology specialists at different language levels,
the rules for their combination (morphotactics), as well as the expressed meanings
(semantic). The feature of this linguistic information presentation is that it is
concentrated, classified and formalized in a single database, as well as equipped with options
for searching and filtering queries.
The expert mode (Fig. 3) provides the user with a wide range of functions for carrying
out research activities such as: collecting core information, classification, comparative
statistical analysis, hypotheses verification. To take advantage of the expert mode, it
is necessary for the user to pass authorization and get confirmation from the
administrator, after which they get access to work with one selected language of their choice.
The expert can also view the other languages in the reader mode. Within the
framework of his chosen language, the expert has the permission to fill in the
languagespecific part of the model as well as some elements of the common part.</p>
        <p>The third mode of portal usage, which is inaccessible for ordinary users, is the
administrator mode (Fig. 4). It gives direct access to the database for any of the Turkic
languages presented in the portal, as well as to the infrastructure part. The
administrator confirms the expert roles, gives them the permission for portal translation, has
access to the system log and is engaged in system technical support. Administrator
mode is implemented using the standard framework tools.</p>
        <p>Fig. 3. Affixal morpheme page in expert mode
Fig. 4. Admin panel, affixal morphemes page
As a service with multilingual Turkic Morpheme Model support, the portal must have
a multilingual interface. The portal framework toolbox provides a localization
mechanism which allows users to translate the portal into other languages using the
localization module (Fig. 5).
The portal provides a forum for collaborative discussions and for user feedback. It
also implements a wiki-like system that gives the user additional information about
the model and the linguistic terminology. In addition, tools for summary overview are
implemented, such as statistics of database and summary tables (Fig. 6), which
provide an interlanguage representation of current model state with HTML and Excel
formats.
The common part of the portal database contains the data for linguistic categories that
are common to all languages presented in the model, such as: grammatical categories,
grammatical values (grammemes, quasigrammemes, derivatemes) and concepts
(objects, actions, attributes). Fig. 7 shows a conceptual database diagram for the common
part of the model. Concepts here express some meanings and are arranged in a
thesaurus with different relations between them. Grammatical categories and values express
some linguistic modifiers, while grammatical values can be composite, i.e., consist of
several nested values.
As an example of working with the common part, let's consider the procedure of
filling in the grammatical categories. By clicking on the left menu item “Grammatical
categories”, the user gets access to the list of grammatical categories (Fig. 8).
The attributes “Typological name (English)” and “Typological name (Russian)” are
filled in by the administrator, and users, both readers and experts, can only view them.
The “National Name” field for a specific language, is intended to be filled by an
expert. To enter information about a category in the expert’s language, he needs to
select the category of interest in the list, after which a view of this category will open,
where a form for national name editing is available (Fig. 9). In addition to the national
name itself, an expert can enter the category description in their language and the
source from which such a description is taken. In reader mode these items are
provided read-only.</p>
        <p>Work with other linguistic elements of the common part is done in a similar way,
that is, their main part is filled in by portal administration, and then the expert has an
opportunity to clarify the linguistic nuances for the corresponding Turkic language. In
addition, a separate user role of typologist expert is set. These users can add their own
concepts to the database (Fig. 10).</p>
      </sec>
      <sec id="sec-4-2">
        <title>Language-specific part</title>
        <p>The language-specific part of the model includes those linguistic categories that
completely depend on the specifics of a particular language. The following categories are
implemented: affixal morphemes, analytical morphemes (particles, adpositions,
auxiliary verbs), root morphemes and morphotactic rules. Affixal and analytical
morphemes express some grammatical values, and root morphemes correspond to
concepts from the common part. Morphotactic rules between root and affixal morphemes
specify possible sequences of roots and allomorphs connections, and those between
two affixal morphemes specify possible connections between pairs of corresponding
allomorphs. Fig. 11 shows the conceptual database diagram for the language-specific
part of the model.
The expert has full access to the language-specific part for their language, they can
add, edit, or delete elements of this part. As an example, filling the category "Affixal
morpheme" is presented. Having selected the corresponding element from the left
menu, the expert goes to a page with affixal morphemes list from system database
(Fig. 12), where they can also add a new affixal morpheme.
When adding or editing an existing affixal morpheme, the following attributes are
available for filling in the morpheme data:
1. Textual value of the morpheme;
2. Grammatical value to which the morpheme corresponds;
3. Digital identifier for use in linguistic software.</p>
        <p>Further, the expert can add specific variants of affixes for the morpheme called
allomorphs. Each allomorph contains such attributes as: value, digital identifier, usage
example, example translation, finality flag for morphotactics. The user interface for
allomorphs is implemented in a table form. Figure 3 presented previously shows the
entire form for editing the affixal morpheme. Editing of root morphemes is
implemented in the same way.</p>
        <p>Morphotactic rules, however, have a different editing interface. For the
"Root+Affix" morphotactics, morphonological type database entity is used, which
connects sets of root morphemes with affixal morphemes. The morphonological type,
therefore, determines which allomorphs can be appended to the root morpheme.
Figure 13 shows the interface for "Root+Affix" morphotactic.</p>
        <p>To fill in "Affix+Affix" morphotactics, a tabular representation of adjacency
matrix for pairs of allomorphs is used, where rows contain allomorphs of one affixal
morpheme, and columns contain allomorphs of another. The checkmark in this matrix
means that the allomorph in column can follow the allomorph in row. Figure 14
shows the interface "Affix+Affix" morphotactics.
4.4</p>
      </sec>
      <sec id="sec-4-3">
        <title>Technical information</title>
        <p>The server part of the portal is written in Python using the Django framework. The
choice of language is explained by its simplicity, broad standard library and good
support of packages for machine learning and natural language processing. This
circumstance will make it possible in the future to facilitate the integration of other
linguistic tools with the portal.</p>
        <p>The Django framework, in turn, allows to effectively develop the standard web
solutions, it automates the database interaction, GUI forms designing, and web-requests
processing. The toolkit of this framework is made with taking into account the typical
problems of web service development.</p>
        <p>PostgreSQL was chosen as the database management system. The choice is
justified on the one hand by the openness of this system, the presence of advanced
functionality and high-grade optimization of queries. On the other hand, this DBMS has
support of many external tools that make it possible to implement the requirements
for extensibility and scalability of research software.</p>
        <p>User interface is implemented separately in form of HTML pages with JavaScript
code, generated on the server side using a template engine provided by the Django
framework.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>The Turkic Morpheme Portal is at the stage of linguistic databases population with
help of 30 experts in turkology and linguistics. The general statistics of the database
are presented in Table 1.
At the time of paper preparation, the most complete are the databases for Tatar,
Bashkir, Crimean Tatar, Kazakh, Uzbek, Kyrgyz languages. Table 2 contains detailed
statistical information on databases for these languages.
Database entity
Tatar language
Affixal morphemes
Analytical morphemes
Root morphemes
Bashkir language
Affixal morphemes
Analytical morphemes
Root morphemes
Crimean tatar language
Affixal morphemes
Analytical morphemes
Root morphemes
Kazakh language
Affixal morphemes
Analytical morphemes
Root morphemes
Uzbek language
Affixal morphemes
Analytical morphemes
Root morphemes
Kyrgyz language
Affixal morphemes
Analytical morphemes
Root morphemes</p>
      <p>Count
The presented computer linguistic models and language processing technologies were
considered in relation to Turkic languages, however, due to structural features, they
are applicable to any agglutinative languages. Therefore, the formulated approaches
to development of multipurpose multifunctional software based on the unified
linguistic models are also applicable outside of context of turkology research.</p>
      <p>Further development of the Turkic Morpheme Portal involves the development of
new and integration of existing tools for Turkic languages processing on the basis of
multilanguage model, which expands the possibilities for unification of linguistic
software between supported languages.</p>
      <p>The authors are confident that this portal will be actively used by the researchers of
Turkic languages and developers of language processors, which will contribute to the
creation and application of new common concepts in Turkic languages, especially in
computer science and computer technology.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Guzev</surname>
            ,
            <given-names>V.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pyotrovski</surname>
            <given-names>R.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sherbak</surname>
            <given-names>A.M.:</given-names>
          </string-name>
          <article-title>O sozdanii mashinnogo fonda tyurkskikh yazykov [About creation of machine fund for Turkic languages]</article-title>
          .
          <source>Sovetskaya tyurkologiya [Soviet turkology]</source>
          ,
          <volume>2</volume>
          ,
          <fpage>92</fpage>
          -
          <lpage>101</lpage>
          (
          <year>1988</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient Estimation of Word Representations in Vector Space</article-title>
          .
          <source>In: ICLR: Proceedings of the International Conference on Learning Representations Workshop Track</source>
          ,
          <fpage>1301</fpage>
          -
          <lpage>3781</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>WordNet: an Electronic Lexical Database</article-title>
          . MIT Press, Cambridge, MA (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. «Tugan Tel» Tatar National Corpus, http://tugantel.tatar,
          <source>last accessed</source>
          <year>2020</year>
          /10/15.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Turkish</given-names>
            <surname>National</surname>
          </string-name>
          <article-title>Corpus (TNC)</article-title>
          , https://www.tnc.org.tr,
          <source>last accessed</source>
          <year>2020</year>
          /10/15.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Bashkir poetical corpus, http://web-corpora.net/bashcorpus/search/, last accessed
          <year>2020</year>
          /10/15.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Turklang Electronic Corpora, http://www.turklang.net/en/resources
          <article-title>-for-turkic-languages/</article-title>
          ,
          <source>last accessed</source>
          <year>2020</year>
          /10/15.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8. Global Lexicostatistical Database, http://starling.rinet.ru/new100/mainr.htm,
          <source>last accessed</source>
          <year>2020</year>
          /10/15.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. Maps for Turkic Languages, http://turk.polycorpora.org,
          <source>last accessed</source>
          <year>2020</year>
          /10/15.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Navigli</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S. P.:</given-names>
          </string-name>
          <article-title>BabelNet: The Automatic Construction, Evaluation and Application of a Wide-Coverage Multilingual Semantic Network</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>193</volume>
          ,
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
          . Elsevier (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>SemLink: Linking PropBank, VN and FrameNet</article-title>
          .
          <source>In: Proceedings of the Generative Lexicon Conference</source>
          , GenLex-
          <volume>09</volume>
          ,
          <fpage>13</fpage>
          -
          <lpage>17</lpage>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Plungyan</surname>
            ,
            <given-names>V.A.</given-names>
          </string-name>
          :
          <article-title>Obshchaya morfologiya: Vvedenie v problematiku: Uchebnoe posobie</article-title>
          . [
          <article-title>General morphology: Problem introduction: Schoolbook]</article-title>
          .
          <string-name>
            <surname>Editorial</surname>
            <given-names>URSS</given-names>
          </string-name>
          , Moscow (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Gorodetsky</surname>
          </string-name>
          , B.Y.:
          <article-title>K probleme semanticheskoy tipologii [Onto semantic typology problem]</article-title>
          . Moscow University Press, Moscow (
          <year>1969</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Suleymanov</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          :
          <article-title>Sistemy i informatsionnye tekhnologii obrabotki estestvennoyazykovykh tekstov na osnove pragmaticheski-orientirovannykh lingvisticheskikh modeley [Systems and information tehcnologies of natural language processing on basis of pragmatically-oriented linguistic models]</article-title>
          .
          <source>Doctorate thesis on technical sciences, Kazan</source>
          State University, Kazan (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Suleymanov</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gatiatullin</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Strukturno-funktsionalnaya kompyuternaya model tatarskikh morfem [Structural and functional computer model of Tatar morphemes]</article-title>
          .
          <source>FEN Tatarstan Academy of Sciences, Kazan</source>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>