<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>S. Albota, A. Peleshchyshyn, O. Markovets, V. Vus, A formal approach to modeling the
characteristics of users of social networks regarding information security issues,
Advances in Intelligent Systems and Computing</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.47839/ijc.17.3.1033</article-id>
      <title-group>
        <article-title>Computer linguistic system modelling for Ukrainian language processing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Victoria Vysotska</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Stepan Bandera 12, 79013 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>902</volume>
      <issue>2020</issue>
      <fpage>309</fpage>
      <lpage>320</lpage>
      <abstract>
        <p>The general structure of the сomputer linguistic system (CLS) processing of textual content in the Ukrainian language and the conceptual scheme/model of the functioning of a typical CLS based on modelling the interaction of the main processes and IS components have been developed. Modelling of the main NLP processes of CLS was carried out due to the interaction of the main processes/components of IS and methods of linguistic processing of text content adapted to the Ukrainian language based on grapheme, morphological, lexical, syntactic, semantic, structural, ontological and pragmatic analysis, which allowed to improve the IT of intellectual analysis of the text flow for solving a specific NLP problem. This ensured the adaptation of NLP processes for the analysis of Ukrainian-language textual content. A formal model of a computer linguistic system for processing Ukrainian-language textual content was developed and described, which made it possible to determine the main structural elements and operators of natural language processing at each level of text analysis such as grapheme/phonological, morphological, syntactic, semantic, referential, structural, ontological and pragmatic. Due to the complexity of the morphology of the Ukrainian language, detailed attention is paid to the description of the model of morphological analysis of textual content. Examples of modelling processes for solving typical NLP problems such as CLS identification of viral news headlines and correction of grammatical and stylistic errors are given.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;computer linguistics</kwd>
        <kwd>system</kwd>
        <kwd>NLP</kwd>
        <kwd>Ukrainian language</kwd>
        <kwd>information resource</kwd>
        <kwd>system modelling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Computer linguistic system (CLS) based on NLP methods for text/audio data analysis is
already an integral part of human everyday life [1-5]. On behalf of the user, some CLSs
browse the large volume of Internet information and offer new personalized
mechanisms/techniques/tools for interacting with the computer [6-9], for example through
spam filters of e-mail traffic [10-15], IISS, virtual personal assistants, automatic translation
IS etc. CLS with the support of natural language analysis is at the intersection of
experimental research [16-21] and practical development of usually commercial software
[22-27]. CLS of speech analysis and text analytics interact directly with the user through the
0000-0001-6417-3689 (V. Vysotska)</p>
      <p>© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
support of feedback, which significantly and continuously affects the functioning of the
software and the results of the analysis [28-33].</p>
      <p>The potential for implementing natural language analysis in CLS/IS/modules is
constantly growing exponentially [34-38]. A disproportionately large volume of NLP
applications is usually implemented by large campaigns due to the complexity of the
projects and the need for their commercialization [39-45]. As the opportunities for
implementation in everyday CLS become more widespread, they become less visible,
masking the complexity of their implementation. In parallel, the development of big data
science and computer linguistics, especially based on non-English natural text corpora, has
not yet reached the level necessary for simplification, optimization, and standardization of
the processes of developing appropriate linguistic software [46-51].</p>
      <p>CLSs for solving a large volume of specific NLP problems are just starting to spread and
will eventually automate more processes that are currently solved through additional forms
and selecting/clicking options/buttons. To develop the IT implementation of the
appropriate linguistic software and ensure high reliability of CLS, it is necessary to take into
account modern promising scientific methods of ML, data analysis, big data, and based on
hypothesis analysis [52-55].</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>To support the functioning of a typical CLS and the operation of the main processes when
solving a specific NLP problem, it is necessary and sufficient to implement the main
subsystems as client, server and technological (Fig. 1) [56-63].</p>
      <p>Input content</p>
      <p>Computer linguistic system
Technological
subsystem</p>
      <p>Server
subsystem</p>
      <p>Client
subsystem</p>
      <p>Relevant
content</p>
      <p>User
requests</p>
      <p>The main processes of the functioning of a typical CLS based on the intellectual analysis
of a text stream for the solution of a specific NLP problem [64-72]:
technological processing of incoming content streams:
a. search and recognition of content from relevant sources;
b. accumulation of analyzed content from the source in the cloud;
c. saving information about the location of the found content in the
corresponding source in the database/datastore;
d. preliminary processing of recognized content in the cloud;
e. analysis and marking/classification of recognized content according to
the degree of relevance to the content and purpose of the CLS;
f. integration of content provided that its degree of relevance/relevance is
greater than the threshold value;
g. saving integrated relevant content in the DB;

h. forming an image (descriptive service data) of integrated relevant
content and saving it in the DB;
content management based on text analysis and processing through the server
subsystem (Fig. 2) based on data from the client subsystem and the content support
module:
a. processing of user request streams from the client subsystem to form the
correct IIS expression and subsequent caching of popular content
through the server subsystem;</p>
      <p>Data from
cloud storage
as a result of
integration
Data from the
moderator
Moderator's
rules</p>
      <p>Conservation /
accumulation
Content data</p>
      <p>store
Intellectual and
informational
content search
Knowledge
base
Rule base</p>
      <p>User
requests</p>
      <p>Relevant
content
Searchable
content DB</p>
      <p>Analysis of
requests from
regular users</p>
      <p>CLS server subsystem</p>
      <p>Organization/
support of</p>
      <p>access
Formation of</p>
      <p>answers
Formation and
filling of content
cache</p>
      <p>Relevant
content
User requests</p>
      <p>Relevant
cached data
in cloud storage
Cached data in
cloud storage
storage files for storing relevant content;
knowledge base for IIS/modification/maintenance of this content;
accumulation and appropriate processing of service content.</p>
      <p>The CLS server subsystem is formed from part of the functional components of the
management/support/integration/content modules, in particular [73-81]:
IIS content based on linguistic analysis of requests;
formation and filling of the cache of information blocks of relevant popular content
frequently requested by users and visitors for quick access;
interactive access to relevant CLS profiles/options;
analysis of user requests to accumulate content cache;
storage and accumulation of information blocks in the cloud;
extraction of cached content from the cloud at the request of the user or its
destruction due to the onset of unpopularity;
replenishment/modernization by the moderator of rules and IIS knowledge
base/content analysis as requests of regular users;
analysis of user requests to generate relevant reports.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Models and methods</title>
      <p>The annotated CLS database is the basis of the Website IIS module. Operational and
highquality IIS in the context of current content ensures its high relevance for the CLS user
[8184]. The use of annotated &lt;L in the IIS module helps to implement effective IIS-relevant
content without information noise (Fig. 3). CLS should provide [85-91]:
generation of Webpage according to the template and content of the Website;
preservation and maintenance of cache/filling of Web page/Website according to
the needs of the target audience;
provision of prompt access to the Website for all types of users.</p>
      <p>Mashup-IS   consists in forming a set of integrated content from Internet resources
according to the needs of the target audience and specific user requests for convenient
navigation on the Website/Web page [81-91]:</p>
      <p>
        =&lt;  ,  ,  ,  , ,  &gt;, (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
where  is a set of simultaneously integrated content from Internet resources  ,  is
user requests to Website/Webpage Mashup-IS   ,  is a set of relevant content as a result
of IIS at the request of a user/visitor of Website/Webpage Mashup-IS   ;  is an operator
in the integration of content from Internet resources W and  is a navigation operator in
databases/data/content/filters storage.
      </p>
      <p>Input data
from various
sources</p>
      <p>Data and
rules from the
moderator</p>
      <p>Technological
subsystem</p>
      <p>Content
integration</p>
      <p>Primary
processing
Content
creation
Content
marking</p>
      <p>Server
subsystem
Content
repository</p>
      <p>Access
management</p>
      <p>Content
search
Annotated
database</p>
      <p>The integration of a set of data  from various sources  , including the Website, consists
of combining them according to the appropriate collection of conditions   in one Website
or web page to use different types of content while preserving its main features,
presentation characteristics and the possibility of further processing :</p>
      <p>
        = ( ,  ). (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <p>The integration should provide the user of the Website/Webpage Mashup-IS   to
perceive the integrated content as a single information space using large DS, including the
cloud, and high-quality/operational IIS of relevant content upon request according to the
collection of IIS conditions  :</p>
      <p>
        = ( ,  ,  ). (
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
      </p>
      <p>
        Convenient navigation in the Website/Webpage Mashup-IS   contributes to the
realization of the possibility of supporting the user to search for relevant and relevant
content for him throughout the available IS information space with the greatest
completeness and accuracy with the least expenditure of effort on his part. CLS is a
specialized IS, DSS or multi-agent system for solving a specific NLP problem based on a set
of integrated content from different sources according to the needs of the target audience
and specific user requests for convenient navigation on the Website/Webpage, taking into
account the statistics of the CLS operation, history of actions and personal profiles of users
and history of requests/transitions from IISS. A typical formal CLS model   will be
presented as a tuple:
 
=&lt;  ,  ,  ,  ,  ,  ,  
,   ,  ,  1,  2,  1 ,  2 ,  3 ,  4 ,  ,
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
, 1, 2, 1, 2, 3, 4,  &gt;,
where  is input data to CLS from various sources of information  ;  is source-relevant
content from CLS as a result of IIS according to user/visitor requests;   is a linguistic
content analysis module as a component of the IATCS subsystem   ;   is a module for
generation/modification of the rules of functioning of all modules from the CLS moderator
(for example, rules for updating the cache, integration of content from various sources of
information, linguistic IIS, etc.);  1 is the module for filling the unstructured DB with
integrated content  ;  2 is filling module of structured DB based on developed integrated
content  ;  1 is the module for generating results according to visitors' requests;  2 is
the module for generating results according to user requests;  3 is a cache processing
module for generating reports on popular requests from CLS users;  4 is cache
filling/modification module;   is a module for generating statistical results of the
functioning of CLS/modules and user activity  ;  is operator of generation/modification of
the rules of operation of all modules from the CLS moderator; 1 is the operator of filling
unstructured DB with integrated content  ; 2 is the operator of filling structured DB based
on processed integrated content  ; 1 is the operator for generating results according to
visitors' requests; 2 is the operator for generating results according to CLS user requests;
3 is cache processing operator for generating  reports on popular requests from CLS
users; 4 is CLS cache filling/modification operator with  data;  is the operator for
generating statistical results of CLS/modules functioning and CLS user activity. Fig. 4 shows
the general structure of the proposed typical CLS solution of a specific NLP problem based
on functionality and interaction with clouds [81-93].
      </p>
      <p>Data repository
of processed
content
Integrated
raw content</p>
      <p>Cloud storage
with cached content</p>
      <p>User</p>
      <p>CLS
Website</p>
      <p>Visitor
IATCS subsystem</p>
      <p>Moderator</p>
      <p>A collection of integrated raw content  is contained in a database based on Non-SQL. A
collection of integrated processed  content is contained in a SQL-based
repository/database. From  , the filling  for the cloud is formed based on the statistics 
пof popular requests from users for a certain period. Collection  is a specific DB/DS of
cached current popular content  to optimize the functioning of CLS based on
builtin/modified/additional services in the cloud. These services are the result of the work of
the moderator, who updates the caching rules in CLS and/or Website, updates data in the
SQL database, IIS/IATCS/ management/ integration/ support of integrated and service
content, integration of unprocessed content in Non-SQL database and
collection/accumulation of CLS/Website operation statistics.</p>
      <p>The basis of the IATCS subsystem is the main NLP processes of CLS.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments, results and discussions</title>
      <sec id="sec-4-1">
        <title>4.1. Formal modelling of the main NLP processes of CLS</title>
      </sec>
      <sec id="sec-4-2">
        <title>4.1.1. Formal model of a computer linguistic system for processing Ukrainianlanguage textual content</title>
        <p>
          Natural languages are determined not by rules, but by the context of the application, which
is reconstructed for computer processing. We often identify the meaning of the words used
in combination with other interlocutors. The phrase золота рибка [zolota rybka] (goldfish)
means both a sea creature and a person with a short memory, or a wish-fulfilling
creature/person, so the interlocutors must agree to a common understanding of the
context. Accordingly, speech/language is limited by mentality, region, society and level of
education. Conveying content/meaning is easiest for interlocutors with similar life
experiences, education, place of residence, etc. Therefore, automating the understanding of
speech to solve a specific NLP problem through the appropriate CLS is a rather complex and
painstaking process, especially for synthetic languages, particularly for the Ukrainian
language. The general formal model of CLS is given by the collection:
 
=&lt;  ,  ,  ,  ,  , , , , , , , , ,  &gt;,
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          )
where  is the input text data array;  is a tuple of the original processed text according
to CLS purpose;  is a set of intermediate content that is processed at the corresponding
level in CLS;  is auxiliary dictionaries;  is a set of content processing rules;  is PHA or GA
text operator;  is MA text operator;  is LA content operator;  is SYA content operator; 
is semantic analysis operator;  is operator of ontological content analysis;  ia content
reference analysis operator;  is structural content analysis operator;  is PA content
operator.
        </p>
        <p>Compared to formal languages (subject/thing/object), natural languages are more
universal, but less formalized. We often use one word to describe several meanings (for
example, краб [krab] (a crab) is a sea creature, a dish, a nebula in the constellation of Taurus
and a cockade on a sailor's cap) depending on the content of the dialogue (for example, a
description of the emotions of diving, dinner, a book read, visiting a museum, watching a
historical film, etc. only for the word краб). To store multiple meanings for each word, the
language must be redundant (exceeding the amount of information to transmit/store a
message over its entropy). That is, it is not possible to determine in advance the exact
meaning of the content for each association (every linguistic variable is ambiguous by
default). Lexical and structural ambiguity is a great achievement of natural language, for
example, for generating new ideas, and manifestations of creativity.</p>
        <p>Regardless of the NLP task, the process of processing Ukrainian-language texts in
arbitrary CLS is presented as a sequence of mandatory operators for meaningful structural
analysis of input text content:</p>
        <p>information source  input text  grapheme analysis  (GA) or phonological analysis
(PHA) morphological analysis  (MA)  lexical analysis  (LA) syntactic analysis 
(SYA)  semantic analysis  (SEM)  structured text content</p>
        <p>Additional operators are analyzed such as pragmatic  (extraction of knowledge),
ontological  and referential  (formation of interphrase units). Their application depends
on the complexity and purpose of solving the NLP problem.</p>
        <p>
          The main process of linguistic analysis of textual content is presented:
(
          <xref ref-type="bibr" rid="ref6">6</xref>
          )
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          )
 =  ∘  ∘  ∘  ∘  ∘  ∘  ∘  ∘ ,
 = ( ,  ,  , ( ,  ,  , ( ,  ,  , ( ,  ,  , , ( ,  ,  , ( ,  ,  ,
( ,  ,  , ( ,  ,  , ( ,  ,  ,  ))))))))),
where multiple textual content  = { ,  ,  ,  ,  ,  ,  ,  ,  }, linguistic dictionaries
 = { ,  ,  ,  ,  ,  ,  ,  ,  , } and sets of production/association rules  =
{ ,  ,  ,  ,  ,  ,  ,  ,  }.
        </p>
        <p>The main linguistic process of processing textual Ukrainian-language information to
solve a specific NLP problem consists of nine stages:</p>
        <p>
          Stage 1. Grapheme analysis   of textual Ukrainian-language information  :
  = ( ,  ,  ), (
          <xref ref-type="bibr" rid="ref8">8</xref>
          )
  = 7 ∘ 6 ∘ 5 ∘ 4 ∘ 3 ∘ 2 ∘ 1, (
          <xref ref-type="bibr" rid="ref9">9</xref>
          )
where  is the input text data array;  is the GA operator;   is grapheme structure of
the input text;   is grapheme dictionaries and libraries;   is grapheme analysis rules; 1
is OCR operator; 2 is grapheme parsing operator of the input text  into sections
(information blocks), paragraphs and sentences; 3 is grapheme analysis operator of
linguistic chains into separate words; 4 is the operator for forming a set of unrecognized
chains; 5 is the operator of identification and marking of unrecognized chains as numbers,
dates, constant returns, abbreviations, proper and geographical names, etc.; 6 is the
operator for marking non-text strings as special symbols, formulas, figures, tables, etc.; 7
is an operator for generating a marked linear sequence of words   with official signs and
connections. GA is replaced by PHA in the case of human speech content recognition.
        </p>
        <p>Stage 2. Morphological analysis  of text content   consists in the identification, analysis
and determination of the form and structure of words, in particular:</p>
        <p> = ( ,  ,  ),
  = 3 ∘ 2 ∘ 1, or   = 3 ∘ 4 ∘ 1,
where 1 is the morphological segmentation operator of a graphemically recognized
chain of symbols (words/tokens); 2 is lemmatization operator of lexemes; 3 is POST
operator (marking of parts of speech) for segmented words; 4 is words stemming operator.</p>
        <p>Classical general algorithm of morphological analysis.</p>
        <p>Step 1. Morphological segmentation of the input chain of symbols (replacing GA for short
English-language messages, and supplementing GA for large corpora of English-language
texts, and for Ukrainian-language texts of all kinds - a separate step for marking words in
two sets as immediately identifiable (for example, prepositions) or impossible to identify
(the noun is not in the nominative case).</p>
        <p>Step 2. Lemmatization (reduction to normal form based on dictionary analysis) or
stemming – determination of bases (word forms with endings cut off).</p>
        <p>Step 3. Identification of the grammatical category of each word and the collection of their
corresponding properties in relation to the use in a specific place of the text. (for example,
a collection for a noun: gender, case, person, etc.).</p>
        <p>Step 4. Formation of a linear sequence of morphological structures.</p>
        <p>Stage 3. Lexical analysis of text content С in the intermediate stage of token sequence
analysis for generating a parsing tree at the SYA level:</p>
        <p> = ( ,  ,  ), (12)
 ′ = 2 ∘ 1, ,  ′ = 5 ∘ 4 ∘ 3, or  ′ = 5 ∘ 4. (13)
where 1 is the Speech Segmentation operator for identification/clarification of
words/phrases/tokens after MA or in case of incorrect interpretation during PHA (usually
performed in parallel with PHA and MA in a cyclic process); 2 is speech recognition
operator (SR) or speech-to-text (STT) depending on the content of the NLP task; 3 is Optical
Character Recognition operator (OCR) as the second part after GA and MA for clarifying
incorrect moments of recognition taking into account the recognized neighboring tokens;
4 is word tokenization/segmentation operator as data preparation for building a parsing
tree in SYA; 5 is Text-To-Speech operator (TTS).</p>
        <p>Stage 4. The syntactic analysis  of the text content С consists in building a tree for
parsing the dependencies of words in a sequence of tokens based on their categories:
  = ( ,  ,  ), or   = 3 ∘ 2 ∘ 1, (14)
where 1 is the implementation operator of Grammar induction; 2 is operator of
identification/elimination of boundary ambiguity or sentence violation (Sentence Breaking,
Sentence Boundary Disambiguation); 3 is operator of syntactic Parsing of
phrases/sentences for building a SYA tree.</p>
        <p>Stage 5. Semantic analysis  of textual content С is</p>
        <p> = ( ,  ,  ), or   = 2 ∘ 1, (15)
where 1 is the identification operator of lexical semantics with the generation of a
collection of values of each lexeme of the text; 2 is the relational semantics identification
operator of the interdependencies of the lexeme content of the text.</p>
        <p>A classic general algorithm for semantic analysis.</p>
        <p>Step 1. Lexemes are compared with meaningful dictionary values.</p>
        <p>Step 2. Formation of probabilistic sets for each fragment of text/sentence/phrase with
alternative sems, respectively, for lexemes.</p>
        <p>Step 3. Preliminary interconnection of the content of tokens into a single structure.</p>
        <p>Step 4. Generation of an ordered collection of logical records of superpositions from
semantic classes of lexemes and basic lexical functions.</p>
        <p>Step 5. Finding/marking inaccuracies, contradictions, incorrectness and ambiguity of the
content of the obtained result based on the lexical dictionary.</p>
        <p>Semantic analysis is currently not used in most CLS.</p>
        <p>Stage 6. Referential analysis  formation of interphrase units  .</p>
        <p> = ( ,  ,  ).
(16)</p>
        <p>Referential analysis is often a part of semantic analysis. For Slavic languages, when
analyzing large text corpora, it is best to take it as a separate stage (for example, to analyze
the correspondence of a social group/community in social networks or other dialogues to
identify logical meaningful connections between the posts of different participants due to
the subjectivity of the speech of each. The classic general algorithm of reference analysis:</p>
        <p>Step 1. Contextual analysis of marked fragments of textual content  , for example,
analysis of the pronoun/conjunction що [shcho] (that) analysis depending on the context to
separate the center of unity or local references such as його [yoho] (his), який [yakyy]
(which), цей [tsey] (this).</p>
        <p>Step 2. Actual sentence segmentation of marked fragments of textual content  , for
identification of thematic structures based on themes/rem.</p>
        <p>Step 3. Identification of regular recurrence of context/theme/rheme.</p>
        <p>Step 4. Highlighting the duplicated nomination of lexical units of the text.</p>
        <p>Step 5. Identification of synonymization of lexical units of the text.</p>
        <p>Step 6. Isolation of implications based on situational connections.</p>
        <p>Step 7. Identification of the identity of the reference (for example, the comparison of
lexical units of the text with the object/subject/phenomenon of dialogue/image).</p>
        <p>Stage 7. Structural analysis of  text content С based on the degree of coincidence of
lexical terminological units of the unity of text fragments.</p>
        <p> = ( ,  ,  ), or   = ( ,  ,  ).
(17)</p>
        <p>Similarly, referential analysis is often part of SEM for short texts/messages, or not used
at all. For large corpora of texts as an additional stage of elimination of marked inaccuracy
in SEM. Classical general algorithm of structural analysis.</p>
        <p>Step 1. Formation/replenishment of the basic set of rhetorical relations of interphrase
units based on the results of reference analysis and/or SEM.</p>
        <p>Step 2. Generation of a non-linear network/graph of interphrase units.</p>
        <p>Stage 8. Ontological analysis of  textual content С based on SEM results and
reference/structural analyzes if necessary:
С = ( ,  ,  ), С = ( ,  ,  ) or С = ( ,  ,  ).
(18)</p>
        <p>Stage 9. Pragmatic analysis of  text content С is used to determine the structure of the
text taking into account the context of sentences when forming paragraphs, sections and
dialogues. PA is an essential addition to SEM, reference and structural analyzes if they did
not contribute to the elimination of marked inaccuracy. In some cases, it is sufficient to
apply PA immediately after SEM. It is also an indispensable stage of data preparation for
extracting knowledge from text corpora.</p>
        <p>= ( ,  ,  ,  , [С,  ,  ], ), or  = 2 ∘ 1, (19)
where 1 is the semantics identification operator outside individual sentences/phrases;
2 is text processing operator through higher-level NLP applications, for example, to
simulate intelligent behavior and apparent understanding of natural language.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.1.2. Models of grapheme and phonological text analysis in Ukrainian</title>
        <p>Depending on the tricky NLP task, grapheme (text analytics) or phonological (speech
recognition) analysis of text content is used. PHA consists in the study of the structure,
organization, and interpretation of speech sounds  of a specific natural language (Table 1)
based on phonemic, phonetic, and prosodic rules of   and dictionaries of   analogs. GA is
a recursive parsing of the text  , taking into account the linguistic features of graphemes of
various languages (including non-natural ones, for example, mathematical, programming,
pseudo-languages, etc.) based on the rules for recognizing strings of a certain language  
and dictionaries   of reference grapheme models, in particular:
 ′ = ( ,  ,  ,  ),  ′ ⊇  .
(20)</p>
        <p>GA can be an OCR part - encoding/recognizing text on an image into a string of characters
for e-submission.</p>
        <p>Depending on the tasks, there are the following methods for forming   PHA rules:






</p>
        <p>General phonology ′1 (rules of phoneme organization in different languages).
Descriptive phonology ′2 (identification of the phoneme of a language or dialect).
Historical phonology ′3 (changes in phonemes, language structure during the
period).</p>
        <p>Segmental phonology ′4 (analysis of phonemes, syllables, phonetic words,
syntagms and phrases).</p>
        <p>Supersegmental phonology ′5 (analysis of intonation, tone, stress, rhythm, tempo
and pauses).</p>
        <p>Phonetic analysis ′6 (analysis of the sound structure of the language).</p>
        <p>Phonemic analysis ′7 of the smallest unit of the phonological level.</p>
        <p>GA is an elementary text parsing (Fig. 5) taking into account the features of graphs of
different languages and the use of special symbols, objects and marks. A grapheme is an
atomic meaningful linguistic (grapheme) unit of a written text (sign, symbol, special symbol,
object as a picture, etc.). The purpose of GA is to form a model of the grapheme structure of
the input text and generate grapheme rules (regular expressions) for the
identification/classification of grapheme units in the sequence of character
strings/graphemes and the connections between them. The purpose of the first level of GA
– grapheme identification – is to label meaningfully independent sequences of symbols,
tokens in these sequences and to determine the main language of the input text
content/fragments based on a priori grapheme standards (Fig. 5). The tuple of reference
grapheme models is best described by a formal grammar (abbreviations of the criteria are
given in Table 2). In parallel with the parsing/identification of graphemes, they are
classified/marked according to established rules (Table 3).</p>
        <p>Decoding
Official symbol</p>
        <p>Brackets
Mathematical symbol</p>
        <p>Capital letter</p>
        <p>Small letter
Latin capital letter</p>
        <p>Latin small letter
Cyrillic capital letter
Cyrillic small letter
English capital letter
English small letter
German capital letter
German small letter</p>
        <p>Polish letter</p>
        <p>Polish small letter
Ukrainian capital letter</p>
        <p>String classifier</p>
        <p>Sign classifier</p>
        <p>Names dictionary
Geography names dictionary
Abbreviations dictionary</p>
        <p>Software modules
Grapheme analysis as
grapheme parsing
Lexical analysis
Syntactic analysis
Semantic analysis
Reference analysis
Ontological analysis</p>
        <p>Pragmatic analysis
Morphological analysis</p>
        <p>Output data
Highlighted text</p>
        <p>Fragments
Lexemes:
- language,
- speechless,
- conditional language
Grapheme structure</p>
        <p>Fragments
Sentence
Syntagms
Lexemes
Relation
Let's consider the classical Chomsky 
with the 
ℎ

=&lt; 

ℎ
,</p>
        <p>, 
ℎ
=&lt;  , 
&gt;,
,</p>
        <sec id="sec-4-3-1">
          <title>Position 1</title>
        </sec>
        <sec id="sec-4-3-2">
          <title>Position N</title>
        </sec>
        <sec id="sec-4-3-3">
          <title>All positions – Symbol Symbol</title>
          <p>Space
Space
&gt;,</p>
          <p>–
Symbol
Space
Symbol
Space
and 
:
Space
–
–
–
–
(21)
(22)</p>
          <p>Table 4 provides a list of grapheme classification/marking rules for reference models
(Table 2) according to production rules:
The set of russian capital letters;
A set of russian lowercase letters;
Latin capital consonant letters;</p>
          <p>Latin capital vowels;
Terminal Latin small consonant letters;</p>
          <p>Terminal Latin small vowel letters;
The set of terminal Cyrillic capital consonants;</p>
          <p>Cyrillic capital vowels;</p>
          <p>Cyrillic small consonant letters;
Terminal Cyrillic lowercase vowels.
production rules are used to identify, classify and mark meaningful grapheme
units of analysis of the input text of content  (words, abbreviations, stable phrases as
idioms and metaphors, sentence boundaries and quotations/sarcasms by punctuation,
emoticons, geographical and proper names, abbreviations, words with apostrophes, etc.) at
the parsing stage, taking into account the language of text fragments. Requirements for
identifying a grapheme unit in a sequence of symbols for further morphological analysis of
words:
1) the character sequence is easily identified and classified;
2) sequence too large to identify value by dictionary;
3) the sequence is too small to identify many values;
4) the number of grapheme units is too large to split the sample.</p>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>4.1.3. Morphological analysis of the Ukrainian language</title>
        <p>MA consists in identifying, analyzing and determining the form and structure of words in a
natural language text using
methods such as</p>
        <p>Morphological Segmentation  ,
1
Lemmatization  , POST  (marking of parts of speech) and Stemming  (Table 5), in
2 3 4
particular:
 ′ = ( ,  ,  ),
(23)
where  ′ ⊆  ,  ′ =  ∘  ∘  (enough for English-language short texts of a certain
3 2 1
topic) or  ′ =  ∘  ∘  (for most cases of messages of various topics).</p>
        <p>3 4 1
Classification of natural language lexeme stemming algorithms</p>
        <p>Name
Word={приватизаційний} 
Stemming={приватизац};
Word={цивілізаційний} 
Stemming={цивілізац};
Word={інформаційний} 
Stemming={інформац};
…………………………
Word={проголошую,
наголошувати, виголошував}
 Stemming={голошу}.</p>
        <p>Stemming={інформац} 
Word={інформаційний,
інформаційна, інформаційне,
інформаційним,
інформаційними,
інформаційних, інформаційні,
інформаційній,
інформаційнім,
інформаційного,
безпритульної,
інформаційному,
інформаційною,
інформаційну}
KnowledgeBase={чорн,
чорняв}  Word={чорнява}
 Count={4, 6}
Stemming={чорнява}.</p>
        <p>The algorithm will choose the
longer option.</p>
        <p>If English stemming is a simple
task, then Ukrainian stemming
is several levels more difficult.</p>
        <p>A detailed description of the
non-commercial stemming
algorithm for Ukrainian is a
matter of time.</p>
        <p>Word={особистість}
Stemming={особист}
End={ість};
Word={спогади}
Stemming={спогад}
End={и}; Word={дивними}
Stemming={дивн}
A hybrid
approach
part of speech is
determined without
taking into account
the context in
which this word
was used in the
sentence.</p>
        <p>A combination of
the above
algorithms is used.
likely part of speech for
that word is preferred.</p>
        <p>The probability of
stemming errors increases
with an incorrect
description of the rules
and formation of the table
of endings
the learning base, the
better the result of
their work. The
knowledge base for
these algorithms is a
set of logical rules
and IIS tables.</p>
        <p>The table does not
contain all word
forms, but
exceptions to the
rules that are
incorrectly
processed by the
clipping algorithm.
End={ими}, where End – the
result of learning the algorithm,
i.e Word(кияни)  {End(ість)
= FALSE, End(и) = TRUE,
End(ими) = FALSE} Cut (и)
or Word(чуйними)  {End
(ість) = FALSE, End (и) =
TRUE, End (ими) = TRUE}
Cut (и) OR Cut (ими).</p>
        <p>For example, the algorithm can
use the method of cutting off
endings and suffixes, but at the
first stage, perform IIS on the
table.</p>
        <p>The best way for Slavic languages:
(24)
1
Lematization  is the transformation of a word form into a lemma (normal dictionary
2
form). Usually, during the transformation, a dictionary is used to present words in their
actual form. Otherwise, remove only inflections and return to the lemma.</p>
        <p>Morphological segmentation  is the division of words into separate morphemes to
identify their class (Table 6-8). The complexity is directly proportional to the complexity of
the morphology (word structure) of a specific natural language (Table 9-10).
Linguistic characteristics of some classes of verb stem morphemes</p>
        <p>Verb
фарбувати(ся)
усміхнутися
стогнати
спитати(ся)
сміятися
розфарбувати(ся)
привести(ся)
поділити(ся)
побудувати(ся)
нести(ся)
молоти(ся)
малювати(ся)
любити(ся)
кохати(ся)
змарніти
запізнюватися →
запізнитися
досліджувати(ся) →
дослідити(ся)
втручатися
втратити → втрачати
вести(ся)</p>
        <p>Analysis
фарб-ува-ти(-ся)
усміх-ну-ти-ся
стогн-а-ти
спит-а-ти(-ся)
сміj-а-ти-ся
розфарб-ува-ти(-ся)
привес-ти(-ся)
поділ-и-ти(-ся)
побуд-ува-ти(-ся)
нес-ти(-ся)
мол-о-ти(-ся)
мал-юва-ти(-ся)
люб-и-ти(-ся)
кох-а-ти(-ся)
змарн-і-ти
запізн-юва-ти-ся →
запізн-и-ти-ся
дослідж-ува-ти(-ся)
→
дослід-и-ти(-ся)
втруч-а-ти-ся
втрат-и-ти →
втрач-а-ти
вес-ти(-ся)</p>
        <p>Basics
, у, ся − ся)
, н, ся)
фарб-( ,  ̄,  , 
усміх-( ̄,  ,  , 
1. If there is a base with the sign  – TE -и(і,ї)-.
2. If there is a base with the sign  – TE -а-/-я-.
3. In the presence of a base with the sign of  and a suffix of a participle beginning with a
consonant - either TE, or a suffix for the formation of perfect and imperfect forms of verbs
mainly of foreign origin.
4. If there is an infinitive base with the sign -а (-я), -ува- (-юва-), -овува- – the suffix -н- (-ий,
а, -е, -і), is added to it, for example, посія-(ти) – посіяний, чита-(ти) – читаний,
розпиля(ти) – розпиляний, писа-(ти) – писаний, зігна-(ти) – зігнаний; загоювати – загоюваний,
оспівувати – оспівуваний, застосовувати – застосовуваний; the suffix -ува- (-юва-), if
the stress moves to the first vowel, it changes to -ова-, for example, роздрукува(ти) –
роздрукований, сформулюва(ти) – сформульований, реконструюва(ти) –
реконструйований, запрограмува(ти) – запрограмований.</p>
        <p>Morphological
and phonological
rules
A
1
2
B
1
2
3
4
5
6
7
8
9
10
11
12
C
1
2
D
1
2
3
E
1
2</p>
        <p>Rule</p>
        <p>Example</p>
        <p>The basic rules of alternation of consonants in personal forms
Conjugation (declension) I - consonants засвистати – засвищу, хотіти – хочу, чесати – чешу, колихати –
change at the end of the stem, if there is an колишу, мазати – мажу, могти – можу, полоскати – полощу, пекти –
alternation in the 1st person singular – г- печу – печений;
ж, з-ж, к-ч, х-ш, с-ш, т-ч, ст-щ, ск-щ
Conjugation (declension) II - we have
sound changes only in the 1st person
singular – д-дж, т-ч, з-ж, с-ш, зд-ждж,
ст-щ
where  1 and  2 are arbitrary vowels.</p>
        <p>ІХ.  :  + и → і,
where  is sound designation [j] (yot).</p>
        <p>Х.  .  :  +  → я,  → я.
Х.  .  :  + у → ю,  у → ю
Х.  .  :  + е → є,  е → є
Х.  .  : Х′ +  → Х + я,
Х.  .  : Х′ + у → Х + ю,
……………………………………………………………………………..</p>
        <p>ІХ.  : о +  ( ,  ) +  →  +  ( ,  ) +  , (96)
where Z is arbitrary sequence no longer than 3 characters (alternation o/a in the base of
the type скочити/скакати [skochyty/skakaty] (jump), ломити/ламати
[lomyty/lamaty] (break), кроїти/краяти [kroyity/krayaty] (cut), клонити/кланятися
[klonyty/klanyatysya] (bow), котити/катати [kotyty/kataty] (roll), схопити/хапати
[skhopyty/khapaty] (grab), гонити/ганяти [honyty/hanyaty] (chase),
допомогти/допомагати [dopomohty/dopomahaty] (help)); group of consonants after
that alternating symbol -о- (that is, what separates it from TE -а-/-я- before -у(ю)ва- ),
cannot contain more than 3 letters.</p>
        <p>ІХ.  .  : с′ + 
ІХ.  .  : в′ + 
ІХ.  .  : б′ + 
ІХ.  .  : д′ + 
→ ш +  ,
→ вл′ +  ,
→ бл′ +  ,
→ дж′ +  ,
(111)
Х.  .  : Х′ + е → Х + є, (112)</p>
        <p>XI. Rules for erasing the indicator of the boundary between verb morphemes in text
content in the Ukrainian language.</p>
        <p>+  →  , (113)
where  and  are any morphemes that none of the rules of groups IX-X apply to  +  .
Such a restriction on  and  prevents untimely destruction of the boundary between
morphemes before the application of the corresponding morphological rules. If any
morphonological rule can be applied, then it must be applied to prevent the formation of
nonsense words, for example, *котаючий [*kotayuchyy] (*rolling) from котити [kotyty]
(rolling) or *качений [*kachenyy] (*rolling) from катати [kataty] (rolling). Main
properties [404, 882]:</p>
        <p>1. Transitivity of derivability. If there is a sequence  0,  1, . . . ,   , in which each i-th
chain is directly derived from i-1 according to the hypothetical syllogism ( →  ,  →  ├ →
 ),   is derived from  0, and the sequence  0,  1, . . . ,   is the derivation of   from  0
[882].</p>
        <p>2. Direct derivability. If there are 2 sequences  and  :
 =  1  2,  =  1  2,
(114)
where  1 and/or  2 can be empty and the grammar  has the rule  →  , the  is
directly derived from  , for example, from the sequence according to rule VI.3О( ,  ̄,  ,  , ∅,
ся − ся)а + С( , , , ̄) + Ф the sequence is directly derived О( ,  ̄,  ,  , ∅, ся − ся)а +
юч + Ф.</p>
        <p>XII. Rules for marking a lexeme as an adjective with multiple features in the
corresponding sentence/phrase (gender, tense, case, etc.) in the Ukrainian text.
(XІI) А′х,у, → лийх, ,</p>
        <p>This chain cannot be further generated for the Ukrainian language.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.1.4. Lexical analysis of the Ukrainian language</title>
        <p>LA is preliminary processing of text or speech, transformation of a chain of symbols into a
sequence of tokens (tuples of symbols according to appropriate patterns):
where  ′ ⊆  ,  ′ = 2 ∘ 1 (enough to transform the sound series of speech into printed
text) or  ′ = 5 ∘ 4 ∘ 3 (to transform scanned text into speech), but is sufficient for most
cases:
 ′ = ( ,  ,  ,  ),
 ′ = 5 ∘ 4.
(116)
(117)
1. Speech Segmentation 1 is the division of the sound stream of human speech into
separate words. In human conversation or speech, pauses between consecutive
words are almost unidentifiable, so segmentation is a necessary task for speech
recognition. In most spoken languages, sounds as consecutive letters are combined
with each other in the process of coarticulation, so converting an analog signal into
discrete symbols is a rather complex NLP process in CLS systems for technical
implementation.
2. Speech Recognition 2 (SR) or Speech-To-Text (STT) is transformation of the speech
signal into e-text. Different people pronounce words in each language with different
stresses, speeds, and intonations. CLS should recognize a wide range of input
unstructured data as identical to a single equivalent.
3. Optical Character Recognition 3 (OCR) is translation of scanned handwritten or
printed text after GA and MA into a sequence of codes for e-submission with
correction of simple errors based on statistics and theory probabilities of using a
sequence of N-grams/words/endings instead of obscure symbols.
4. Tokenization or Word Segment 4 is the demarcation and categorization of sections
of the chain of input symbols for MA. For English or Ukrainian languages, this is a
fairly trivial situation, since words are usually separated by spaces (problems only
if there are stylistic and grammatical errors in the text). However, some written
languages such as Chinese, Korean, Japanese, and Thai do not mark word boundaries
in this way. Then tokenization is an important task based on vocabulary and word
morphology. Sometimes the method is used to form a Bag-of-words model (BOW) in
Data Mining.
5. Text-To-Speech 5 (TTS) is transformation of handwritten, typewritten or printed
text into a speech signal for oral presentation, for example, for people with visual
impairments.</p>
      </sec>
      <sec id="sec-4-6">
        <title>4.1.5. Syntactic analysis of the Ukrainian language</title>
        <p>SYA is the basis of semantic analysis:
 ′ = ( ,  ,  ),
(118)
where  ′ ⊆  ,  ′ = 3 ∘ 2 ∘ 1:
1. Grammar induction 1 is generation of formal grammar to describe language syntax.
2. Sentence Breaking or Sentence Boundary Disambiguation 2 is analysis of the
presence/absence of appropriate punctuation marks and the text between them (a
dot not only marks the end of a sentence, but also a contraction).
3. Parsing 3 is the generation of sentences from the input sequence of symbols of the
parsing tree (grammatical analysis) for the analysis of the grammatical structure
according to the given formal grammar (Table 16). There are hundreds or thousands
of analyzes for a typical sentence, most of which are absolutely meaningless to a
native speaker. There are two main types of parsing: Dependency Parsing 13 and
Constituency Parsing 32, or in the form of some combination of these methods 33.
Dependency parsing focuses on the relationship between words in a sentence
(primary objects and predicates), and constituent parsing focuses on building a SYA
tree using probabilistic context-free (stochastic) grammar (PCFG).</p>
        <p>For example, when parsing/generating sentences/phrases, the choice of case inflection
of a specific word in the Ukrainian language directly depends on the type of base and part
of speech, in particular, for noun groups taking into account the context (Table 9-10, Table
16):</p>
        <p> 1  1 1 2,   2  2 3 4,</p>
        <p>A set of linguistic units  of one type, together with a set of linguistic units  1 of another
type, are transformed in a different way  1 2, than together with a set of linguistic units  2
of the third type –  3 4. Without taking into account the context, a more fractional
classification should be introduced:</p>
        <p> 1 1 2,   2 3 4,</p>
        <p>Verb group</p>
        <p>Word
replacement
3)  1 ̃ , , ,   2 →  1 з,ай,м,  2,  1    , , ,  2   ̃ ′; 4)  ̃ , , ,3 →   , , .
1)  ̃ ,тепер, →   ,тепер,   ̃ ′, ′,зн, ′  ̃ ″, ″,ор, ″; 2)  ̃ ,тепер, →
  ,тепер,   ̃ ′, ′,ор, ′  ̃ ″, ″,зн, ″;</p>
        <p>Each record is a rules set, for example, II.1 forms 648 rules:  ̃ч,од,н,3 →  ̃ч,од,н,3  ̃ч,од,р,1;
 ̃ч,од,р,3 →  ̃ч,од,р,3  ̃ч,од,р,1; ...;  ̃сер,мн,м,3 →  ̃сер,мн,м,3  ̃сер,мн,род,3 (Table 17). To generate a
sentence tree in Ukrainian, the inflections agreement rules corresponding are used (Fig.
6Fig. 7).</p>
        <p>Each step is an expansion of a symbol of the sequence (for example,  ̃ од,тп,3 – Rод,тп,3
Šч,од,з,1 Šс,од,о,3) or replacement (so, Šч,од,з,1 –  чз,аойдм,з,1). For such deployment, it is necessary
to form more detailed types of linguistic units to take into account the location in the context
of the sentence, for example:
 мн,род →    мн,род,
(121)
, , (4.1)
where  is the word form of a sentence,   is the base of a word of type 
( = 1,2,3, . . . ),  мн,род is the inflection of the genitive plural in Ukrainian languages, for
example:
2.(I) # Sж,од,н,3 Vод,тепер,3 #
3.(III.1) # Sж,oд,н,3 Vод,тепер,3 Sч,од,зн,1 Sсер,од,ор,3 #
4.(II.1) # Sж,од,н,3 Sч,од,р,3 Vод,тепер,3 Sч,од,зн,1 Sсер,од,ор,3 #
5.(II.2) # Аж,од,н Sж,од,н,3 Sч,од,р,3 Vод,тепер,3 Sч,од,зн,1 Sсер,од,ор,3 #
6.(II.2) # Аж,од,н Sж,од,н,3 Ач,од,р Sч,од,р,3 Vод,тепер,3 Sч,од,зн,1 Sсер,од,ор,3 #
7.(II.2) # Аж,од,н Sж,од,н,3 Ач,од,р Sч,од,р,3 Vод,тепер,3 Sч,од,зн,1 Асер,од,ор Sсер,од,ор,3 #
8-9 ..............................................................................................................................
10.(II.4) # Аж,од,н Sж,од,н Ач,од,р Sч,од,р Vод,тепер,3 Sч,од,зн,1 Асер,од,ор Sсер,од,ор #
11.(II.3) # Аж,од,н Sж,од,н Ач,од,р Sч,од,р Vод,тепер,3 Sчз,аойдм,зн,1 Асер,од,ор Sсер,од,ор #
12-18 ............................................................................................................................</p>
        <p>IV.1 IV.7 IV.4 IV.6 IV.3
сина наповнює мене безмежним щастям</p>
        <p>IV.6 IV.2 IV.6
# весела посмiшка твого
#
 1 мн,род →   ів(друз − ів),  м1н,род → ів,
 2 мн,род →   ок(іграш − ок),  м2н,род → ок,
 3 мн,род →   ей(діт − ей),  м3н,род → ей,
the Ukrainian language on the presence/absence of negation:</p>
        <p>̃ →    ̃ ,  ̃ →    ̃1,
 ̃ → ¬   ̃ ,  ̃ → ¬   ̃2.</p>
        <p>≠¬
where  ̃ is a verb group in a sentence,   is a transitive verb in a verb group,  ̃ is a direct
object,  ̃ is a noun group, ¬ is a negation, in particular, for sentences, школяр пише есе
[shkolyar pyshe ese] (the student writes an essay) and школяр не пише есе [shkolyar ne
pyshe ese] (the student does not write an essay) with relevant conclusions:
The use of the instrumental subject  ̃</p>
        <p>with a verbal noun depends on the presence of
the object  ̃ (аналіз змісту системою [analiz zmistu systemoyu] - content analysis by
 ≠¬
the system):
 4 мн,род →   их(знайом − их),  м4н,род → их,
 5 мн,род →   (машин −),  м5н,род →  .</p>
        <p>…………………………………………………………………
The choice of the case of the direct complement in a word form or sentence depends in
(126)
(127)
(128)
(129)
(130)
(131)
(132)
 ̃ →  ̃′ ̃</p>
        <p>̃ ,  ̃ →  ̃′ ̃  ̃  1,</p>
        <p>It is impossible to completely abandon the semantics of the context during the correct
identification and correction of grammatical and stylistic errors, it is necessary to take into
account not only one symbol in the left part of the rules (30)-(40), which ensures the
permutation of symbols.</p>
        <p>Semantic features are investigated using a set of lexical and linguistic resources in the
form of  dictionaries and libraries,   dictionary management tools, semantic role marking
23 and the process of word embeddings 3. Word embedding consists in mapping words,
phrases or phrases from the dictionary  into vectors of real numbers   for ease of
processing, for example, based on Word2Vec.</p>
        <p>= 23(3( ,   ),   ).</p>
        <p>Taking semantics into account significantly simplifies the grammatical tree (Fig. 8). If we
connect the symbols (ancestors) directly to the final results (descendants) during expansion,
replacement or rewriting, we get a tree of components, or a syntactic structure. The
specified rules (Table 16) are capable of generating other phrases that are not necessarily
meaningful, as rules II.1 and II.2 are cyclical. Along with the sequence as весела посмішка
[vesela posmishka] (funny smile), you can get весела весела посмішка [vesela vesela
posmishka] (funny funny smile), etc. The number of phrases in natural language must be
finite. With the correct construction of the sentence parsing tree (Fig. 6-Fig. 8) and further
shortening (Fig. 9), it is possible to match cases.
Аж,од,н Sж,од,н,3</p>
        <p>Ач,од,р
Sж,од,н</p>
        <p>Sч,од,р,3
Sч,од,р
Vод,тепер,3 Sч,од,зн,1</p>
        <p>Sсер,од,ор,3
Асер,од,ор</p>
        <p>Sсер,од,ор,3
Sчз,аойдм,зн,1</p>
        <p>Sсер,од,ор</p>
        <p>сина наповнює мене безмежним щастям
безмежним Sсер,од,ор
(133)
весела</p>
        <p>посмiшка наповнює безмежним щастям</p>
        <p>The agreement of cases between the linguistic units of the sentence affects the
subsequent semantic analysis of the text. For example, in the Ukrainian language it is
possible to generate sequences of linguistic units of the type  1 2
 2 2 3  2′
 1′
 2′
 3′,  1 3 2 1  1′
 3′
 2′ 1′ etc. (or as   ′): Саша, Софія, Катя, Данило, … –
а
спортсмен, співачка, художниця,поет, … відповідно. In particular, 
sequence of proper names;  ′  ( ′ ′ ′ ′. . . ) is a sequence of professions agreed with
proper names in the family;  is a punctuation mark.
с
. . . ) is a
1.   →   
 ′ ,</p>
        <p>2.   ′  →    ′ ,  ,  = 1,2,3,</p>
        <p>3.   
4.  
 →    ,
→  .</p>
        <p>}
where  ,  ′,  are basic linguistic units;  ,   are auxiliary linguistic units;  is the initial
symbol as an indicator of the chain generation type. For a more effective study of errors,
meta-data analysis of linguistic features and characteristics of the original text is used, in
particular, content genre, presence/absence of dialect, slang, terminology, and the
probability of writing by a native speaker or the result of a translation.</p>
      </sec>
      <sec id="sec-4-7">
        <title>4.1.6. Semantic analysis of the Ukrainian language</title>
        <p>Semantic analysis is currently not used in most CLS, but with the gradual introduction of AI
into the everyday life of the average person, this task should be solved and simplified (Fig.
10). The more complex the grammar of the language, the more difficult it is to conduct CEM:
 ′ = ( ,  ,  ),
where  ′ ⊆  ,  ′ = 2 ∘ 1.</p>
        <p>Linguistic semantics 1 (individual words in context) consists of:
 ′′ = 16 ∘ 15 ∘ 14 ∘ 13 ∘ 12 ∘ 1,  ′′ ⊆  ′.</p>
        <p>1
Semantic analysis
(134)
(135)
Linguistic semantics
Лексична семантика</p>
        <p>Lexical semantics
Recognition of named</p>
        <p>entities</p>
        <p>Sentiment analysis
Extracting terminology
Ambiguity of the word's
meaning</p>
        <p>Categorization and
analysis of words</p>
        <p>Template
Common and</p>
        <p>distinctive
characteristics</p>
        <p>Relational semantics</p>
        <p>Removal of relations</p>
        <p>Semantic parsing
Semantic marking of
roles
coherence (integrity) of the text (Coherence), Anaphora Resolution and final
conclusion (Eng. Inference).
4. Named-Entity Recognition 14 (NER) or identification of an object entity,
fragmentation of an object entity, extraction of an object entity is extraction from
unstructured text of information about the presence of certain named entities of the
corresponding categories such as dates, percentages, monetary values, quantities,
time, geographical locations, names of organizations or people, etc. Having a capital
letter at the beginning of a word does not solve this problem - the beginning of a
sentence or a line of poetry also begins with a capital letter. In German, all nouns
begin with a capital letter. In addition, named entities often include several words,
only some of which are capitalized. French, Ukrainian, and Spanish do not capitalize
adjective names. In German, all nouns are capitalized. Chinese, Korean, Japanese, or
Arabic have no capital letters at all.
5. Terminology Mining 15 (Terminology Extraction, Term Recognition, Glossary
Extraction, Term Extraction) is the automatic extraction of relevant terms from the
relevant corpus. One of the first steps towards knowledge domain modeling is the
collection of a vocabulary of domain terms as linguistic features of text content.
Approaches to term extraction use linguistic processors (marking parts of speech,
fragmentation of parts of texts) to extract terminological candidates, that is,
syntactically plausible terminological noun phrases or keywords, for example for
rubrication.
6. Tonality analysis of the text or multimodal sentiment analysis 16 (Opinion Mining,
Sentiment Analysis) is the detection of a subjective emotionally-colored set of
content (positive, negative or neutral) in a text array data of a specific author in
relation to the corresponding object, subject, event or phenomenon of thematic
subject area. It is useful for identifying trends of public opinion for online marketing,
in social networks, forming political propaganda, etc.</p>
        <p>Relational semantics 2 (semantics of individual sentences) consists of:
 ′′′ = 2 ∘ 22 ∘ 2,  ′′′ ⊆  ′.</p>
        <p>3 1
(136)
1. Relationship Extraction 12 is identification of relationships of nominal entities (for
example, family ties, colleagues, enemies, etc.).
2. Semantic Parsing 22 is presentation of the semantics of a part of the text (usually a
sentence) in the form of a logical formalism (DRT parsing, Discourse Representation
Theory) or a graph (AMR parsing, Abstract Meaning Representation), for example:
 ,  ,  :  ( , хотіти01) ( , гуляти01) ( , дитина)</p>
        <p>g0(a, b)g1(a, e)g0(e, b).</p>
        <p>AMR format: (a / хотіти-01 : arg0 (b / дитина) : arg1 (e / гуляти-01 : arg0 b)).
(137)
3. Semantic Role Labelling 23 (Implicit Semantic Role Labeling Below) is assignment of
semantic role labels to words or word combinations in a sentence, for example, the
role of goal, agent, or result according to the algorithm:
 ′′′′= 325 ∘ 234 ∘ 233 ∘ 232 ∘ 231,  ′′′′ ⊆  ′′′.
(138)
a. Selection of a sentence or a fragment of text 231.
b. Definition of semantic predicates in a sentence232 (verb and noun frames).
c. Disjunction of defined semantic predicates 233.
d. Identification of the elements of the defined frame 234.
e. Classification of the identified elements of the frame 235– assigning a
semantic role in the sentence.</p>
      </sec>
      <sec id="sec-4-8">
        <title>4.1.7. Pragmatic analysis of the Ukrainian language</title>
        <p>PA is used to determine the structure of the text taking into account the context of sentences
when forming paragraphs, sections and dialogues. The main task is the identification of the
context of such linguistic units as sentences and the formation of the semantic relationship
of these linguistic elements. PA is a mandatory component in AIS for the interpretation of
human dialogue, speech and its corresponding analysis (Fig. 11).</p>
        <p>= ( ,  ,  ),  = 2 ∘ 1.</p>
        <p>Pragmatic analysis</p>
        <p>Discourse
Coreference resolution</p>
        <p>Discursive analysis
Non-semantic marking of</p>
        <p>roles
Identification of text</p>
        <p>indentation</p>
        <p>Arguments mining
Thematic segmentation and
recognition</p>
        <p>Correction of
grammatical errors
Dialogue management</p>
        <p>Natural language</p>
        <p>generation
AI document</p>
        <p>High-level NLP applications</p>
        <p>Summarizing the text</p>
        <p>Book generation
Understanding natural</p>
        <p>language
Machine translation
Question-and-answer
systems
(139)</p>
        <p>Discourse 1 (semantics outside individual sentences) consists of:
 ′ = 6 ∘ 15 ∘ 14 ∘ 13 ∘ 12 ∘ 1,  ′ ⊆  .</p>
        <p>1 1
(140)
1. Coreference Resolution 11 is identification in the next text of the following
wordsreferences or expressions about objects, subjects, phenomena and events, which are</p>
        <p>Natural language processing through higher-level NLP-applications 2 simulate
intelligent behavior and obvious understanding of natural language and are currently
generally divided into the following classes:
7. Text Summarization 12 (Automatic Summarization) is generation of a readable
digest/annotation as a summary in the form of a text fragment from the general
analyzed text (for example, a scientific article, publication in a newspaper or
magazine).
8. Grammatical Error Correction 22 is identification and correction of
grammatical/stylistic errors at all levels of linguistic analysis
(phonology/orthography, morphology, vocabulary, syntax, semantics, pragmatics).
9. Machine Translation 32 is automatic translation of text from one human language to
another using all levels of linguistic analysis, especially grammar, semantics and
facts about the real world, etc., based on the solution of AI-complete (AI-hard) tasks
of abstract content translation.
10. Dialogue Management 42 (Dialogue System, Conversational Agent, CA) is the
organization of communication between a program and a person, other than
chatbots, using text, speech, graphics, tactility, gestures, etc. for two-way
communication.
11. Question Answering systems 52 is determination of the correct answer based on the
analysis of a typical human question (for example, "What is the capital of Ukraine?")
and an open/complex question (for example, "In what the meaning of existence?").
12. Natural Language Generation 62 (NLG) is conversion of content from
databases/databases or semantic intentions into readable specific human language
through an algorithm:
a. Determination of the content of 621 (what information to mention in the
text).
b. Document structuring 622 (content transfer template).
c. Aggregation 623 (combining similar sentences to improve the readability and
naturalness of the text content).
d. Lexical selection 624 (adding words to concepts).
e. Generation of reference expressions 65 (generation of expressions
2
identifying objects and regions). This task also involves making decisions
about pronouns and other types of anaphora.
f. Implementation of 66 (text generation taking into account the rules of
2
syntax, morphology and spelling, for example, the use of verbs in the
necessary tenses).
g. When necessary, use 627 machine learning methods (most often LSTM) on a
large set of input data and corresponding (human-written) output texts, for
example to generate text captions for images.
13. Natural Language Understanding 72 (NLU) is transformation of text into logical
structures for controlling NLP programs, i.e. identification of semantics from a set of
possible notations in the form of natural language concepts. The introduction and
creation of speech meta-model and ontology is efficient and empirical. To build a
formalization of semantics, explicit formalization is used in contrast to implicit
assumptions, for example, about a Closed-World Assumption (CWA - any statement
whose truth is not known is false) against an open one (Open-World Assumption
OWA – the truth of a statement does not depend on the observer's knowledge of its
truth), or subjective YES/NO versus objective TRUE/FALSE.
14. Book Generation 82 is creation of full-fledged books based on NLG technology 62,
production/associative rules, neural network, factual knowledge and generalization
of text 12.
15. Document AI 92 is a platform for training an agent to extract specific necessary data
from various types of documents. Designed for users without AI, ML and NLP
experience to quickly access the necessary content hidden in texts, for example, for
lawyers, business analysts and accountants.
(141)
(142)
(143)</p>
      </sec>
      <sec id="sec-4-9">
        <title>4.2. Examples of modeling the processes of solving typical NLP problems</title>
      </sec>
      <sec id="sec-4-10">
        <title>4.2.1. A formal CLS model for identifying viral news headlines</title>
        <p>The CLS model for identifying viral news headlines is presented as:
  ℎ = 16 ∘ </p>
        <p>∘ 1 ∘  ∘ 3 ∘ 14 ∘ 4,
where 4 is tokenization; 14 is recognition of named entities; 3 is marking of parts of
speech;  is Ngrams (sequences of elements and their frequencies); 1 is clustering; 
is ML based on Neural Networks (NN); 16 is application of SentiWordNet (a lexical-semantic
thesaurus for the analysis of text tonality).</p>
      </sec>
      <sec id="sec-4-11">
        <title>4.2.2. Correction of grammatical and stylistic errors</title>
        <p>Another relevant NLP technology is error correction, the main tasks of which are error
identification, error correction, and user training. Grammatical error correction 22 is one of
the sub-processes of correcting various types of errors. Bug fixes provided:
  = 2( ),</p>
        <p>2
where   is the result of identifying and correcting grammatical/stylistic errors at all
levels of linguistic analysis (phonology/orthography, morphology, vocabulary, syntax,
semantics, pragmatics); 22 is the basic error correction process in the text data array X.</p>
        <p>The detailed error correction process is given as:</p>
        <p>= 224 ∘ 223 ∘ 222 ∘ 221,
coreference resolution 11;</p>
        <p>where  ′ = 221( ,  ,   ,   ,   ,  , 3, 3, 11) is a pattern matching check based on
complex multilayer NLP- rules   based on regular expressions   , lexical vand
grammatical victionaries, POS-tags   and part-of-speech marking 3, parsing trees 3 and
 ′′ = 222( ′,   ,  ,   ,  , 2212, 2222) statistical methods for refining corrections ( is–

Ngrams analysis on based on the set of tokens   , POS-tags  , and the history of analogues
  ; 2212 is error identification through the usage statistics with similar words/errors in the
text; 2222 is correction to the most likely variant of possible analogues);
classifiers   for a specific language, analysis N-grams  (for example, bigrams with
appropriate analysis of left and right context followed by replacement of the best option
from POS Ngrams of the model), machine learning rules   as a choice from several correct
options or identification correct but rare application, ML-based error detection 2213 and
error correction 2223 processes, data annotation   for training 2233, feature selection for
training 2243 and classifier training 2253 using random method forest, logistic regression or
other (sometimes pattern matching and simple statistical data cannot generalize decision
options, for example, identification and selection of a preposition or article in English, and
an adjective in Ukrainian; derivational word-formation morphology; run-on sentences, i.e.
independent or subordinate clauses are not joined by a conjunction or punctuation);
 
= 224( ′′′,   ,  ,   ,  , 2214, 2224) is neural machine translation based on Noisy

channel translation processes 2241 (spelling check, question answers, speech recognition
and machine translation based on finding a predicted word with a given word, where
symbols are somehow encrypted) and Round-trip translation 2224 (two-way translation
from the source language to the target language to assess the quality/accuracy of the result).</p>
        <p>The N-gram model assigns probabilities to sentences and sequences of words based on
counting N-grams, for example, according to the Markov assumption:

 =1
 (   ) ∏  (  |  −1),
(144)
where    is the  -th chain or sentence of  words;   is the current word in the  -th chain
or sentence;   −1 is the previous word in the  -th chain or sentence. To identify an error, it
is necessary to determine grammatical or stylistic features using the rules of marking parts
of speech 3, parsing: dependencies 13 and constituents 32 or in the form of some
combination of these methods 33.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>The general structure of CLS processing of textual content in the Ukrainian language and
the conceptual scheme/model of the functioning of a typical CLS based on modeling the
interaction of the main processes and IS components have been developed.</p>
      <p>
        Modeling of the main NLP processes of CLS was carried out due to the interaction of the
main processes/components of IS and methods of linguistic processing of text content
adapted to the Ukrainian language on the basis of grapheme, morphological, lexical,
syntactic, semantic, structural, ontological and pragmatic analysis, which allowed to
improve the IT of intellectual analysis of the text flow for solving a specific NLP problem.
This ensured the adaptation of NLP processes for the analysis of Ukrainian-language textual
content. A formal model of a computer linguistic system for processing Ukrainian-language
textual content was developed and described, which made it possible to determine the main
structural elements and operators of natural language processing at each level of text
analysis such as grapheme/phonological, morphological, syntactic, semantic, referential,
structural, ontological and pragmatic. Due to the complexity of the morphology of the
Ukrainian language, detailed attention is paid to the description of the model of
morphological analysis of textual content. Examples of modeling processes for solving
typical NLP problems such as CLS identification of viral news headlines and correction of
grammatical and stylistic errors are given.
[11] N. Liubchenko, A. Podorozhniak, V. Oliinyk, Research Application of the Spam Filtering
and Spammer Detection Algorithms on Social Media, CEUR Workshop Proceedings
3171 (2022) 116-126.
[12] V.B. Fernandes, et al., Corrigendum to “A spam filtering multi-objective optimization
study covering parsimony maximization and three-way classification”, Applied Soft
Computing Journal 55 (2017) 565.
[13] V. Basto-Fernandes, et al., A spam filtering multi-objective optimization study covering
parsimony maximization and three-way classification, Applied Soft Computing Journal
48 (2016) 111–123.
[14] I Yevseyeva, V. Basto-Fernandes, D. Ruano-Ordás, J.R. Méndez, Optimising anti-spam
filters with evolutionary algorithms, Expert Systems with Applications 40(
        <xref ref-type="bibr" rid="ref10">10</xref>
        ) (2013)
4010–4021.
[15] D. Ruano-Ordás, V. Basto-Fernandes, I. Yevseyeva, J.R. Méndez, Evolutionary
multiobjective scheduling for anti-spam filtering throughput optimization, Lecture Notes in
Computer Science 10334 (2017) 137–148.
[16] O. Cherednichenko, O. Kanishcheva, Readability Evaluation for Ukrainian Medicine
      </p>
      <p>Corpus(UKRMED), CEUR Workshop Proceedings 2870 (2021) 402-412.
[17] V. Bobicev, O. Kanishcheva, O. Cherednichenko, Sentiment Analysis in the Ukrainian
and Russian News, in: Procedeengs of the First Ukraine Conference on Electrical and
Computer Engineering (UKRCON), 2017, pp. 1050-1055.
[18] O. Cherednichenko, N. Babkova, O. Kanishcheva, Complex Term Identification for</p>
      <p>
        Ukrainian Medical Texts, CEUR Workshop Proceedings 2255 (2018) 146-154.
[19] I. Balush, V. Vysotska, S. Albota, Recommendation System Development Based on
Intelligent Search, NLP and Machine Learning Methods, CEUR Workshop Proceedings
2917 (2021) 584-617.
[20] V. Vysotska, S. Mazepa, L. Chyrun, O. Brodyak, I. Shakleina, V. Schuchmann, NLP Tool
for Extracting Relevant Information from Criminal Reports or Fakes/Propaganda
Content, in: Procedeengs of the IEEE 17th International Conference on Computer
Sciences and Information Technologies (CSIT), 2022, pp. 93-98, doi:
10.1109/CSIT56902.2022.10000563.
[21] O. Prokipchuk, V. Vysotska, P. Pukach, V. Lytvyn, D., Uhryn, Y. Ushenko, Z. Hu, Intelligent
Analysis of Ukrainian-language Tweets for Public Opinion Research based on NLP
Methods and Machine Learning Technology, International Journal of Modern Education
and Computer Science 15(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) (2023) 70–93.
[22] V. Husak, O. Lozynska, I. Karpov, I. Peleshchak, S. Chyrun, A. Vysotskyi, Information
System for Recommendation List Formation of Clothes Style Image Selection According
to User’s Needs Based on NLP and Chatbots, CEUR workshop proceedings 2604 (2020)
788-818.
[23] J. Deriviere, T. Hamon, A. Nazarenko, A scalable and distributed NLP architecture for
web document annotation, Lecture Notes in Computer Science 4139 (2006) 56–67.
[24] M. Boyè, T.M. Tran, N. Grabar, NLP-oriented contrastive study of linguistic productions
of alzheimer’s and control people, Lecture Notes in Computer Science 8686 (2014)
412–424.
[25] A. Berko, V. Andrunyk, L. Chyrun, M. Sorokovskyy, O. Oborska, O. Oryshchyn, M.
      </p>
      <p>Luchkevych, O. Brodovska, The Content Analysis Method for the Information Resources
Formation in Electronic Content Commerce Systems, CEUR Workshop Proceedings
2870 (2021) 1632-1651.
[26] R. Bekesh, L. Chyrun, P. Kravets, A. Demchuk, Y. Matseliukh, T. Batiuk, I. Peleshchak, R.</p>
      <p>Bigun, I. Maiba, Structural modeling of technical text analysis and synthesis processes,
CEUR Workshop Proceedings 2604 (2020) 562–589.
[27] O. Artemenko, V. Pasichnyk, N. Kunanets, K. Shunevych, Using sentiment text analysis
of user reviews in social media for e-tourism mobile recommender systems, CEUR
workshop proceedings 2604 (2020) 259-271.
[28] N. Grabar, T. Hamon, Automatic Detection of Temporal Information in Ukrainian</p>
      <p>General-language Texts, CEUR Workshop Proceedings 2136 (2018) 1-10.
[29] V. Shyrokov, "Accuracy" vs "Unambiguity" in Linguistics, CEUR Workshop Proceedings
2870 (2021) 1-5.
[30] O. Kozlov, Structural Optimization of Fuzzy Systems based on Determination of</p>
      <p>Linguistic Terms Number, CEUR Workshop Proceedings 2917 (2021) 174-185.
[31] O. Yurchenko, N. Ugolnikova, Linguistic Methods in Social Media Marketing, CEUR</p>
      <p>Workshop Proceedings 2870 (2021) 743-754.
[32] E. Fedorov, O. Nechyporenko, Linguistic Constructions Translation Method Based on</p>
      <p>Neural Networks, CEUR Workshop Proceedings 3396 (2023) 295-306.
[33] E. Fedorov, O. Nechyporenko, Dynamic Stock Buffer Management Method Based on</p>
      <p>Linguistic Constructions, CEUR Workshop Proceedings 2870 (2021) 1742-1753.
[34] A. Taran, The Role of Keyword Language in the Database of World Slavic linguistics
"iSybislaw", CEUR Workshop Proceedings 3171 (2022) 266-276.
[35] A. Taran, Terminology of Computational Linguistics in Terms of Indexing and
Information Retrieval in the System "iSybislaw", CEUR Workshop Proceedings 2870
(2021) 225-234.
[36] А. Taran, Information-retrieval System "Base of the World Slavic Linguistics
(iSybislaw)" in Language Education, CEUR workshop proceedings 2604 (2020)
590599.
[37] I. Bekhta, N. Hrytsiv, Computational Linguistics Tools in Mapping Emotional</p>
      <p>Dislocation of Translated Fiction, CEUR Workshop Proceedings 2870 (2021) 685-699.
[38] D. Sitnikov, P. Sytnikova, A. Kovalenko, Methods of Eliminating Features from Linguistic</p>
      <p>Equations, CEUR Workshop Proceedings 2870 (2021) 877-889.
[39] S. Albota, Linguistically Manipulative, Disputable, Semantic Nature of the Community</p>
      <p>Reddit Feed Post, CEUR Workshop Proceedings 2870 (2021) 769-783.
[40] S. Albota Resolving conflict situations in reddit community driven discussion platform.</p>
      <p>CEUR Workshop Proceedings 2604 (2020) 215–226.
[41] S. Albota, Linguistic and Psychological Features of the Reddit News Post, in:
Procedenngs of the IEEE 15th International Scientific and Technical Conference on
Computer Sciences and Information Technologies, CSIT, 2020, 1, pp. 295–299.
[42] S. Albota, A. Peleshchyshyn, Contradictory statement as a basis for conflict resolution
strategies, CEUR Workshop Proceedings 2588 (2020) 336–345.
[58] N. Hrytsiv, I. Bekhta, M. Tkachivska, V. Byalyk, Sylvia Plath’s I felt-Narrative Label of
The Bell Jar in Ukrainian Translation: Tagging Textness Features, CEUR Workshop
Proceedings 3171 (2022) 240-255.
[59] Z. Kunch, O. Lytvyn, I. Mentynska, Modern Ukrainian Electronic Dictionaries: the
Problem of Implementing Spelling Changes, CEUR Workshop Proceedings 3396 (2023)
32-47.
[60] O. Levchenko, M. Dilai, Key Colour Terms in the Ukrainian Prose Fiction of the 21st</p>
      <p>Century, CEUR Workshop Proceedings 3171 (2022) 49-60.
[61] O. Levchenko, N. Romanyshyn, D. Dosyn, Method of Automated Identification of
Metaphoric Meaning in Adjective + Noun Word Combinations (Based on the Ukrainian
Language), CEUR Workshop Proceedings 2386 (2019) 370-380.
[62] V. Starko, A. Rysin, VESUM: A Large Morphological Dictionary of Ukrainian As a</p>
      <p>Dynamic Tool, CEUR Workshop Proceedings 3171 (2022) 61-70.
[63] V. Starko, Implementing Semantic Annotation in a Ukrainian Corpus, CEUR Workshop</p>
      <p>Proceedings 2870 (2021) 435-447.
[64] O. Synchak, V. Starko, Ukrainian Feminine Personal Nouns in Online Dictionaries and</p>
      <p>Corpora, CEUR Workshop Proceedings 3171 (2022) 775-790.
[65] N. Romanyshyn, Application of Corpus Technologies in Conceptual Studies (based on
the Concept Ukraine Actualization in English and Ukrainian Political Media Discourse),
CEUR workshop proceedings 2604 (2020) 472-488.
[66] M. Shvedova, The General Regionally Annotated Corpus of Ukrainian (GRAC,
uacorpus.org): Architecture and Functionality, CEUR workshop proceedings 2604
(2020) 489-506.
[67] M. Shvedova, N. Prydvorova, I. Skibina, Normalization of Early Modern Ukrainian in
GRAC: the Case of Lesia Ukrainka's Works, CEUR Workshop Proceedings 3171 (2022)
71-80.
[68] S. Kubinska, R. Holoshchuk, S. Holoshchuk, L. Chyrun, Ukrainian Language Chatbot for
Sentiment Analysis and User Interests Recognition based on Data Mining, CEUR
Workshop Proceedings 3171 (2022) 315-327.
[69] N. Borysova, K. Melnyk, N. Babkova, Z. Kochuieva, V. Melnyk, Gender Classification of</p>
      <p>Surnames: Ukrainian aspect, CEUR Workshop Proceedings 3171 (2022) 354-364.
[70] A. Dmytriv, S. Holoshchuk, L. Chyrun, R. Holoshchuk, Comparative Analysis of Using
Different Parts of Speech in the Ukrainian Texts Based on Stylistic Approach, CEUR
Workshop Proceedings 3171 (2022) 546-560.
[71] K. S. Mandziy, U. V. Yurlova, M. P. Dilai, English-Ukrainian Parallel Corpus of IT Texts:
Application in Translation Studies, CEUR Workshop Proceedings 3171 (2022)
724736.
[72] K. Vyrodov, A. Chupryna, R. Kotelnykov, Detecting of Anti-Ukrainian Trolling Tweets,</p>
      <p>CEUR Workshop Proceedings 3396 (2023) 48-62.
[73] L. Kobylyukh, Z. Rybchak, O. Basystiuk, Analyzing the Accuracy of Speech-to-Text APIs
in Transcribing the Ukrainian Language, CEUR Workshop Proceedings 3396 (2023)
217-227.
[74] K. Datsyshyn, Z. Haladzhun, N. Kunanets, O. Hotsur, N. Veretennikova, Neologisms with
the Prefix Anti- in the Ukrainian Online Media in the Covid-19 Pandemic Period, CEUR
Workshop Proceedings 3171 (2022) 192-211.
[75] Z. Haladzhun, K. Datsyshyn, Y. Bidzilya, N. Kunanets, N. Veretennikova,
“Antivaccinationists&amp;Anti-vax”: Linguistic Means of Actualizing Assessment in the
Headlines and Leads of Ukrainian Text Media, CEUR Workshop Proceedings 3396
(2023) 118-129.
[76] Y. Chychkarov, O. Zinchenko, Handwritten Ukrainian Character Recognition using a
Convolutional Neural Networks and Synthetic Dataset, CEUR Workshop Proceedings
3426 (2023) 109-121.
[77] T. Basyuk, A. Vasyliuk, Peculiarities of an Information System Development for
Studying Ukrainian Language and Carrying out an Emotional and Content Analysis,
CEUR Workshop Proceedings 3396 (2023) 279-294.
[78] A. Dmytriv, V. Vysotska, M. Bublyk, The Speech Parts Identification for Ukrainian Words
Based on VESUM and Horokh Using, in: Proceedings of the IEEE 16th International
Conference on Computer Sciences and Information Technologies (CSIT), 22-25 Sept.,
Lviv, Ukraine, 2021, vol. 2, pp. 21–33.
[79] V. Vysotska, O. Markiv, S. Teslia, Y. Romanova, I. Pihulechko, Correlation Analysis of
Text Author Identification Results Based on N-Grams Frequency Distribution in
Ukrainian Scientific and Technical Articles, CEUR Workshop Proceedings 3171 (2022)
277-314.
[80] V. Lytvyn, P. Pukach, V. Vysotska, M. Vovk, N. Kholodna, Identification and Correction
of Grammatical Errors in Ukrainian Texts Based on Machine Learning Technology,
Mathematics 11 (2023) 904. doi:10.3390/math11040904.
[81] P. Kryndach, V. Vysotska, S. Chyrun, L. Chyrun, S. Goloshchuk, R. Holoshchuk, Analysis
of Semantic Relationships in Ukrainian Text Content Based on Word2Vec and Machine
Learning, in: Proceedings of the IEEE 18th International Conference on Computer
Sciences and Information Technologies (CSIT), Lviv, 19-21 October, 2023.
[82] B. Rusyn, V. Vysotska, L. Pohreliuk, Model and architecture for virtual library
information system, in: Proceedings of the International Conference on Computer
Sciences and Information Technologies, CSIT, 2018, pp. 37-41. doi:
10.1109/STCCSIT.2018.8526679.
[83] B. Rusyn, V. Lytvyn, V. Vysotska, M. Emmerich, L. Pohreliuk, The Virtual Library System
Design and Development, Advances in Intelligent Systems and Computing 871 (2019)
328-349. doi: 10.1007/978-3-030-01069-0_24.
[84] L. Chyrun, Y. Burov, B. Rusyn, L. Pohreliuk, O. Oleshek, A. Gozhyj, І. Bobyk, Web resource
changes monitoring system development, CEUR Workshop Proceedings 2386 (2019)
255–273.
[85] B. Rusyn, L. Pohreliuk, A. Rzheuskyi, R. Kubik, Y. Ryshkovets, L. Chyrun, S. Chyrun, A.</p>
      <p>Vysotskyi, V.B. Fernandes, The mobile application development based on online music
library for socializing in the world of bard songs and scouts’ bonfires, Advances in
Intelligent Systems and Computing 1080 (2020) 734–756. doi:
10.1007/978-3-03033695-0_49.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Khairova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grabar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Burov</surname>
          </string-name>
          , Preface: Computational Linguistics Workshop, CEUR Workshop Proceedings 3396 (
          <year>2023</year>
          ). URL: https://ceur-ws.org/Vol3396/.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Jonek-Kowalska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kowalska-Styczen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kupriianov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kanishcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cherednichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grabar</surname>
          </string-name>
          ,
          <source>Preface: Computational Linguistics and Intelligent Systems, CEUR Workshop Proceedings</source>
          <volume>3171</volume>
          (
          <year>2022</year>
          ). URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3171</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cherednichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kupriianov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kanishcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grabar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kowalska-Styczen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Jonek-Kowalska</surname>
          </string-name>
          ,
          <source>Preface: Computational Linguistics and Intelligent Systems, CEUR Workshop Proceedings</source>
          <volume>2870</volume>
          (
          <year>2021</year>
          ). URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2870</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grabar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cherednichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kanishcheva</surname>
          </string-name>
          ,
          <source>Preface: Computational Linguistics and Intelligent Systems, CEUR Workshop Proceedings</source>
          <volume>2604</volume>
          (
          <year>2020</year>
          ). URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2604</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Cherednichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grabar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kowalska-Styczen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <source>Preface: Computational Linguistics and Intelligent Systems, CEUR Workshop Proceedings</source>
          <volume>2362</volume>
          (
          <year>2019</year>
          ). URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2362</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hlibko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Vnukova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Davydenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Pyvovarov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Avanesian</surname>
          </string-name>
          ,
          <article-title>The Use of Linguistic Methods of Text Processing for the Individualization of the Bank's Financial Servise</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3403</volume>
          (
          <year>2023</year>
          )
          <fpage>157</fpage>
          -
          <lpage>167</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Vnukova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Davydenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hlibko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Shorokh</surname>
          </string-name>
          ,
          <article-title>Mastering Computer Linguistics for the Designation of Risks in Cooperation Communications</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>376</fpage>
          -
          <lpage>386</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Savytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Vnukova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Bezugla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Pyvovarov</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Turgut Sübay, Using Word2vec Technique to Determine Semantic and Morphologic Similarity in Embedded Words of the Ukrainian Language</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>2870</volume>
          (
          <year>2021</year>
          )
          <fpage>235</fpage>
          -
          <lpage>248</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Batiuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Holoshchuk</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <article-title>Holoshchuk, Intelligent System for Socialization of Individual's with Shared Interests based on NLP, Machine Learning</article-title>
          and
          <source>SEO Technologies, CEUR Workshop Proceedings</source>
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>572</fpage>
          -
          <lpage>631</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mykytiuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Albota</surname>
          </string-name>
          ,
          <article-title>Spam Filtration System with the Use of Machine Learning Technology</article-title>
          ,
          <source>in: Proceedings of the IEEE 16th International Conference on Computer Sciences and Information Technologies (CSIT)</source>
          ,
          <fpage>22</fpage>
          -
          <lpage>25</lpage>
          Sept.,
          <string-name>
            <surname>Lviv</surname>
          </string-name>
          , Ukraine.
          <year>2021</year>
          , vol.
          <volume>1</volume>
          , pp.
          <fpage>124</fpage>
          -
          <lpage>130</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>