<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Model of Ukrainian-Language Internet Communication Content</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vitalii</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Slobodzian</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kovalchuk</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maryna</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Molchanova</string-name>
          <email>momolchanova@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olena</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sobko</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olexander Mazurets</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olexander Barmak</string-name>
          <email>lexander.barmak@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iurii Krak</string-name>
          <email>yuri.krak@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Glushkov Cybernetics Institute</institution>
          ,
          <addr-line>Kyiv, 40, Glushkov ave., 03187</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Khmelnytskyi National University</institution>
          ,
          <addr-line>Khmelnytskyi, 11, Institutes str., 29016</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Taras Shevchenko National University of Kyiv</institution>
          ,
          <addr-line>Kyiv, 64/13, Volodymyrska str., 01601</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper proposes a model for analyzing the content of the modern Ukrainian everyday language. The model is built by combining the known and significant, for meaning transfer, frequency dictionaries, which cumulatively cover spheres of activities and types of content. Studies established that the best separability, when classifying texts using the proposed model, is observed at a word vector length of 1500 units; it determines the model's optimal dimension. According to the results of the test classification of more than 400 texts, the Text Rank keyword search method has been established as the most suitable for work with the proposed model of the modern Ukrainian everyday language. To compare the effectiveness of different keyword search methods, it was used the method of visual analytics MDS, which is considered an effective and sufficient tool for visual verification of the results of classification of digital texts into different categories, including both thematic classes categories of emotionally charged texts. The conducted research provides opportunities of using the created model of modern Ukrainian everyday language to solve the text analysis issues and their classification by various grounds. It is essential to prevent suicidal tendencies, detect bullying on social networks, determine the negative emotional charged texts, and warn users about potentially harmful content. text data vectorization, web-content classification, Text Rank, YAKE, TF-IDF, dispersion evaluation, MDS, corpus of the everyday Ukrainian language, Internet communication.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        People have a deep and universal need to interact with others, and the greater their communicative
ability, the more satisfying and rewarding will be their lives [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Nowadays, Internet communication
has a
prominent place in the
means of nonverbal
communication, with more than 4 billion users. Due to the COVID-19 pandemic, social media
audiences were significantly increased. On average, Ukrainians spend more than 2 hours a day on
social networks, forming a voluminous amount of information available to many users. This is
comparable to public address because, on average, one post may be read by a large number of people.
Also, recently, alternative communication technologies are being developed, which attract people
with disabilities to communicate [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], expanding the audience and volume of communication.
      </p>
      <p>2022 Copyright for this paper by its authors.</p>
      <p>
        Internet content is not only a means of communication but also an critical factor of influence to
people because it can cause various collective reactions, create fear, panic, especially in the context of
the spread of Covid-19 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Also, internet content has a significant impact on adolescents' risky
behaviors. The study [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] shows that adolescents with friends who drink alcohol and promote such
behavior on social networks have a higher risk of alcohol consumption than adolescents whose
environment has not published photos and posts with similar content. Also, we must not forget the
recent regrettable incidence of suicides of teenagers worldwide from the so-called game "Blue Whale
Challenge," which is closely associated with depression, bullying, self-mutilation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Therefore,
classifying Internet content to prevent socially dangerous manifestations is an essential issue of
computational linguistics. Using such a classification will help prevent suicidal behavior, detect
bullying on social networks, identify negative emotional charged texts and warn users about possible
harmful content.
      </p>
      <p>This research aims to develop a model of the everyday Ukrainian language, within which a
hyperplane classification of Internet content will be possible for the issue of preventing socially dangerous
manifestations that occur when communicating on the Internet.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>We will review recent publications that correlate with the considered problem in one way or
another. Reviewed publications present approaches that correlate with the approaches of this study.
The authors plan to use a vector model to model the Ukrainian-language segment of Internet
communication, which will be based on statistical measures used to assess the importance of the word
in the context of the message, which in turn is part of the collection of messages or corpus (TF-IDF,
Dispersion evaluation). On the other hand, it is necessary to consider works that offer tasks related to
preventing socially dangerous manifestations that occur when communicating on the Internet.</p>
      <p>
        In recent years, the recognition of abuse on social networking platforms has been an active
research issue. In non-native English-speaking countries, social media texts are primarily mixed. The
study [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] presents experiments using several machine learning models, deep learning, and transfer
learning to detect offensive content on Twitter. The experiment results showed that the features of
TFIDF are more suitable for this task. In study [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], it was compared the usage of TF-IDF, n-grams, and
pre-trained MuRIL. The issue is to identify the offensive content from YouTube's mixed-comment
set. The TF-IDF showed the best results for two languages out of three. In study [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the research
focused on identifying offensive content in Tanglish, Manglish, and Malayalam languages, using four
classifiers (SVM, Random Forest, k-nearest neighbors, and Naive Bayes). The proposed model
achieved an accuracy of 76.96% when using a linear SVM with the TF-IDF feature presentation
technique.
      </p>
      <p>
        In study [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], it was conducted experiments in different languages, which show that images
complement natural language processing models (including BERT), taught without prior external
training. Text classification studies were focused on Wikipedia articles, as images usually are
complemented with text, and Wikipedia pages can be written in different languages. In study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], it
was researched online conversations: their course, arguments, and how they are resolved. The main
emphasis in the work is on the identification of "sarcasm." The authors show that identifying sarcasm
in the message helps to understand whether the author of this message agrees or disagrees with the
statement under discussion. The experiment was performed based on functions, using the logistic
regression model from Scikit-learn. In study [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], was considered the impact and effective measures
to counteraction the spread of disinformation in the context of the COVID-19 pandemic. The research
was conducted among Twitter users (analysis of posts). The study explains how to use the
BERTbased model to match facts to tweets and identify misinformation.
      </p>
      <p>
        In study [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], was carried out the allocation of emotional moods and classification of their polarity.
The publication conducted experiments with 8 data sets in English. The study shows the potential
importance of annotating phrases for small data sets about emotional moods. At the same time, the
results show that the performance of modern models for the prediction of polar expressions of
language is poor, which prevents the use of this information in practice. The authors [13] proposed
using emojis to represent abusive words to reveal abusive meaning in social networks because emojis
are extralinguistic information. This approach does not depend on manual annotation and does not
require expensive resources to download (for example, WordNet). In many experiments, the authors
used BERTLARGE as a basis for the most modern text classification to detect offensive posts and the
BERT model.
      </p>
      <p>Social networks face a severe issue of forming rumors and fake news. This is due to their internal
nature of connecting millions of users to millions of others without proper identification. Therefore,
the authors of [14] propose automated detection of rumors by the method of constructing semantic
opposites. The Glove was used to embed and initialize word vectors and construct opposites.</p>
      <p>Therefore, considering the research, this issue is relevant for the Ukrainian everyday language. It is
also prospective to use a vector model to achieve this goal, which will be characterized by statistical
measures used to find the importance of the word in the content of the message, which in turn is part
of the message collection or corpus (TF-IDF, Dispersion evaluation, etc.).</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods and materials</title>
      <p>The Ukrainian language used for communication on the Internet is mixed with other languages.
This mixed language is named "surzhyk." There are no statistically significant corpora for this
language. This fact leads to constraints in using standard approaches in vectorization of language and
searching for features of a text that can be used for further processing by machine learning.</p>
      <p>In this study, it is proposed to use the variant of the bag of words (BOW) model [15, 16] to
vectorize the text. In this variant, the list of feature words is selected from various frequency
dictionaries that are currently available [17-21]. The set of words is generated by discarding stop
words and words that are not often used from the frequency dictionaries. The threshold of such
discarding is defined experimentally. Different frequency dictionaries are combined to get a more
complete list of words.</p>
      <p>Statistical measures are considered in the work for digitizing texts (determination of numerical
values for words included in BOW). Statistical measures are used for assessing the importance of
words in document context, which is a part of the document collection or corpus: TF-IDF [22],
dispersion evaluation [23] etc.</p>
      <p>Visual analytics methods are used to verify the obtained sets of words and numerical values
(measures) of vectors that represent texts in the ability to classify texts. The Multidimensional scaling
(MDS) [24, 25] method is used in this work. Visual assessment of the ability of the proposed
implementation of the model allows assessing the quality of results.
3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>The description of the proposed approach</title>
      <p>The scheme of the proposed approach for building a model for the content analysis of the
Ukrainian-language segment of Internet communication is shown in Figure 1.</p>
      <p>Next, consider each of the stages in detail.
3.2.</p>
    </sec>
    <sec id="sec-5">
      <title>Building of a vector of keywords for the common lexicon of the</title>
    </sec>
    <sec id="sec-6">
      <title>Ukrainian-language segment of the Internet</title>
      <p>The first stage of the approach is to build a vector of keywords for the transmission of meaning in
Internet communication. Building a vector of keywords is a separate subtask because of the
mentioned feature of language (using "surzhyk," distorted words, and profanity). The vector should
not include only keywords from papers. It should be a balanced data set because we need to analyze
the inflected Ukrainian language. That is why frequency dictionaries of the Ukrainian language were
used in works [17-21] (  = { 1,  2, . . . ,   },  = 1. .  ), where j – serial number of the dictionary, n
– number of dictionaries. Every dictionary is a set of words 
= {
1, 
2, . . . , 
},  =
1. .  , where n – number of words in the dictionary. In the study, words were removed from each
frequency dictionary that would not significantly affect the model in the authors' opinion (Words are
inherent in a vast number of texts (stop words that have official meaning and are used to connect
words in the text), infrequent words, etc.). As a result,   contains only selected words. The keyword
vector</p>
      <p>will be a combination of such frequency dictionaries:</p>
      <p>The second stage is the selection of texts that are subject to ideal classification when it is
unambiguously possible to conclude that the selected text belongs to a specific category. A set of texts
D is formed where each text 
∈</p>
      <p>can correspond to a specific category  ∈  , where C is a set of
categories. In this case, binary classification will be used, as our goal is to identify texts with negative
color content as a category.</p>
      <p>In the third stage, each text document is vectorized by known keyword search methods.</p>
      <p>At the preprocessing stage, each text document   ∈  , is converted into a word vector, here i – is
the number of documents in the collection. After that, a corresponding vector of estimates of
(3)
search methods.
occurrences of keywords    , which are in the 
methods are proposed to use for forming estimates.</p>
      <p>Classical TF-IDF has the following formula:
vector, is formed by using each of the keyword
TF-IDF [22], Dispersion evaluation [23], Text Rank [26, 27, 28], YAKE [29] keyword searching
( ,  ,  ) = 
( ,  ) × 
( ,  ,  )
where IDF(t,d,D) is the weight of the term t of document d of the corpus D, and TF(t,d) is the value of
the frequency of the term t in document d.</p>
      <p>Dispersion evaluation is presented as an evaluation of the importance of each word in the analyzed
text by using the dispersion evaluation method. This method evaluates the discriminant force of words
and allows getting words that are evenly distributed from the general set of commonly used words in
the text [23]. According to [23] if a word A in a text consisting of N words is denoted as   , where the
index k is the number of occurrences of the word in the test, and n is its position in the text then the
interval between sequential occurrences of the word is calculated by the following formula:
  
=   +1 −   =</p>
      <p>−  ,
 =
√(  2 ) − (
)2⁄
(
,
)
where m is iteration and n is a position of A word that met k+1 and k time in the text. So, the
dispersion evaluation is calculated by the formula:
where (A) is the average value of sequence A1, A2,.., Ak . К is the number of occurrences of the</p>
      <p>The Text Rank method is intended for modeling text as an undirected weighted graph G=(V,E,W).
In this method, the keyword candidates are vertices V of the graph, and the relationship between two
words is considered as an edge E. W represents the frequency of occurrence in relation to E [26, 27].
The following formula is used for iterative calculation of weights of vertices:
word А.
others.
(4)
(5)
where d is damping factor (default value is 0.85) [26].  (   )  (   ) is the inbound link of   ,
(  ) is the outgoing link of   . Formula (5) shows that the weight of   vertex depends on the
weight of the edge between   to   vertices and the sum of the weight of outgoing edges from   to</p>
      <p>The YAKE algorithm consists of the following four steps [29]: preprocessing and term-candidate
generation; determining the features of terms; counting points of the term; association of similar
terms. At the first stage, the division is performed at the level of sentences, which are further divided
into terms. At the stage of determining the features of terms, each term is evaluated by using special
functions [29]. At the stage of calculating points for the term, the following formula is used:
 ( ) = (
∗ 
)/
+ ((
/ 
) + (
/
)),
(6)
Where Tcase – the importance of capitalization and acronyms, Tposition – more importance is given
to the words that are present at the beginning of the document, Tnorm – word frequency, Trel –
checks for the diversity of context in which this word is used, Tsentence – determines how often the
candidate word occurs with different sentences. The highest score is given to words that are often
found in different sentences. The last step is to combine the evaluation of morphologically similar
words. According to this method, the better word has minimal evaluation.</p>
      <p>The approach to the interpretation of the obtained results (stage 4) is given below.
3.3.</p>
    </sec>
    <sec id="sec-7">
      <title>Criteria for the quality of visual analysis modeling</title>
      <p>Multidimensional scaling (MDS) [24, 25] method is proposed to assess the quality of the obtained
models for the classification task. MDS is one of the methods of reducing the dimension of vector
space. The purpose of the method is to reduce the dimension to make it possible to visualize (3 or
2dimensional). For example, the criterion for reducing dimensionality is the Euclidean distance
between the vectors. Solving the optimization problem, we find the mapping  
→  2, which makes
it possible to obtain a two-dimensional graph of the relative position of the points-vectors and visually
assesses the quality of the model for the classification task.</p>
      <p>Visual criteria for assessing the quality of modeling are proposed (Figure 2).</p>
      <p>Criterion 1 – High level model for text classification. The Figure 2 (a) shows, that the two classes
are clearly separated. It indicates the correctness of the proposed model.</p>
      <p>Criterion 2 – An acceptable level of model for classifying texts. Figure 2 (b) shows that two
classes are tangent with each other. This result can be considered as workable, but it will require
additional expert review to confirm the classification.</p>
      <p>(a)
(b)
(c)</p>
      <p>Criterion 3 – Unsatisfactory level of model for text classification. The Figure 2 (c) shows that the
two classes are almost inseparable. The distance between classes is insignificant, and there is an
intersection in some places. This model can't be considered workable, and it needs refinement.</p>
      <p>It is proposed to use the above criteria to check the quality of the proposed model of the everyday
Ukrainian language. We will consider the model correct if the results' value is in the range between
the first and second criteria.
3.4.</p>
    </sec>
    <sec id="sec-8">
      <title>Dataset</title>
      <p>The following corpora of the Ukrainian categorized texts and frequency dictionaries were used to
build the model and its validation (ability to classify texts of communication on the Internet):
1. БрУК – an open genre-balanced corpus of modern Ukrainian language with volume of 1 million
word usages. The corpus is built on the foundations that formed the basis of the famous English
corpus Brown [17]. This corpus also include the VESUM dictionary (URL:
https://r2u.org.ua/vesum/), which includes defective words, profanity, "surzhik", which is an integral
part of the everyday Ukrainian language.</p>
      <p>2. Corpus of Ukrainian language MOVA.info – designed to search for tokens and word forms in
Ukrainian texts of a specific style (for some part of corpus can be used to search for morphemes and
syntactic structures) [18]. In this paper, it was used to construct a keyword vector.</p>
      <p>3. UA-GEC – the first annotated GEC-corpus of the Ukrainian language. This is a collection of
texts written by ordinary people. It includes texts like essays, blog posts, social networks, reviews,
letters, etc. These texts contain grammatical, stylistic, and spelling mistakes that bring them as close
as possible to everyday language [19].</p>
      <p>4. Ukrainian News collection – a collection of over 150,000 news articles collected from over 20
news resources. Data sets are divided into the following 5 categories: politics, sports, news, business,
technology. The data set was provided by the non-profit student organization FIdo.ai (FIdo Machine
Learning Research Department of the National University of Kyiv-Mohyla Academy) for research
purposes of data analysis (classification, clustering, keyword selection, etc.) [20].</p>
      <p>5. Ukrainian web-building of the University of Leipzig – the corpus of Ukrainian language
texts. Contains frequency dictionaries. Dictionaries were created based on Wikipedia, news sites, web
documents. Text can be downloaded with different size (words): 10 000, 30 000, 100 000, 300 000,
1000 000 [21].</p>
      <p>6. Karpaty bud karkas – set of articles with information about the progress of construction, news
of modern architecture, technology. The corpus contains more than 200 texts containing 500 words in
an average (https://karpatybud.com.ua/statti/).</p>
      <p>7. Gardener's blog – site contains more than 200 texts of about 500 words each, dedicated to the
topic of gardening (https://agro-market.net/ua/news/).</p>
      <p>This number of sources is needed because the everyday Ukrainian language covers many areas of
life. That is why taking only one of the corpora for the study cannot cover the entire vocabulary of
everyday language.</p>
    </sec>
    <sec id="sec-9">
      <title>4. Results and discussion</title>
      <p>The software application in C# was developed to validate the proposed model in its ability to
classify texts in the everyday Ukrainian language. This application converts the textual content of
files from the training set into a digital representation. The main window of the developed application
is shown in Figure 3.
The created program allows setting following parameters:
1. The path to the file containing the key terms.
2. Choose a method of evaluating terms in the text.
3. Choose the corpora of texts to be processed.</p>
      <p>The result of the application is a file that contains a digital representation of each text from the
selected corpora.</p>
      <p>Parameters of the text data analysis environment. The resulting file is sent for processing to a
software application developed in the Python programming language using the Manifold library. This
application reads data from the input file from the previous step and processes the received data using
the MDS method from the Manifold library. The result obtained from the MDS method is visualized
on the graphical interface.</p>
      <p>The basis of the proposed generalized vector  is a dictionary MOVA.info [18], as it is
closest to the tasks of classifying Internet content with the addition of words from other dictionaries.
Filtering by noun was also performed to achieve this goal because the dictionary had different parts of
speech and was heterogeneous. The resulting vector consists of 1500 words.</p>
      <p>(a) (b)
Figure 4: Visualization of binary classification of texts by TF-IDF and vector with (a) 3000 words and
(b) 2000 words length</p>
      <p>The words vector length of the model for vectorization of the Ukrainian-language segment of the
Internet in 1500 units were established based on research results. For example, Figure 4 (a) shows the
binary classification of the described set of texts by TF-IDF by vector length of 3000 words, and
Figure 4 (b) shows the visualization of the binary classification of the described set of texts by
TFIDF by vector lengths of 2000 words. Studies have shown that the best ability to separate in the
classification of texts is observed at a word vector length of 1500 units.</p>
      <p>Texts of two following categories were collected for experimental research: construction
(https://karpatybud.com.ua/statti/) and gardening (https://agro-market.net/ua/news/). These collections
contain 200 texts each and an average length of 500 words for each text. It is enough to use two
categories to validate the proposed model and conduct experimental research because the research
goal is binary classification.</p>
      <p>Result 1. The result of validation of the quality of the model of everyday lexicon of the
Ukrainianlanguage segment of the Internet using the TF-IDF method is shown in Figure 5 (a). The result can be
assessed as unsatisfactory because the categories of some texts were defined incorrectly and the area
of division/delimitation is blurred. It is due to the characteristics of this method because this method is
high sensitivity to the selection of texts of alternative categories. Texts of alternative categories may
be outside the classes used for classification. It may lead to cases of incorrect classification of texts.</p>
      <p>Result 2. The results of validation of the quality of the model of everyday lexicon obtained by
using the method of dispersion evaluation are shown in Figure 5 (b). The result can be assessed as
acceptable because there are texts that are contained on the class boundaries. The method of
dispersion evaluation requires as many appearances of meaningful words in individual texts for the
effective computing of value of the semantic importance of words. Everyday communication is
characterized with using the small number of meaningful words, that is why some time unsatisfactory
statistical values of unique words of the text do not provide sufficient separation of values.</p>
      <p>Result 3. The result of using TextRank method is shown in Figure 5 (c). The quality of the result
can be assessed as high because both categories have a clear separation. To calculate the semantic
importance of words, the Text Rank method uses not only the actual position of words in the text, but
also the relationships between words and the relationship between the frequencies of occurrence of
words in relation to the relationships between words. This allows us to take into account more
parameters of the text than other methods, but the topics of everyday communication do not allow
obtaining many primary indicators for evaluation. That is why this method has shown a fairly high
efficiency of division of texts by category.</p>
      <p>Result 4. The result of using YAKE method is shown in Figure 5 (d). Such results can be
interpreted between high and satisfactory. The categories are divided, but there is no clear boundary.
This method takes into account a number of important indicators of the text, such as word frequency,
variety of context of word occurrences and frequency of word occurrences in different sentences.
However, the characteristic features of everyday communication do not allow the method to use its
advantages. Such advantages include taking into account using of capital letters and abbreviations, the
presence of words at the beginning of the text. Also, when the number of texts increases, the
capabilities of the YAKE method decrease and its use may lead to fuzzy classification. That is why
this method is effective for texts of another type, like scientific articles. The method showed a less
satisfactory result during analysis of texts on everyday communication.</p>
      <p>Result 5. This research was performed by using the set of words that consists of the intersection of
sets of keywords found by different methods. The resulting set of words was about 500 words, which
was common for the methods considered. The obtained words can be considered sufficient for
modeling the task of classification by the sets of texts. The obtained set can be used as a basis for the
formation of a vector model in the future. The formation of the model can be done by supplementing
the basic set of words with words that are specific to the specific tasks like detection of suicidal
ideation, bullying, negative emotional coloring of texts, negative content and etc.</p>
      <p>The results of validation of the model of the everyday lexicon of the Ukrainian-language segment
of the Internet show that the classification of everyday texts is the most effective with using the Text
Rank method for keyword searching. Further research will aim at improving and modifying the
described approaches by testing assumptions, using methods for vectorization of texts, and so on.
Further research will also be aimed at improving the general vector of words of the everyday
Ukrainian language and solving problems of determining the negatively colored text messages in the
segment of Internet communication.</p>
    </sec>
    <sec id="sec-10">
      <title>5. Conclusion</title>
      <p>The paper proposes a modern model of everyday Ukrainian language built by combining
meaningful frequency dictionaries. The model should be used to analyze the content of the
Ukrainianlanguage segment of the Internet, which includes using swear words and grammatically and
syntactically incorrect words because existing frequency dictionaries do not fully cover this segment.
To create a correct model of the modern everyday Ukrainian language, frequency dictionaries of the
existing corpora of the Ukrainian language were used, covering different areas of activity, types of
text, different areas of communication (including everyday communication on the Internet). The
vector of words is obtained by combining frequency dictionaries of these corpora of the Ukrainian
language. The resulting vector was filtered by a noun and limited in number.</p>
      <p>Two orthogonal sets of texts, more than 200 in each, were taken to validate the proposed model.
Each text was vectorized by one of the four proposed keyword search methods (TF-IDF, variance,
text rank, YAKE) and was assigned to a specific category. To assess the quality of the obtained
models, the MDS method was used, and the criteria for the interpretation of the obtained results on a
three-level scale were proposed. This allowed us to determine the keyword searching method that is
best suited for use with the proposed model of modern Ukrainian everyday language.</p>
      <p>The studies identified the following results:
1. It is established that the best ability for separation during classifying texts using the proposed
model of modern Ukrainian everyday language is observed at a word vector length of 1500 words. It
determines the optimal dimension of the model.</p>
      <p>2. According to the results of test classification of more than 400 texts and using visual verification
of classification results by MDS, it was determined that the Text Rank method for keyword searching
is best suited for using the proposed model of modern Ukrainian everyday language.</p>
      <p>3. It was confirmed that the MDS method for visual analysis is an effective and sufficient tool for
visual verification of the classification results of digital texts into various categories, which include
both thematic categories and categories of the emotional coloring of texts.</p>
      <p>The points above allow arguing about the possibility of using the created model of the modern
everyday Ukrainian language to solve the tasks of analysis of textual content of Internet
communication and its classification according to various grounds. This is especially important to
prevent suicidal tendencies, detect bullying on social networks, determine the negative emotional
color of texts, and warn users about possible harmful content.</p>
      <p>A characteristic feature of the approach considered in the paper is high efficiency in working with
the Ukrainian language as a characteristic representative of inflectional languages. Also, the features
of the approach allow being effective for analytical languages, including English. This increases the
value of the approach to working with Ukrainian-language content of everyday communication
because it often contains English words and proper names as borrowings. The approach's
effectiveness for agglutinative languages such as Hungarian is lower because identifying and working
with formats is somewhat different from working with inflections of inflected languages and requires
additional solutions. Thus, it creates a separate area of further research on working with everyday
communication texts with mixed, multilingual content.</p>
      <p>In further research, it is planned to improve the above approaches to vectorization of text with
verification and assumption about the use of the composition of methods and so forth. There is also a
need to focus research on improving the general vector of words of the everyday Ukrainian language
and solving problems of determining the negative color of text messages during online
communication.</p>
    </sec>
    <sec id="sec-11">
      <title>6. References</title>
      <p>Association for Computational Linguistics, Association for Computational Linguistics, 2021, pp.
49–62. doi:10.18653/v1/2021.eacl-main.5.
[13] M. Wiegand, J. Ruppenhofer, Exploiting Emojis for Abusive Language Detection, Proceedings
of the 16th Conference of the European Chapter of the Association for Computational
Linguistics, Association for Computational Linguistics, 2021, pp. 369–380.
doi:10.18653/v1/2021.eacl-main.28.
[14] N. de Silva, D. Dou, Semantic Oppositeness Assisted Deep Contextual Modeling for Automatic
Rumor Detection in Social Networks, Proceedings of the 16th Conference of the European
Chapter of the Association for Computational Linguistics, Association for Computational
Linguistics, 2021, pp. 405–415. doi:10.18653/v1/2021.eacl-main.31.
[15] D. Yan, K. Li , Sh. Gu, L. Yang, Network-Based Bag-of-Words Model for Text Classification,</p>
      <p>IEEE Access, Volume 8 (2020). doi:10.1109/ACCESS.2020.2991074.
[16] C. Macdonald, N. Tonellotto, S. MacAvaney, IR From Bag-of-words to BERT and Beyond
through Practical Experiments, CIKM '21: Proceedings of the 30th ACM International
Conference on Information &amp; Knowledge Management, Association for Computing Machinery,
New York, United States, 2021. doi:10.1145/3459637.3482028.
[17] Corpus of Modern Ukrainian Language (BRUK). URL: https://r2u.org.ua/corpus.
[18] MOVA.info: about the Ukrainian language, linguistics and more. URL: http://www.mova.info/.
[19] UA-GEC: the first annotated GEC-corpus of the Ukrainian language. URL:
https://ua-gecdataset.grammarly.ai/.
[20] Ukrainian News is a collection. URL:
https://github.com/fido-ai/uadatasets/tree/main/ua_datasets/src/text_classification.
[21] Deutscher Wortschatz. Corpora Ukrainian. URL:
https://wortschatz.unileipzig.de/en/download/Ukrainian#ukr_mixed_2014.
[22] Zh. Jiang, Bo Gao, Y. He, Y. Han, P. Doyle, Q. Zhu, Text Classification Using Novel Term
Weighting Scheme-Based Improved TF-IDF for Internet Media Reports, Mathematical Problems
in Engineering (2021). doi:10.1155/2021/6619088.
[23] I. Krak, O. Barmak, O. Mazurets, The practice investigation of the information technology
efficiency for automated definition of terms in the semantic content of educational materials.</p>
      <p>CEUR Workshop Proceedings, 2016, vol.1631, pp. 237–245. doi:10.15407/pp2016.02-03.237.
[24] I. Krak, O. Barmak, E. Manziuk, Using visual analytics to develop human and machine-centric
models: A review of approaches and proposed information technology. Computational
Intelligence. 2020; pp. 1–26. doi:10.1111/coin.12289.
[25] E. L. Fink, D. A. Cai, Multidimensional Scaling, The International Encyclopedia of Media</p>
      <p>Psychology. Hoboken, NJ: Wiley, 2020. doi:10.1002/9781119011071.iemp0282.
[26] A. Kazemi, V. P'erez-Rosas, R. Mihalcea, Biased TextRank: Unsupervised Graph-Based Content
Extraction, Proceedings of the 28th International Conference on Computational Linguistics,
Barcelona, Spain (Online), 2020, pp. 1642–1652. doi: 10.18653/v1/2020.coling-main.144.
[27] Zh. Huang, Zh. Xie, A patent keywords extraction method using TextRank model with prior
public knowledge, Complex Intell. Syst. (2021). doi:10.1007/s40747-021-00343-8.
[28] M. Zhang, X. Li , Sh. Yue, L. Yang, An Empirical Study of TextRank for Keyword Extraction,</p>
      <p>IEEE Access, Volume 8 (2020). doi: 10.1109/ACCESS.2020.3027567.
[29] R. Campos, V. Mangaravite, A. Pasquali, A. Jorge, C. Nunes, A. Jatowt, YAKE! Keyword
extraction from single documents using multiple local features, Information Sciences, Volume
509, 2020, pp. 257–289. doi: 10.1016/j.ins.2019.09.013.
[30] R. A. Yunmar, A. Setiawan, H. Tantriawan, The Combination of YAKE and Language
Processing for Unsupervised Term Extraction Ontology Learning, International Conference on
Science, Infrastructure Technology and Regional Development, 2019.
doi:10.1088/17551315/537/1/012023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>O.</given-names>
            <surname>Hargie</surname>
          </string-name>
          ,
          <source>Skilled Interpersonal Communication. Research, Theory and Practice</source>
          , 7th. ed., London,
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .4324/9781003182269.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Kryvonos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Krak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>Barmak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bagriy</surname>
          </string-name>
          ,
          <article-title>Predictive text typing system for the Ukrainian language</article-title>
          ,
          <source>Cybern. Syst. Anal</source>
          .
          <volume>53</volume>
          (
          <issue>4</issue>
          ) (
          <year>2017</year>
          ) pp.
          <fpage>495</fpage>
          -
          <lpage>502</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10559-017-9951-5.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Katz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hancock</surname>
          </string-name>
          ,
          <source>The Role of Subjective Construals on Reporting and Reasoning about Social Media Use, Social Media + Society</source>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1177/20563051211035350.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Unger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Soto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Fujimoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Pentz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jordan-Marsh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. W.</given-names>
            <surname>Valente</surname>
          </string-name>
          ,
          <source>Offline Friendship Networks on Adolescent Smoking and Alcohol Use</source>
          ,
          <volume>54</volume>
          (
          <issue>5</issue>
          ) (
          <year>2014</year>
          ) pp.
          <fpage>508</fpage>
          -
          <lpage>514</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.jadohealth.
          <year>2013</year>
          .
          <volume>07</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>R. J</surname>
          </string-name>
          . Moreira de Freitas,
          <string-name>
            <given-names>T. N. Carvalho</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Lopes de Melo</surname>
          </string-name>
          , J. do V. e
          <string-name>
            <surname>Silva</surname>
            , K. C. de Oliveira e Melo,
            <given-names>S. Fontes</given-names>
          </string-name>
          <string-name>
            <surname>Fernandes</surname>
          </string-name>
          ,
          <article-title>Adolescents' perceptions about the use of social networks and their influence on mental health</article-title>
          ,
          <source>Enfermería Global</source>
          ,
          <volume>20</volume>
          (
          <issue>64</issue>
          ) (
          <year>2021</year>
          ) pp.
          <fpage>324</fpage>
          -
          <lpage>364</lpage>
          . doi:
          <volume>10</volume>
          .6018/eglobal.462631.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Saumya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <article-title>Offensive language identification in Dravidian code mixed social media text</article-title>
          ,
          <source>Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages, Association for Computational Linguistics</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>36</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Dave</surname>
          </string-name>
          , Sh.
          <string-name>
            <surname>Bhat</surname>
          </string-name>
          , P. Majumder, IRNLP DAIICT@
          <article-title>DravidianLangTech-EACL2021: Offensive Language identification in Dravidian Languages using TF-IDF Char N-grams and MuRIL</article-title>
          ,
          <source>Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages, Association for Computational Linguistics</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>266</fpage>
          -
          <lpage>269</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Sivalingam</surname>
          </string-name>
          , S. Thavareesan, OffTamil@DravideanLangTech-EACL2021:
          <article-title>Offensive Language Identification in Tamil Text</article-title>
          ,
          <source>Proceedings of the First Workshop on Speech and Language Technologies for Dravidian Languages, Association for Computational Linguistics</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>351</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Ch</surname>
            . Ma,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Yoshikawa</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Iwakura</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Beck</surname>
          </string-name>
          , T. Baldwin,
          <article-title>On the (In)Effectiveness of Images for Text Classification, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics</article-title>
          , Association for Computational Linguistics,
          <year>2021</year>
          , pp.
          <fpage>42</fpage>
          -
          <lpage>48</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .eacl-main.
          <volume>4</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Shrivastava</surname>
          </string-name>
          , S. Muresan, «
          <article-title>Laughing at you or with you»: The Role of Sarcasm in Shaping the Disagreement Space</article-title>
          ,
          <year>2021</year>
          , pp.
          <fpage>1998</fpage>
          -
          <lpage>2010</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .eacl-main.
          <volume>171</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Zh. Zhu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Meng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Caraballo</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Jaradat</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Akrami</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Arslan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Jimenez</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          <string-name>
            <surname>Saeef</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Pathak</surname>
          </string-name>
          , Ch. Li,
          <article-title>A Dashboard for Mitigating the COVID-19 Misinfodemic, Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations</article-title>
          ,
          <source>Association for Computational Linguistics</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>105</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2021</year>
          .eacl-demos.
          <volume>12</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Barnes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Øvrelid</surname>
          </string-name>
          , E. Velldal,
          <article-title>If you've got it, flaunt it: Making the most of fine-grained sentiment annotations</article-title>
          ,
          <source>Proceedings of the 16th Conference of the European Chapter of the</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>