<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Styles in Scientific and Technical Publications</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Volodymyr Motyka</string-name>
          <email>volodymyr.motyka@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yaroslav Stepaniak</string-name>
          <email>yaroslav.stepaniak@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariia Nasalska</string-name>
          <email>mariia.nasalska@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victoria Vysotska</string-name>
          <email>victoria.a.vysotska@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>S. Bandera Street, 12, Lviv, 79013</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Osnabrück University</institution>
          ,
          <addr-line>Friedrich-Janssen-Str. 1, Osnabrück, 49076</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A self-developed dataset based on the analysis of more than 300 Ukrainian-language scientific and technical publications from the specialized Bulletin of the Lviv Polytechnic National University of the Information Systems and Networks series for 2001-2016 is selected. It contains information about the lexical and syntactic development of the author's styles of scientific publications -technical direction. Namely are the total number of words in this text, the number of words in a certain text (without repetitions), the number of words with a frequency of 1, the number of words with a frequency of 10 or more, the number of separate sentences, the number of prepositions, the number of conjunctions, Lexical diversity, Syntactic complexity, coefficient of speech coherence, exclusivity index, concentration index. The purpose of the research is to find the differences and dependence of the given data. For this, various methods of visualization and data processing, smoothing methods and correlation analysis are used. Стиль автора, Лексична різноманітність, Синтаксична складність, Коефіцієнт зв'язності мовлення, Індекс винятковості, Індекс концентрації, кореляційний аналіз, згладжування Such variables as Lexical Diversity, Syntactic Complexity, Cohesion of Speech, Index of Exclusiveness, Index of Concentration on Dependency and Distinction were studied. We will describe the variables to improve the further understanding of the work done: Lexical diversity is the ratio of the number of words to the total number of word forms of the text, Syntactic complexity is the ratio of the number of sentences to the number of words of a certain text, Cohesiveness of speech is the ratio of the number of prepositions and conjunctions to the number of individual sentence Exclusiveness index the variability of the vocabulary, i.e. the share of the text occupied by words that occurred 1 time, Concentration index - the share of the text occupied by words that occurred 10 times or more.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In this work, a rather interesting dataset was investigated, which can be described as statistics of the
lexical and syntactic development of literature. The self-developed dataset is based on the results of the
research of more than 300 Ukrainian-language scientific and technical publications from the specialized
Bulletin of the National University "Lviv Polytechnic" of the "Information Systems and Networks"
series for 2001-2016. This dataset reminded us of the popular application Grammarly, the essence of
which is to increase the quality of written communication, offering guidance on correctness, clarity,
appeal and tone of message.</p>
      <p>2023 Copyright for this paper by its authors.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        The scheme of the combination of methods for determining the author of Ukrainian-language textual
content of a scientific and technical direction shows that it consists of lexical and syntactic levels [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
The use of the syntactic level involves the calculation of linguistic relationships in combinations of
words [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The work [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] proposed a model for building an author's style profile, which consists of a
characteristic author's vocabulary and author's syntax [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. To describe the syntax, it is necessary to use
a formalized description of the linguistic relationships between the lexical units of a phrase in a
pluraltheoretical language [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In work [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a formalized description of any text is put forward, but the
formalized description of linguistic relationships between lexical units is not updated. A formalized
description of the text is also found in references [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        In the handbook [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], a formalized textual presentation was compiled for the automation of
procedures for the analysis of scientific and educational texts in order to identify semantically
significant fragments [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] sets out a plural-theoretical description of linguistic relations
in phrases. Such models can be used to describe images of author's vocabulary and author's syntax, but
they do not take into account statistical information about vocabulary frequency and syntax [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The
formalized description, which was used to analyze the text of the terminological dictionary in order to
build a semantic network of its terms, is presented in the reference book [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. However, the proposed
model also does not include accounting for statistical information about the frequency of vocabulary
and syntax [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        Methods of determining the author of Ukrainian-language textual content of a scientific and
technical direction are proposed and investigated in works [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5">1–5</xref>
        ]. Various algorithms [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], in particular
quantitative ones [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], can be used to implement these methods. Therefore, there is a problem of
analyzing such algorithms in order to find the most effective one [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Authorization of authorship is a technique for determining the author of a text when it is not clear
who wrote it [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. This is useful when several people claim authorship of the same publication [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] or
in cases where no one claims authorship of textual content [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], such as so-called trolls in social
networks during information warfare [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The complexity of the problem of the author's text is
obviously exponentially higher, the number of probable authors is greater [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The availability of
author's text samples is also essential in advancing this problem [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Attribution of the author's text
includes the following three problems [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]:
 identifying the author of the text author from the group of probable or expected authors, where
the author is always in the group of suspects [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ];
 not identifying the author of the textual author from the group of probable or expected authors,
where the author may not be in the group of suspects [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ];
 assessment of the possibility of a given text, written by a given author or not [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>
        Therefore, the task of automatically determining the author of scientific and technical textual content
is urgent and requires new (more advanced) approaches to its solution [
        <xref ref-type="bibr" rid="ref27 ref28 ref29 ref30 ref31 ref32 ref33 ref34 ref35">27-36</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods and materials</title>
      <p>We will use the methods of visual presentation of data, smoothing, correlation method to perform
the tasks. Methods of visual presentation of data - methods of presenting data in the form of graphs,
charts and/or other subtypes of them (histograms, pie charts, etc.), time series, etc. Depending on the
specific task, a specific method of data presentation will be used. We will implement these methods
using Microsoft Power BI and/or R tools.</p>
      <p>Smoothing methods are used to reduce the influence of the random component (random fluctuations)
in time series. They make it possible to obtain more "pure" values, which consist only of deterministic
components. Some of the methods are aimed at highlighting some components, for example, the trend
[37-39]. We will implement these methods using Microsoft Excel, R and/or Microsoft Power BI.</p>
      <p>Correlation method (Correlation - analysis) - a method of studying the interdependence of
characteristics in the general population, which are random variables with a normal distribution [40-44]
for different NLP-talks based on text analysis [45-54].</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>Let's open the generated dataset using R Studio:
 Speech coherence coefficient - the ratio of the number of prepositions and conjunctions to the
number of separate sentences: Kz=(Z+S)/(3P), where Z is the number of prepositions, S is the
number of conjunctions, P is the number of separate sentences.
 Exclusiveness index - the variability of the vocabulary, i.e. the share of the text occupied by
words that occurred 1 time, i.e. Iwt =W1/W, where Iwt is the exclusiveness index of the text, W1 is
the number of words with a frequency of 1, W is the number of words in the entire text.
 Concentration index - the share of the text occupied by words that occur 10 times or more: Ikt
=W10/W, where Ikt is the text concentration index, W10 is the number of words with a frequency
of 10 or more, W is the number of words in the entire text.</p>
      <p>Let's group the data by the Year field:
gr&lt;-dt %&gt;%
group_by(Year)%&gt;%
select(Year,W)
new_gr&lt;- gr %&gt;% summarise(avg = mean(W10))</p>
      <p>Let's calculate the quantitative characteristics by choosing the data column w1, which characterizes
the number of words with a frequency of 1, by means of R:
 Sample size – the number of units in the sample: nrow(new_gr)
 Sample mean. We find using the built-in method mean(): median(new_gr$avg, na.rm = FALSE)
 The median of the sample is the number that "divides" "in half" the ordered set of all the values
of the sample, that is, the average value of the changing characteristic, which is contained in the
middle of the series, placed in the order of increasing or decreasing of the characteristic. For this,
we will use the median() method: median(new_gr$avg, na.rm = FALSE)
 Mode - the value that occurs most often in the sample. Since there is no built-in method for
finding it in R, we will define our modes function:
## modes function
modes &lt;-function(v) {
uniqv &lt;- unique(v)
uniqv[which.max(tabulate(match(v, uniqv)))]}##
modes(new_gr$avg)
 Sample size – the difference between the maximum and minimum value of the sample. To find
the maximum and minimum, use the built-in methods max() and min():</p>
      <p>max(new_gr$avg)-min(new_gr$avg)
 Standard deviation - the amount of spread relative to the arithmetic mean. To find, we will use
the built-in method sd(): sd(new_gr$avg)
 Coefficient of variation – an indicator that determines the percentage ratio of the average
deviation to the average value: sd(new_gr$avg)*100/mean(new_gr$avg, na.rm = FALSE)
 Asymmetry reflects the skewness of the distribution relative to the mode. Let's use the built-in
skewness() method: skewness(new_gr$avg)
 The kurtosis coefficient characterizes the "steepness", that is, the steepness of the rise of the
distribution curve compared to the normal curve. Let's use the kurtosis() method:
kurtosis(new_gr$avg)
 Standard error is the deviation of the sample from the actual mean. To find it, we will use the
formula for calculating the standard error and the sd() method for calculating the standard deviation:
sd(new_gr$avg)/sqrt(nrow(new_gr))</p>
      <p>To find the number of intervals, we will use Sturges' formula, and to find the width of the interval
Scott's formula. Cumulative – a continuous curve is displayed graphically, which gives a more accurate
result compared to a histogram. For construction, we will use the ecdf() function. Finding the number
of intervals and the interval width for the avg attribute:
k&lt;-1+log2(nrow(new_gr)) #Number of intervals
h&lt;-3.5*sd(new_gr$avg)*(nrow(new_gr))^(-1/3) #Interval width</p>
      <p>Construction of a histogram: hist(new_gr$avg, breaks = k, xlab = "", main = "Histogram of w")
Construction of cumulata:
plot(ecdf(new_gr$avg), main="Cumulate", xlab="", ylab = "Frequency", verticals = FALSE)</p>
      <p>Smoothing methods are used to reduce the influence of the random component (random fluctuations)
in time series. They make it possible to obtain more "pure" values, which consist only of deterministic
components. Some of the methods are aimed at highlighting some components, for example, a trend.
Smoothing methods can be conventionally divided into two classes based on different approaches:
analytical and algorithmic.</p>
      <p>The simplest method of forecasting is considered to be an approach that determines the forecast
estimate from the actually achieved level using the average level, average growth, average growth rate.
Extrapolation based on the average level of the series. The resulting confidence interval takes into
account the uncertainty hidden in the estimate of the average value. However, the assumption remains
that the predicted indicator is equal to the sample mean, that is, this approach does not take into account
the fact that individual values of the indicator have fluctuated around the average in the past, and this
will also happen in the future.</p>
      <p>Analytical smoothing methods include regression analysis together with the method of least squares
and its modifications. To identify the main trend by analytical method means to give the studied process
the same development throughout the entire observation period. Therefore, for 4 of these methods, it is
important to choose the optimal function of the deterministic trend (growth curve), which smoothes a
number of observations.</p>
      <p>Forecasting methods based on regression methods are used for short- and medium-term forecasting.
They do not allow for adaptation: with the receipt of new data, the forecast construction procedure must
be repeated from the beginning. The optimal length of the lead-up period is determined separately for
each economic process, taking into account its statistical instability.</p>
      <p>The most widely used are the methods of smoothing time series using moving averages. For moving
average smoothing, we will use Kendel's formulas to calculate the lost levels at the beginning and end
of the smoothed series. Let's prepare the data for using smoothing methods:
ma &lt;- new_gr %&gt;% select(Year,avg) %&gt;%
mutate(ma1 = rollmean(avg, k = 3, fill = NA), ma2 = rollmean(avg, k = 5, fill = NA),
ma3 = rollmean(avg, k = 7, fill = NA))</p>
      <p>
        The method of smoothing according to Kendel's formulas:
k_ma1&lt;-matrix(c(5,2,-1,6,3,6),byrow = TRUE,nrow=2)
ma$ma1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]&lt;-0
ma$ma1[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]&lt;-0
for(i in 1:3){
ma$ma1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]&lt;-ma$ma1[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]+k_ma1[1,i]*ma$avg[i]/k_ma1[nrow(k_ma1),1]
ma$ma1[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]&lt;-ma$ma1[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]+k_ma1[1,i]*ma$avg[17-i]/k_ma1[nrow(k_ma1),3]}
k_ma2&lt;-matrix(c(3,2,1,0,-1,4,3,2,1,0,5,10,5,10,5),byrow = TRUE,nrow=3)
ma$ma2[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]&lt;-0
ma$ma2[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]&lt;-0
ma$ma2[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]&lt;-0
ma$ma2[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]&lt;-0
for(j in 1:2){
for(i in 1:5) {
ma$ma2[j]&lt;-ma$ma2[j]+k_ma2[j,i]*ma$avg[i]/k_ma2[nrow(k_ma2),j]
ma$ma2[17-j]&lt;-ma$ma2[17-j]+k_ma2[j,i]*ma$avg[17-i]/k_ma2[nrow(k_ma2),j] }}
k_ma3&lt;-matrix(c(seq(13,-5,by=-3),seq(5,-1,by=-1),seq(7,1,by=-1),28,14,28,7,28,14,28),
byrow = TRUE,nrow=4)
alpha&lt;-0.1
exp_smooth&lt;-1:16
exp_smooth[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]&lt;-ma$avg[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
for(i in 2:16){ exp_smooth[i]&lt;-ma$avg[i]*alpha +(1-alpha)*exp_smooth[i-1]}
      </p>
      <p>Visualization:
Median filtering:</p>
      <p>
        Visualization:
ggplot(ma,mapping= aes(x=Year)) + geom_line(mapping= aes(y=avg, col="Real"),lwd=1.5) +
geom_line(mapping= aes(y=exp_smooth, col="es"),lwd=1.5)+
scale_color_manual(values= c("Real"="blue","es"="red"))+ labs(x="",y="",title ="alpha = 0.30")+
theme(legend.title = element_blank(),plot.title = element_text(hjust = 0.5))
med_fil&lt;-1:16
med_fil[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]&lt;-(5*ma$avg[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]+2*ma$avg[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]-ma$avg[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ])/6
med_fil[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]&lt;-(-ma$avg[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]+2*ma$avg[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]+5*ma$avg[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ])/6
for(i in 2:15){
      </p>
      <p>med_fil[i]&lt;-max(min(ma$avg[i-1],ma$avg[i]),min(ma$avg[i],ma$avg[i+1]),min(ma$avg[i-1],ma$avg[i+1]))}
ggplot(ma,mapping= aes(x=Year)) + geom_line(mapping= aes(y=avg, col="Real"),lwd=1.5) +
geom_line(mapping= aes(y=med_fil, col="Median"),lwd=1.5)+
scale_color_manual(values= c("Real"="blue","Median"="red"))+
labs(x="",y="Views",title ="Median filter")+
theme(legend.title = element_blank(),plot.title = element_text(hjust = 0.5))</p>
      <p>Turning points:
tp_mf&lt;-turnpoints(med_fil)
summary(tp_mf)
plot(ma$avg, type = "l")
lines(tp_mf)</p>
      <p>Visualization of turning points:
Correlation coefficient: cor(ma$avg,med_fil)</p>
      <p>In general, correlation can be described as any statistical relationship of data. Correlation allows us
to see the trends of changes in the average values of the functions depending on the parameter changes.
Correlation can be positive or negative. Negative correlation is a correlation in which an increase in one
variable is associated with a decrease in another, and the correlation coefficient is negative. Positive
correlation is a correlation in which an increase in one variable is associated with an increase in another,
and the correlation coefficient is positive.</p>
      <p>Construction of the correlation field (plot)
plot(dt$Kl, dt$W, main="Correlation field", xlab="lexical diversity",</p>
      <p>ylab="Word count without duplicates")
plot(dt$Ks, dt$P, main="Correlation field", xlab="Syntax complexity", ylab="Sentance count")
plot(dt$Kz, dt$P, main="Correlation field", xlab="Coefficient of coherent speech",</p>
      <p>ylab="Sentance counts")
plot(dt$Iwt, dt$W1, main="Correlation field", xlab="Coefficient of coherent speech",
ylab= "Count of words that have only one duplicate")
plot(dt$Ikt, dt$W10, main="Correlation field", xlab="Coefficient of coherent speech",
ylab="Count of words that have 10 or more duplicates")</p>
      <p>Finding multiple correlation coefficients:
numericData &lt;- cbind(dt$N,dt$W,dt$P,dt$Ks)
chart.Correlation(numericData, histogram=FALSE, pch=19)
numericData &lt;- cbind(dt$P,dt$Z,dt$S,dt$Kz)
chart.Correlation(numericData, histogram=FALSE, pch=19)
numericData &lt;- cbind(dt$N,dt$W, dt$W1,dt$Iwt)
chart.Correlation(numericData, histogram=FALSE, pch=19)
numericData &lt;- cbind(dt$N,dt$W,dt$W10,dt$Ikt)
chart.Correlation(numericData, histogram=FALSE, pch=19)</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>Let's present the dataset in the form of a table and group the data by years:</p>
      <p>The number of words in a certain text without repetitions and the number of words with a frequency
of more than 10:</p>
      <p>Let's find the statistical parameters for the attribute (Table 1).
After executing the code, we have histograms and corresponding cumulates:</p>
      <p>Let's analyse the change in time series trends using smoothing methods.</p>
      <p>Using Kendel's formulas, we obtained the initial and final values that were lost in the calculation of
the averages, depending on the average from which we calculate the data.</p>
      <p>It can be noted that graphs with k&gt;5 are not very suitable for us to identify trends, since we do not
have a large date interval, only 16 years. For more accurate detection of trends, it is desirable to take
k=3.5. At k=7 was plotted to show that the data is smoothed too much.</p>
      <p>The correlation coefficients are not large and positive. This is probably because we took annual
averages everywhere.</p>
      <p>It can be noted that ma4, ma5, ma6, ma7 are not very suitable for detecting trends, since we do not
have a large date interval, only 16 years. For more accurate detection of trends, it is advisable to take
ma1, ma2 or ma3.</p>
      <p>It can be noted that ma4, ma5, ma6, ma7 are not very suitable for identifying trends, since we do not
have a large date interval, only 40 days. For more accurate detection of trends, it is advisable to take
ma1, ma2 or ma3. The number of turning points allows better analysis of trends.</p>
      <p>It can be noted that ma4, ma5, ma6, ma7 are not very suitable for detecting trends, since we do not
have a large date interval, only 16 years. For more accurate detection of trends, it is advisable to take
ma1, ma2 or ma3. The number of turning points allows better analysis of trends.</p>
      <p>Correlation coefficients approach 1 and decrease as the step increases, as less and less data will
influence the average.</p>
      <p>Exponential smoothing directly depends on the latest data, i.e. how the weighted average will react
quickly to changes.</p>
      <p>Median smoothing completely removes single extreme or anomalous values of levels that are
separated from each other by at least half of the smoothing interval; preserves sharp changes in the trend
(moving average and exponential smoothing smooth them); effectively removes single levels with very
large or very small values that are random in nature and stand out sharply from other levels.</p>
      <p>As can be seen from fig. 26, median filtering removed random levels that are random in nature. As
a result, we have a more stable schedule.</p>
      <p>From fig. 30, it can be seen that the average number of words without repetitions remains
approximately at the same level. This means that the "jumps" of the graph are not so important, but are
only isolated cases and simply related to the texts. Note that the correlation is high, because the median
filtering does not calculate, does not generalize, but shows the median on a certain interval. That is why
median filtering is very effective when studying time series.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>We will investigate in detail the dependence of attributes on the basis of correlation analysis of time
sequences. To do this, we will construct multiple correlation graphs to find the most significant
variables by analyzing correlation relations and construct correlation graphs of the most significant
variables found.</p>
      <p>This visualization is built on three attributes, namely "total number of words of this text" (var 1),
"number of words in a certain text (without repetitions)" (var 2), "Lexical diversity" (var 3). It can be
seen from this visualization that the correlation coefficient between the second and third variables is the
largest, so let's take a closer look at their correlation graph:</p>
      <p>The graph shows the linear dependence of the variables - when one variable grows, the other grows
accordingly.</p>
      <p>This visualization is built on four attributes, namely the total number of words of the text (var 1),
the number of words in a certain text (without repetitions) (var 2), the number of separate sentences
(var 3) and Syntactic complexity (var 4). It can be seen from this visualization that the correlation
coefficient between the third and fourth variables is the most significant, so let's take a closer look at
their correlation graph:</p>
      <p>Thanks to the graph, you can make sure that when the dependent variable increases, the independent
variable drops rapidly, which corresponds to this negative correlation coefficient.</p>
      <p>This visualization is built on four attributes, namely the number of separate clauses (var 1), the
number of prepositions (var 2), the number of conjunctions (var 3) and the Speech Coherence Factor
(var 4). It can be seen from this visualization that the correlation coefficient between the first and fourth
variables is the most significant, so let's take a closer look at their correlation graph:</p>
      <p>This graph visualizes the almost identical logic of dependence as in the previous case.</p>
      <p>This visualization is built on four attributes, namely the total number of words of this text (var 1),
the number of words in a specific text (without repetitions) (var 2), the number of words with a
frequency of 1 (var 3), and the Uniqueness Index (var 4). It can be seen from this visualization that the
correlation coefficient between the third and fourth variables is the most significant, so let's take a closer
look at their correlation graph:</p>
      <p>The graph shows the linear dependence of the variables - as one variable grows, the other grows,
which is why the positive correlation coefficient shows.</p>
      <p>This visualization is built on four attributes, namely the total number of words of this text (var 1),
the number of words in a certain text (without repetitions) (var 2), the number of words with a frequency
of 10 or more (var 3) and the Concentration Index (var 4 ). It can be seen from this visualization that
the correlation coefficient between the third and fourth variables is the most significant, so let's take a
closer look at their correlation graph:</p>
      <p>The graph shows the linear dependence of the variables - as one variable grows, the other grows,
which is why the positive correlation coefficient shows.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusions</title>
      <p>A simple moving average is suitable for identifying trends in the past, which will help us predict the
future with less error. This will allow us to predict how succinct texts will be in the future. To do this,
they need methods that quickly respond to the latest data. When performing work on such methods, we
used exponential smoothing.</p>
      <p>During data analysis, it was found that the larger the text, the fewer words it contains without
repetitions, which is logical, since it is difficult to pick up new words every time. Over time, the number
of words without repetitions does not increase and does not decrease significantly, although it is not
immediately visible on the graph. We reached this conclusion using median filtering</p>
      <p>It is also worth noting that the relationship between the number of words, the number of words
without repetitions, lexical diversity, syntactic complexity, the coefficient of speech coherence, the
exclusivity index and the concentration index was investigated. There is a direct relationship between
them, so when one of these attributes increases, the others will also increase.</p>
    </sec>
    <sec id="sec-8">
      <title>8. References</title>
      <p>[36] C. Luckman, S. A. Wagovich, C. Weber, B. Brown, S. E. Chang, N. E. Hall, N. B. Ratner, Lexical
diversity and lexical skills in children who stutter, Journal of Fluency Disorders 63 (2020) 105747.
[37] A. Fronzetti Colladon, C. A. D’Angelo, P. A. Gloor, Predicting the future success of scientific
publications through social network and semantic analysis, Scientometrics 124 (2020) 357-377.
[38] S. Kumar, M. Yadava, P. P. Roy, Fusion of EEG response and sentiment analysis of products
review to predict customer satisfaction. Information Fusion, 52 (2019) 41-52.
[39] V.V. Hnatushenko, P. I. Kogut, M. V. Uvarov, On Optimal 2-D Domain Segmentation Problem
via Piecewise Smooth Approximation of Selective Target Mappings. Journal of Optimization,
Differential Equations and Their Applications 27(2) (2019). 60–95. DOI: 10.15421/141908.
[40] Hnatushenko V., Kogut P., Uvarov M. On Satellite Image Segmentation via Piecewise Constant</p>
      <p>Approximation of Selective Smoothed Target Mapping, Applied Mathematic
[41] L. P. Morency, R. Mihalcea, P. Doshi, Towards multimodal sentiment analysis: Harvesting
opinions from the web. In Proceedings of the 13th international conference on multimodal
interfaces, 2011, pp. 169-176.
[42] N. Romanyshyn, Algorithm for Disclosing Artistic Concepts in the Correlation of Explicitness and</p>
      <p>Implicitness of Their Textual Manifestation, CEUR Workshop Proceedings 2870 (2021) 719-730.
[43] Y. Yusyn, T. Zabolotnia, Methods of Acceleration of Term Correlation Matrix Calculation in the</p>
      <p>Island Text Clustering Method, CEUR workshop proceedings, Vol-2604 (2020) 140-150.
[44] B. Rusyn, V. Ostap, O. Ostap, A correlation method for fingerprint image recognition using
spectral features, in: Proceedings of the International Conference on Modern Problems of Radio
Engineering, Telecommunications and Computer Science, TCSET, 2002, pp. 219–220.
[45] S. Voloshyn, O. Markiv, V. Vysotska, I. Dyyak, L. Chyrun, V. Panasyuk, Emotion Recognition
System Project of English Newspapers to Regional E-Business Adaptation, in: IEEE 17th
International Conference on Computer Sciences and Information Technologies (CSIT), 2022, pp.
392-397, doi: 10.1109/CSIT56902.2022.10000527.
[46] N. Kholodna, V. Vysotska, S. Albota, A Machine Learning Model for Automatic Emotion</p>
      <p>Detection from Speech, CEUR Workshop Proceedings, Vol-2917 (2021) 699-713.
[47] M. Hryntus, M. Dilai, Translating emotion metaphors from English into Ukrainian: based on the
parallel corpus of fiction, CEUR Workshop Proceedings, Vol-3171 (2022) 737-750.
[48] O. Bisikalo, V. Kovenko, I. Bogach, O. Chorna, Explaining Emotional Attitude Through the Task
of Image-captioning, CEUR Workshop Proceedings, Vol-3171 (2022) 1056-1065.
[49] K. Smelyakov, O. Bohomolov, M. Kizitskyi, A. Chupryna, Identification of Modern Facial</p>
      <p>Emotion Recognition Models, CEUR Workshop Proceedings, Vol-3171 (2022) 1267-1281.
[50] D. Nazarenko, I. Afanasieva, N. Golian, V. Golian, Investigation of the Deep Learning Approaches
to Classify Emotions in Texts, CEUR Workshop Proceedings, Vol-2870 (2021) 206-224.
[51] I. Bekhta, N. Hrytsiv, Computational Linguistics Tools in Mapping Emotional Dislocation of</p>
      <p>Translated Fiction, CEUR Workshop Proceedings, Vol-2870 (2021) 685-699.
[52] I. Spivak, S. Krepych, O. Fedorov, S. Spivak, Approach to Recognizing of Visualized Human
Emotions for Marketing Decision Making Systems, CEUR Workshop Proceedings, Vol-2870
(2021) 1292-1301.
[53] P.C. Thoumelin, N. Grabar, Subjectivity in the medical discourse: On uncertainty and emotional
markers, Revue des Nouvelles Technologies de l'Information, E.26 (2014) 455–466.
[54] N. Grabar, L.O. Dumonet, Automatic computing of global emotional polarity in French health
forum messages, Lecture Notes in Computer Science 9105 (2015) 243–248.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>V.</given-names>
            <surname>Lytvyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pukach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Nytrebych</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Demkiv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Senyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Malanchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sachenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kovalchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Huzyk</surname>
          </string-name>
          ,
          <article-title>Analysis of the developed quantitative method for automatic attribution of scientific and technical text content written in Ukrainian</article-title>
          , volume
          <volume>6</volume>
          (
          <issue>2</issue>
          -
          <fpage>96</fpage>
          ) of
          <source>EasternEuropean Journal of Enterprise Technologies</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>31</lpage>
          . DOI:
          <volume>10</volume>
          .15587/
          <fpage>1729</fpage>
          -
          <lpage>4061</lpage>
          .
          <year>2018</year>
          .
          <volume>149596</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Sharonova</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kyrychenko</surname>
          </string-name>
          , I. Gruzdo, G. Tereshchenko,
          <source>Generalized Semantic Analysis Algorithm of Natural Language Texts for Various Functional Style Types, CEUR Workshop Proceedings</source>
          , Vol-
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>16</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Hrytsiv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shestakevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shyyka</surname>
          </string-name>
          ,
          <article-title>Quantitative Parameters of Lucy Montgomery's Literary Style</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          , Vol-
          <volume>2870</volume>
          (
          <year>2021</year>
          )
          <fpage>670</fpage>
          -
          <lpage>684</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <article-title>Authorship and Style Attribution by Statistical Methods of Style Differentiation on the Phonological Level</article-title>
          ,
          <source>Advances in Intelligent Systems and Computing</source>
          <volume>871</volume>
          (
          <year>2019</year>
          )
          <fpage>105</fpage>
          -
          <lpage>118</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -01069-
          <issue>0</issue>
          _
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Holovatyy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Morushko</surname>
          </string-name>
          ,
          <article-title>Development of Methods, Models and Means for the Author Attribution of a Text</article-title>
          .
          <source>Eastern-European Journal of Enterprise Technologies</source>
          <volume>3</volume>
          /2 (
          <issue>93</issue>
          ) (
          <year>2018</year>
          )
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <article-title>The Method of Statistical Analysis of the Scientific, Colloquial, BellesLettres and Newspaper Styles on the Phonological Level</article-title>
          .
          <source>Advances in Intelligent Systems and Computing</source>
          ,
          <volume>512</volume>
          (
          <year>2017</year>
          )
          <fpage>149</fpage>
          -
          <lpage>163</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -45991-2_
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <article-title>Specifics of Phonostatistical Structure of the Scientific Style in English Style System</article-title>
          ,
          <source>in Proceedings of the XIth Scientific and Technical Conference on CSIT, Lviv</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>129</fpage>
          -
          <lpage>131</lpage>
          . doi:
          <volume>10</volume>
          .1109/stc-csit.
          <year>2016</year>
          .
          <volume>7589887</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Buk</surname>
          </string-name>
          , Osnovy statystychnoi lingvistyky,
          <source>Lviv</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Berko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matseliukh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ivaniv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chyrun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Schuchmann</surname>
          </string-name>
          ,
          <article-title>The text classification based on Big Data analysis for keyword definition using stemming</article-title>
          ,
          <source>in: proceedings of IEEE 16th International conference on computer science and information technologies</source>
          , Lviv, Ukraine,
          <fpage>22</fpage>
          -
          <lpage>25</lpage>
          September,
          <year>2021</year>
          , pp.
          <fpage>184</fpage>
          -
          <lpage>188</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>O.</given-names>
            <surname>Hladun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Berko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bublyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chyrun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Schuchmann</surname>
          </string-name>
          ,
          <article-title>Intelligent system for film script formation based on artbook text and Big Data analysis in:</article-title>
          <source>proceedings of IEEE 16th International conference on computer science and information technologies</source>
          , Lviv, Ukraine,
          <fpage>22</fpage>
          -
          <lpage>25</lpage>
          September,
          <year>2021</year>
          , pp.
          <fpage>138</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Perebyinis</surname>
          </string-name>
          ,
          <article-title>Matematychna linhvistyka</article-title>
          .
          <source>Ukrainska mova. Kyiv</source>
          ,
          <year>2000</year>
          ,
          <fpage>287</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>V.</given-names>
            <surname>Perebyinis</surname>
          </string-name>
          ,
          <article-title>Statystychni metody dlia linhvistiv</article-title>
          .
          <source>Vinnytsia</source>
          ,
          <volume>176</volume>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dyriv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Andrunyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Burov</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Karpov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chyrun</surname>
          </string-name>
          ,
          <article-title>The user's psychological state identification based on Big Data analysis for person's electronic diary</article-title>
          ,
          <source>in: proceedings of IEEE 16th International conference on computer science and information technologies</source>
          , Lviv, Ukraine,
          <fpage>22</fpage>
          -
          <lpage>25</lpage>
          September,
          <year>2021</year>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zhyhalo</surname>
          </string-name>
          ,
          <article-title>Pidkhid do rishennia problem poshuku dvomovnoho plahiatu</article-title>
          .
          <source>Problemy informatyzatsii ta upravlinnia</source>
          <volume>2</volume>
          (
          <issue>24</issue>
          ), (
          <year>2008</year>
          )
          <fpage>125</fpage>
          -
          <lpage>129</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Oborska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Andrunyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chyrun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hasko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vysotskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mushasta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Petruchenko</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Shakleina</surname>
          </string-name>
          ,
          <article-title>The Intelligent System Development for Psychological Analysis of the Person's Condition</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          , Vol-
          <volume>2870</volume>
          (
          <year>2021</year>
          )
          <fpage>1390</fpage>
          -
          <lpage>1419</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Victana</surname>
          </string-name>
          . URL: http://victana.lviv.ua/nlp/linhvometriia
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Boyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dolamic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Grabar</surname>
          </string-name>
          ,
          <source>Automated Detection of Health Websites' HONcode Conformity: Can N-gram Tokenization Replace Stemming? Studies in Health Technology and Informatics</source>
          <volume>216</volume>
          (
          <year>2015</year>
          )
          <fpage>1064</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dmytriv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Holoshchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chyrun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Holoshchuk</surname>
          </string-name>
          ,
          <source>Comparative Analysis of Using Different Parts of Speech in the Ukrainian Texts Based on Stylistic Approach, CEUR Workshop Proceedings</source>
          , Vol-
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>546</fpage>
          -
          <lpage>560</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kubinska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Holoshchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Holoshchuk</surname>
          </string-name>
          , L. Chyrun,
          <article-title>Ukrainian Language Chatbot for Sentiment Analysis and User Interests Recognition based on Data Mining</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          , Vol-
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>315</fpage>
          -
          <lpage>327</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kholodna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Markiv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chyrun</surname>
          </string-name>
          ,
          <source>Machine Learning Model for Paraphrases Detection Based on Text Content Pair Binary Classification, CEUR Workshop Proceedings, Vol3312</source>
          (
          <year>2022</year>
          )
          <fpage>283</fpage>
          -
          <lpage>306</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>V.</given-names>
            <surname>Hryhorovych</surname>
          </string-name>
          ,
          <article-title>Analysis of Scientific Texts by Semantic Inverse-Additive Metrics for Ontology Concepts</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          , Vol-
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>801</fpage>
          -
          <lpage>816</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Albota</surname>
          </string-name>
          ,
          <article-title>Modelling the Impact of the Pandemic on Online Communication: Textual Semantic Analysis</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          , Vol-
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>471</fpage>
          -
          <lpage>486</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kunanets</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Oliinyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Myhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shunevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rzheuskyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shcherbyna</surname>
          </string-name>
          ,
          <article-title>Enhanced LSA Method with Ukraine Language Support</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>2870</volume>
          (
          <year>2021</year>
          )
          <fpage>129</fpage>
          -
          <lpage>140</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mobasher</surname>
          </string-name>
          ,
          <article-title>Data mining for web personalization</article-title>
          .
          <source>The adaptive web</source>
          , (
          <year>2007</year>
          )
          <fpage>90</fpage>
          -
          <lpage>135</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>540</fpage>
          -72079-
          <issue>9</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Dinucă</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ciobanu</surname>
          </string-name>
          , Web Content Mining.
          <source>Annals of the University of Petroşani. Economics</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ) (
          <year>2012</year>
          )
          <fpage>85</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>G.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          , Web Content Mining.
          <source>Web Mining and Social Networking</source>
          (
          <year>2010</year>
          )
          <fpage>71</fpage>
          -
          <lpage>87</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-1-
          <fpage>4419</fpage>
          -7735-
          <issue>9</issue>
          _
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          , Modelling of Phonostatistical Structures of English Backlingual Phoneme Group in Style System,
          <source>in CADMS : Proceedings of the 14th International Conference. Polyana</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>324</fpage>
          -
          <lpage>327</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <article-title>Modelling of Phonostatistical Structures of the Colloquial and Newspaper Styles in English Sonorant Phoneme Group</article-title>
          ,
          <source>in CSIT : Proceedings of the XIIth Scientific and Technical Conference. Lviv</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>67</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <article-title>Authorship Attribution by Differentiation of Phonostatistical Structures of Styles, in CSIT : Proceedings of the XIIIth Scientific</article-title>
          and Technical Conference. Lviv,
          <year>2018</year>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Teslyuk</surname>
          </string-name>
          ,
          <source>The Software for Authorship and Style Attribution in CADMS : Proceedings of the 15th International Conference. Polyana</source>
          ,
          <year>2019</year>
          ,pp.
          <fpage>23</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>I.</given-names>
            <surname>Khomytska</surname>
          </string-name>
          , V. Teslyuk,
          <article-title>Mathematical Methods Applied for Authorship Attribution on the Phonological Level</article-title>
          ,
          <source>in CSIT : Proceedings of the XIVth Scientific and Technical Conference. Lviv</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>V.</given-names>
            <surname>Vysotska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Markiv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Teslia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Romanova</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Pihulechko</surname>
          </string-name>
          ,
          <article-title>Correlation Analysis of Text Author Identification Results Based on N-Grams Frequency Distribution in Ukrainian Scientific and Technical Articles</article-title>
          .
          <source>In CEUR Workshop Proceedings</source>
          , Vol-
          <volume>3171</volume>
          (
          <year>2022</year>
          )
          <fpage>277</fpage>
          -
          <lpage>314</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>C.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Torvik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schnaars</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. Zhang,</surname>
          </string-name>
          <article-title>Examining scientific writing styles from the perspective of linguistic complexity</article-title>
          ,
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>70</volume>
          (
          <issue>5</issue>
          ) (
          <year>2019</year>
          )
          <fpage>462</fpage>
          -
          <lpage>475</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. Zhang,</surname>
          </string-name>
          <article-title>Exploring linguistic characteristics of highly browsed and downloaded academic articles</article-title>
          .
          <source>Scientometrics</source>
          <volume>122</volume>
          (
          <year>2020</year>
          )
          <fpage>1769</fpage>
          -
          <lpage>1790</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mitsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Repariuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sharkan</surname>
          </string-name>
          ,
          <article-title>Identification of authorship of Ukrainian-language texts of journalistic style using neural networks</article-title>
          ,
          <source>Eastern-European Journal of Enterprise Technologies</source>
          <volume>1</volume>
          (
          <issue>2</issue>
          (
          <issue>103</issue>
          )) (
          <year>2020</year>
          )
          <fpage>30</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .15587/
          <fpage>1729</fpage>
          -
          <lpage>4061</lpage>
          .
          <year>2020</year>
          .
          <volume>195041</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>