<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>International Journal of Information Management
64 (2022) 102474.
[21] B.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/CSIT56902.2022.10000563</article-id>
      <title-group>
        <article-title>Cluster Analysis of Discussions Change Dynamics on Twitter about War in Ukraine</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sofiia Mainych</string-name>
          <email>sofiia.mainych.sa.2020@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alina Bulhakova</string-name>
          <email>alina.bulhakova.sa.2020@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victoria Vysotska</string-name>
          <email>victoria.a.vysotska@lpnu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>S. Bandera Street, 12, Lviv, 79013</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Osnabrück University</institution>
          ,
          <addr-line>Friedrich-Janssen-Str. 1, Osnabrück, 49076</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>3</volume>
      <fpage>93</fpage>
      <lpage>98</lpage>
      <abstract>
        <p>The article analyzes the dynamics of changes in the pace and directions of discussion of the russian-Ukrainian war on Twitter based on a set of thematic hashtags from February 23, 2022. A cluster analysis of popular tweets was conducted. The object of the study is the analysis of the discussion of the russian-Ukrainian war on Twitter. Data on the number of tweets and hashtags related to this topic of discussion from the beginning of the war were used for the analysis. The largest share of all hashtags was #Ukraine, followed by #NATO and #UkrainerussiaWar, respectively. At the beginning of the war (the first week), information about the situation in Ukraine became a kind of explosion in the media space of the world. The number of comments, posts, "tweets" on this topic reached a record high of 4,000,000. After a week, this trend changed dramatically - the number of discussions decreased significantly. This is due in particular to the fact that the crisis stage has passed (there has been a lot of information, it is no longer so "interesting" for the world). Although it is also worth saying that this decrease continued to a specific mark. For 50 days from the start of the invasion, the number of publications decreased from 4,000 thousand to 500 thousand. At the level of 500 thousand, the number is maintained until now (with insignificant deviations during certain military events). We can assume that under constant circumstances, without significant changes, this trend will take on a downward trend. At the same time, if the circumstances change, it is impossible to make objective assumptions. Another interesting conclusion can be that during the cluster analysis, it was found that the world focuses more on Ukraine, as a victim of the war, than on russia, as the aggressor. Also, a large share of foreigners supports the initiative to help Ukraine (in particular, the issue of joining NATO). On the contrary, it should not be forgotten that the assumptions "support" - "do not support" are not objective, since people tagging the NATO hashtag can also write negative posts. A fair statement would be the following: foreigners are concerned about the situation in Ukraine and its request for protection from NATO.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Хештег</kwd>
        <kwd>пост</kwd>
        <kwd>Twitter</kwd>
        <kwd>#Ukraine</kwd>
        <kwd>#NATO</kwd>
        <kwd>#UkrainerussiaWar</kwd>
        <kwd>кластерний аналіз</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Continuous consumption of information content from the Internet over the last year has already
become a new daily habit for most Ukrainians, especially from YouTube and various social networks
such as Telegram, Twitter, Facebook and others [1].</p>
      <p>The Kyiv International Institute of Sociology conducted a survey commissioned by the OPORA
Civic Network from May 19 to 24, 2022, among 2,009 adult citizens of Ukraine who at that time were
in the territory of Ukraine controlled by the Ukrainian authorities [2]. The results showed the following
distribution of popularity of sources of operational news among surveyed Ukrainians: 76.6% - social
networks (77.9% of men and 75.5% of women), 66.7% – television (predominance among women
70.4%), 61.2% – Internet sources without social networks (predominance among men – 63.5%), 28.4%
– radio and 15.7% – print mass media. Young people generally get their news from social media (92%),
compared to the older, middle-aged generation (64%). The latter also prefer the Internet without social
networks (60%). Older people prefer television (78%), radio (36%) and sometimes print media (20%).</p>
      <p>Rural residents are more likely to watch television (71.3% versus 64.4% in cities) and read print
media (21.5% versus 12.7%). On the other hand, townspeople more often consume news via the Internet
(64% vs. 55.6% in the village), social networks (79.2% vs. 71.4%), and radio (28.5% vs. 28.2%) [3].</p>
      <p>In Western Ukraine, more people watch television (73.5%), read print media (23.2%) and listen to
the radio (34.6%). On the other hand, in eastern Ukraine, internet media are read the most (63.2%).
Social networks as a source of news are most actively used in southern Ukraine (77.8%) [1-3].</p>
      <p>The main question to be investigated is the verification of the trend of interest regarding the situation
in Ukraine in the world (by time). First of all, we can put forward the following hypothesis: "At the
beginning of the war (the first few days-weeks), the number of publications (tweets) on the Internet
increased rapidly, after that it remained at approximately the same level for a certain time, later (let's
say a month later) interest is still higher, than it was before the reference point (the beginning of the
war), and yet it has already significantly decreased compared to the maximum value." The presented
hypothesis is based on completely logical ideas about the state of the media when riots, wars, and coups
suddenly begin. The news, which originated in one media space, spreads around the world with
extraordinary speed. People are interested in knowing details and predictions, both true and fake. Later,
when the state of affect passes and the world is already better informed about the situation, this interest
gradually subsides. But since this hypothesis is still purely empirical, a study was conducted, the results
of which are presented in this report. The purpose of this work is to verify the proposed hypothesis
using data analysis from the selected dataset on the topic "Discussion of the russian-Ukrainian war on</p>
      <sec id="sec-1-1">
        <title>Twitter".</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>After the escalation of russia's war with Ukraine on February 24, 2022, the top 3 social networks,
where Ukrainians get information, changed somewhat [1-2]. Yes, YouTube remains the most stable in
the lives of Ukrainians - year after year, starting in 2019, it ranks second in Ukraine as a source of news.
In 2019-2021, the ranking of social networks remained almost unchanged, with the exception of 2019,
when the third most popular messenger was Viber. However, after the start of the full-scale invasion,
Facebook for the first time gave up its positions to Telegram and moved all the way to third place. Thus,
among 76.6% of citizens who use social networks as a source of information, 66% choose Telegram,
61% choose YouTube, and another 58% choose Facebook. Less popular news sources for Ukrainians
during the war were Viber (48%), Instagram (29%), TikTok (19.5%) and Twitter (8.9%).</p>
      <p>The war exposed russian propaganda and changed the perception of Ukraine in the world, says Nina
Yankovich, vice-president of the "Center for Information Resilience" project [3]. russian disinformation
has played a crucial role in the war against Ukraine since 2014, although this role has changed
somewhat. Now we have become more aware of russia's intentions in Ukraine. Starting from 2022, the
russian information war in Ukraine is actually failing. Ukraine has shown the world its extraordinary
resilience, amazing creativity and authenticity in communications. The aggressor, russia, has become
less sophisticated in disseminating its disinformation, as social networks have learned to counter bots
and trolls. Therefore, there is less of such an artificial increase in informational emissions now. When
the full-scale invasion first began, people had a lot of questions about Ukraine. American publications
covered this conflict qualitatively: almost all of them now have correspondents in Ukraine. Ukraine's
constant presence in the international media has been incredibly important in filling the gaps that many
Americans have. There is a false impression that fighting disinformation means stifling freedom of
speech, as imposing restrictions on expression. This is not always true.</p>
      <p>Social media is an important public platform for many democracies and autocracies around the
world, and without it, people resisting authoritarian regimes would have no voice [3]. Therefore, it is
very important that these multi-billion dollar corporations invest in people who actually understand the
context of what is happening in the country and develop regional offices. If Twitter, for example, claims
to be a voice for free speech and democracy, then it should support the democratic side of this conflict,
which is fighting for freedom. Also, if we're talking about content sharing and freedom of speech, social
networks should definitely allow users to share and document what they experience in their daily lives.</p>
      <sec id="sec-2-1">
        <title>And any platform that doesn't, in my opinion, violates freedom of expression.</title>
        <p>There is a huge interest in what is happening in Ukraine, and authentic content from "real people" is
one of the most effective ways to fight disinformation [1-3].</p>
        <p>Many Ukrainians are rapidly gaining hundreds of thousands of followers by telling stories about
Ukraine. There is a huge interest in what is happening in Ukraine, and authentic content from "real
people" is one of the most effective ways to combat disinformation. Every well-told story is the perfect
antidote to propaganda. Because it's not just fact-checking or myth-busting, it's the real experience of
ordinary people. And it allows people living in the West with their simple, predictable lives to put
themselves in the shoes of people living in Ukraine, and it's extremely powerful. It is now much more
difficult to create a fake account and impersonate someone. Social networks have found ways to
recognize such accounts. And Americans have also become more savvy: they show this behavior more
often. It is still an actual threat. Aggressor russia finds ways to manipulate our discourse on the Internet.</p>
      </sec>
      <sec id="sec-2-2">
        <title>But this threat is not as straightforward as before.</title>
        <p>American businessman, founder of SpaceX and CEO of Tesla, Elon Musk, reported on May 18 that,
according to his calculations, 50% of Twitter accounts are bots [4-7]. Twitter's new identity verification
system has led to a boom in fakes. Perhaps Elon Musk is all too aware of how Twitter can influence
politics, politicians and political debates around the world, and is taking advantage of it. Previously,
Twitter said that only 5% of accounts on the social network are fake. After that, Musk suspended the
deal to buy the company. He also stated that "the deal cannot move forward" unless Twitter provides
evidence that the number of bots does not exceed 5% of daily active users per month during the quarter
[8]. At a conference in Miami, the businessman said that fake users make up at least 20% of all Twitter
accounts, and suggested that the figure could be as high as 90%. He also revealed that his team will
conduct its own audit of 100 random Twitter accounts to identify the number of fake ones. After these
statements, according to Musk, Twitter lawyers accused him of violating the rules of non-disclosure of
information (NDA) [9]. The head of Twitter's security department noted that the new policy is designed
to "help protect discussions on Twitter, starting with a focus on the russian invasion of Ukraine." On
May 19, the Twitter company introduced a new policy "to combat misinformation during a crisis",
which provides that false content about the war in Ukraine will be marked with special marks [10].</p>
        <p>Twitter may add a warning to a post if it contains:
 false information that incorrectly characterizes the conditions on the ground during the
development of the conflict;
 false statements about the use of force, encroachment on territorial sovereignty or the use of
weapons;
 knowingly false or misleading statements about war crimes or mass atrocities against certain
population groups;
 false information about the reaction of the international community, sanctions, defensive
actions or humanitarian operations.</p>
        <p>Under the new policy, the social network will not recommend or share tweets that have been
identified as false. Twitter will also limit the spread of such information: users will not be able to like,
retweet or share content that violates the new rules. This change is part of a broader promotion of
truthful information during a conflict or crisis following the Soviet invasion of Ukraine.</p>
        <p>Twitter's head of security and corporate ethics, Yoel Roth, noted that the new policy is designed to
help protect discussions on Twitter, starting with a focus on the russian invasion of Ukraine [11]. These
rules will focus on potentially dangerous misinformation about alleged war crimes, armed conflicts,
humanitarian crises, etc.</p>
        <p>We will remind, on February 26, the Twitter social network blocked the possibility of registering
accounts in russia [12]. For users from Ukraine and russia, the social network has suspended advertising
and some recommendations for tweets from people they are not following. This was done to reduce the
spread of offensive content. On February 28, Twitter began flagging russian state media. Then, in four
days, the social network recorded more than 45,000 tweets per day with links to russian state media,
but as of March 11, Twitter saw a drop in impressions. On March 11, the social network began flagging
the accounts of state media in Belarus [13]. On March 17, the social network deleted or marked as
unreliable more than 50,000 messages containing false information about russia's war against Ukraine
[14-21]. That is why this social network itself was chosen to analyse the discussion of the
russianUkrainian war on Twitter (the number of hashtags related to this topic from February 23, 2022 - from
the beginning of the full-scale war of russia against Ukraine).</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methods and materials</title>
      <sec id="sec-3-1">
        <title>The following datasets were used for this work:</title>
        <p> general statistics of the daily number of tweets during 23.02-22.09:
https://github.com/alexdrk14/RussiaUkraineWar/blob/main/data/daily_stats.csv (Fig. 1);
 statistics of the use of various hashtags related to the russian-Ukrainian war of 2022:
https://github.com/alexdrk14/RussiaUkraineWar/blob/main/data/daily_hashtags.csv.</p>
        <p>We will process the data in several stages, in particular: graphic representation of trends, calculations
of correlations (relationships between data) and cluster analysis. First of all, in any work with data, it is
necessary to properly organize them for further processing. The first type of presentation of the received
information is a tabular presentation. In addition to tables, graphs are a widely applicable way of
presenting data. With their help, you can visually see certain patterns, establish the type of connection,
and even roughly predict the behavior of a particular process. The main difference from a table is that
in a graph we have to clearly specify some relationship that we want to investigate, unlike a table, which
helps to quickly navigate a large volume of data.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Data for table presentation: number of tweets per day (9 subtables in Fig. 2).</title>
      </sec>
      <sec id="sec-3-3">
        <title>Data for graphic representation: number of tweets per day since the start of the war:</title>
        <sec id="sec-3-3-1">
          <title>Ukrainian war (x is day of war, y is the amount of tweets inthousands</title>
          <p>Formulas for the transition to polar coordinates:
 =</p>
          <p>{



2
3
2
 2 =  2 +  2




( )
( ) + 2


( ) + 
 &gt; 0,  ≥ 0
 &gt; 0,  &lt; 0
 &lt; 0
 = 0,  &gt; 0
 = 0,  &lt; 0</p>
          <p>Data for graphical representation: ratio of hashtags #StopRussian and #StandWithUkraine.</p>
          <p>As we can already see from Fig. 1-4, the number of "references" to the war is much greater at the
beginning of the war and significantly decreases over time. The graph of the Cartesian coordinate
system clearly shows this rapid drop in quantity to a certain number, after which the quantity is
maintained at a roughly constant level. The next important step in research is computation. As already
mentioned, the hypothesis that was put forward has no "numerical" basis, and it is this problem that is
solved by quantitative characteristics. We will use for work (Fig. 5):
</p>
          <p>Mean – arithmetic mean (a measure of central tendency that reflects the most characteristic
value for this sample):
 ̅=

Sd – standard deviation (a measure of variation of a characteristic that reflects the amount of</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Median – median (the value that divides an ordered set of variables in half):</title>
        <p>( ) =
 +1 (or
2
Skew – asymmetry (reflects the skew of the distribution relative to the fashion to the left or</p>
      </sec>
      <sec id="sec-3-5">
        <title>Kurtosis – kurtosis (displays the height of the distribution);</title>
      </sec>
      <sec id="sec-3-6">
        <title>Se – standard error.</title>
        <p>For calculations, we use the R studio environment and its built-in function for constructing
descriptive statistics: describe.
time, feature intervals are plotted on the abscissa axis, and frequencies are plotted on the ordinate axis.</p>
        <p>For construction, it is necessary to find the number of partition intervals. For this, we will use the
Sturges formula  = 1 + 
2 ,  ℎ</p>
        <p>is number of intervals.
intervals.</p>
        <p>Cumulatives are also a way of presenting data. It has a certain similarity with a histogram, but for
its construction it is necessary to calculate the accumulated frequencies. Accumulated frequencies show
how many units of the population have a characteristic value no greater than that under consideration.
You can build a cumulate in several ways, in particular, the number of #UkraineRussiaWar hashtags
throughout the war (Fig.6). Number of split intervals: Sturges formula  = 1 + 
2 ,  is number of</p>
        <p>As can be seen in fig. 6-7, the built cumulates are similar, but the one built by integral percentage is
smoother, so it gives a smaller error.
library(readr) # dataset loading
dataset1 &lt;- read.csv("daily_stats.csv")
dataset2 &lt;- read.csv("daily_hashtags.csv")
dataset &lt;- cbind(dataset1[1], dataset1[3], dataset2[2:11])
n &lt;- dim(dataset)[1] # the number of elements
table &lt;- as.data.frame(cbind(dataset$day, dataset$daily_tweets)) # formation of the report table
m &lt;- dim(table)[2] # the number of stacks of the initial table
t &lt;- 6# the number of subtables into which to split the original table
k &lt;- round(n/t) # determining the number of rows in each subtable
l &lt;- k + 1# row number that will move to the next subtable
a &lt;- m + 1# the number of the first empty column of the initial table
for (j in 1:(t-1)){
for (i in 1:k){
for (q in 0:(m - 1)){ # table compression
#if (table[l, 1 + q] != is.null)
table [i, a + q] &lt;- table[l, 1 + q]
#else
# stop
}
table &lt;- table[-l, ]
}
a &lt;- a + m # determining the number of the next empty column
}
i &lt;- 1# renaming the columns of the created table
while(i &lt;= 2*t){
names(table)[i] = "Date"
names(table)[i + 1] = "Amount of tweets"
i &lt;- i + 2
}# plotting data in the Cartesian coordinate system
plot(1:dim(dataset)[1], dataset$daily_tweets/1000, xlab = "Day of war",
ylab = "The amount of tweets (in thousands)", main = "Dynamics of discussion about russian-Ukranian war",
type = "o", pch = 16, lwd = 3, col = "darkorchid4")
polardataset &lt;- cbind(as.numeric(dataset$StopRussia), as.numeric(dataset$StandWithUkraine)) # plotting data in polar coordinate system
for(i in 1:n){ # creating a table for a polar graph regarding the ratio of hashtags StopRussian and StandWithUkraine
r &lt;- sqrt(polardataset[i, 1]^2 + polardataset[i, 2]^2) # transition to polar coordinates
if (polardataset[i, 1] &gt; 0 &amp;&amp; polardataset[i, 2] &gt;= 0)
f &lt;- atan(polardataset[i, 2]/polardataset[i, 1])*180/pi
if (polardataset[i, 1] &gt; 0 &amp;&amp; polardataset[i, 2] &lt; 0)
f &lt;- atan(polardataset[i, 2]/polardataset[i, 1])*180/pi + 2*pi
if (polardataset[i, 1] &lt; 0)
f &lt;- atan(polardataset[i, 2]/polardataset[i, 1])*180/pi + pi
if (polardataset[i, 1] == 0 &amp;&amp; polardataset[i, 2] &gt; 0)
f &lt;- pi/2
if (polardataset[i, 1] == 0 &amp;&amp; polardataset[i, 2] &lt; 0)
f &lt;- 3*pi/2
polardataset[i, 1] &lt;- r
polardataset[i, 2] &lt;- f
}
library(plotly) # graph construction
plot_ly( type = "scatterpolar", r = polardataset[, 1], theta = polardataset[, 2], mode = 'markers')
library("psych")# construction of descriptive data statistics
statistics &lt;- describe(dataset[2:12])
hist(as.numeric(dataset$Ukraine[7:37]), breaks = round(1 + log2(n)), # histogram construction of #Ukraine hashtags distribution in March
xlab = "Grouping intervals", ylab = "Frequency of values",
col = "seagreen4", main = "Histogram of hastag #Ukraine in March")
k &lt;- round(1 + log2(n)) # construction of cumulate according to histogram data
h1 &lt;- hist(as.numeric(dataset$UkraineRussiaWar), breaks = k, xlab = "Grouping intervals",
ylab = "Frequency of values", col = "mediumpurple2",
main = "Histogram of #UkraineRussiaWar hastags during all war")
plot(h1$mids, cumsum(h1$counts)/n, xlab = "Intervals", ylab = "Frequency (probability)",</p>
        <p>main = "Cumulative plot by histogram", 'l', lwd = 3)
m &lt;- dim(dataset)[1] - 1# construction of the cumulate by integral interest
h2 &lt;- hist(as.numeric(dataset$UkraineRussiaWar), breaks = m, xlab = "Grouping intervals",
ylab = "Frequency of values", col = "mediumpurple2",
main = "Histogram of #UkraineRussiaWar hastags during all war")
plot(h2$mids, cumsum(h2$counts)/n, xlab = "Intervals", ylab = "Frequency (probability)",</p>
        <p>main = "Cumulative plot by integral percentage", 'l', lwd = 3)
4. Experiments, results and discussions</p>
        <p>The study of a numerical series is mainly based on the analysis of its graphical representation. That
is, if we have a graph (a certain line), we can use it to provide conclusions about the state of our initial
data. This works well if the basic patterns are not affected by any other, random, external factors. If
such an influence is monitored, sometimes in order to "clean" the main trend from "obstacles",
smoothing methods are used. Time series for smoothing: number of #NATO hashtags for each day of
the war. The moving average method gives an estimate of the average level for a certain period of time
(the longer the time interval to which the average belongs, the more the level will be smoothed, but the
less accurately the trend of the original series of dynamics will be described).</p>
      </sec>
      <sec id="sec-3-7">
        <title>The built-in function sma() was used in the programming environment.</title>
        <p>The method of the weighted moving average on each active area of the value of the central level is
replaced by the calculated one, which is determined by the formula of the weighted arithmetic average
(the weighting factors are determined using the method of least squares).
library("zoo")# smoothing with a moving average
plot.ts(dataset$NATO,main = "Simple moving average", ylab = "Count of #NATO")
lines(rollmean(dataset$NATO,5),col = 'blue')
lines(rollmean(dataset$NATO,3),col = 'red')
lines(rollmean(dataset$NATO,7),col = 'green')
legend(50,150000,col = c('black','blue', 'red','green'), legend = c('main', 'SMA 5', 'SMA 3','SMA 7'),lty = 1,cex = 0.8)
plot.ts(dataset$NATO,main = "Locally weighted smoothing", ylab = "Count of #NATO")# lowess smoothing curve
lines(lowess(dataset$NATO,f = 0.5),col = "blue")
lines(lowess(dataset$NATO,f = 0.05),col = "red")
legend(50,150000,col = c('black','blue','red'), legend = c('main', 'LOWESS 0.5', 'LOWESS 0.05'),lty = 1,cex = 0.8)</p>
        <p>When applying the method of moving averages, the selection of the smoothing interval value should
be made on the basis of meaningful considerations and be tied to the period of possibly existing
oscillatory processes. If the moving average procedure is used to smooth the time series in the absence
of any fluctuations, then most often the value of the smoothing interval is chosen equal to three, five or
seven. The larger the averaging interval, the smoother the trend graph looks. Time series for smoothing:
number of #UkraineRussiaWar hashtags for each day of the war (in thousands). Linear smoothing for
w = 3:  ̅1 =
5 1+2 2− 3
6
;  ̅ =   −1+  +  +1 ,  ̅ =
3
−  −2+2  −1+5  ,  = 2, 3, … ,  − 1.</p>
        <p>6
makes it possible to compare the indicators obtained for different objects;
linear transformation, which consists in the fact that the values of the levels of the time series
lead to the interval of values [0,1] according to the formula:  н =
the time series.
normalized value,   is level value,  
and</p>
        <p>−  
 
−</p>
        <p>, where  н is
– the smallest and largest value of the levels of
norm_sequence &lt;- function(x){ # normalization of time sequences
x_norm &lt;- vector()
for (i in 1:length(x))
return(x_norm)
x_norm[i] &lt;- (x[i] - min(x))/(max(x) - min(x))</p>
        <p>Time series smoothing effectiveness criteria;



}
}
}
}
}
}
}
t &lt;- 0
for (i in 2:(n - 1)){</p>
        <p>t &lt;- t + 1
return(t)
v &lt;- vector()
for (i in 2:(n - 1)){</p>
        <p>v &lt;- cbind(v, i)
return(x[v])
S1 &lt;- 0
S2 &lt;- 0
S3 &lt;- 0
S4 &lt;- 0
S5 &lt;- 0
for (i in 1:n){
S1 &lt;- S1 + x[i]
S2 &lt;- S2 + y[i]
S3 &lt;- S3 + x[i]*y[i]
S4 &lt;- S4 + (x[i])^2
S5 &lt;- S5 + (y[i])^2
</p>
        <p>Criterion of turning points: ( ( 3 &gt;  2)&amp;&amp; ( 3 &gt;  4) ) || ( ( 3 &lt;  2)&amp;&amp; ( 3 &lt;  4) ).
Correlation coefficient:  
=</p>
        <p>∑ =1     − ∑ =1   ∑
.
turn_points &lt;- function(x, n){ # the search function for the number of turning points</p>
        <p>if ((x[i] &gt; x[i - 1]&amp;&amp; x[i] &gt; x[i + 1]) || (x[i] &lt; x[i - 1] &amp;&amp; x[i] &lt; x[i + 1]))
turn_points_value &lt;- function(x, n){ # turning point search function</p>
        <p>if ((x[i] &gt; x[i - 1]&amp;&amp; x[i] &gt; x[i + 1]) || (x[i] &lt; x[i - 1] &amp;&amp; x[i] &lt; x[i + 1]))
correlation_coeficient &lt;- function(x, y, n){ # correlation coefficient determination function
return((n*S3 - S1*S2)/(sqrt((n*S4 - (S1)^2)*(n*S5 - (S2)^2))))</p>
        <p>Smoothing according to Kendel's formulas is a type of weighted moving average. Made smoothing
according to Kendel's formulas for various w (3, 5, 7, 9, 11, 13, 15) for #NATO hashtags. According
to fig. 12, we can conclude that the property of linear smoothing (the larger the coefficient, the stronger
the graph is smoothed) is also true for smoothing according to Kendel's formulas. Correlation analysis
refers to a set of methods that make it possible to detect the presence and degree of relationship between
several randomly changing parameters (more details about correlation analysis will be found in clause
4, at this stage we are only interested in such concepts as "correlation table" (table of ratios) and turning
points (levels whose values are greater or less than two adjacent ones)).</p>
        <p>Determination of efficiency criteria for performed smoothing according to formulas from Kendel:</p>
        <p>Conclusion: the larger the parameter, the fewer turning points, the smaller the correlation coefficient.
Therefore, the more "smoothed" is the resulting graph (for ordinary smoothing according to Kendel's
formulas).</p>
        <p>Re-smoothing is when we apply the data obtained in the previous layer to the input data of the next
layer. Re-smoothing according to Kendel's formulas for different w (3, 5, 7, 9, 11, 15):</p>
        <p>Construction of a correlation table for all intervals of re-smoothing according to formulas from</p>
      </sec>
      <sec id="sec-3-8">
        <title>Kendel:</title>
        <p>Construction of turning point diagrams for all re-smoothing intervals according to Kendel's
formulas:</p>
        <p>Determination of efficiency criteria for repeated smoothing according to Kendel's formulas:</p>
        <p>Number of turning points Correlation coefficient</p>
        <p>Compared to the results of "no re-smoothing", with re-smoothing, the number of turning points
decreases more rapidly, as does the correlation coefficient. Here already on parameter "11" the number
of turning points was 7, when there - 12 (and the smallest value is 9, when here - 5).
cat("\t\tKendall smoothing")# smoothing according to formulas from Kendel
library(irr)
kSmooth1 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=3)
kSmooth2 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=5)
kSmooth3 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=7)
kSmooth4 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=9 )
kSmooth5 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=11)
kSmooth6 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=13)
kSmooth7 &lt;- ksmooth(time(dataset$NATO), dataset$NATO, 'normal', bandwidth=15 )
plot.ts(dataset$NATO, main = 'Kendall Smoothing', ylab = '#NATO') # image of smoothing according to Kendel's formulas on the graph
lines(dataset$NATO, col = '1')
lines(kSmooth1, type = 'l', col = '2')
lines(kSmooth2, type = 'l', col = '3')
lines(kSmooth3, type = 'l', col = '4')
lines(kSmooth4, type = 'l', col = '5')
lines(kSmooth5, type = 'l', col = '6')
lines(kSmooth6, type = 'l', col = '7')
lines(kSmooth7, type = 'l', col = '8')
legend(180,150000, col = c('1','2','3','4','5','6','7','8'),</p>
        <p>legend = c('original', 'KSM3', 'KSM5','KSM7', 'KSM9','KSM11', 'KSM13','KSM15'),lty = 1, cex = 0.8)
kendallSmoothing &lt;- cbind(dataset$NATO, kSmooth1[[2]], kSmooth2[[2]], kSmooth3[[2]], kSmooth4[[2]],</p>
        <p>kSmooth5[[2]],kSmooth6[[2]], kSmooth7[[2]]) # construction of a generalized correlation table
colnames(kendallSmoothing) &lt;- c("Y", "Yn, w = 3", "Yn, w = 5", "Yn, w = 7", "Yn, w = 9", "Yn, w = 11", "Yn, w = 13", "Yn, w = 15")
View(kendallSmoothing)
k_t &lt;- vector()# the number of turning points when smoothing according to formulas from Kendel
for (i in 1:7){
k_t[i] &lt;- turn_points(kendallSmoothing[, i + 1], dim(kendallSmoothing)[1])
}# construction of diagrams of turning points for smoothing according to formulas from Kendel
barplot(turn_points_value(kSmooth1[[2]], length(kSmooth1[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[1]), main = "Turned points diagram, w = 3")
barplot(turn_points_value(kSmooth2[[2]], length(kSmooth2[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[2]), main = "Turned points diagram, w = 5")
barplot(turn_points_value(kSmooth3[[2]], length(kSmooth3[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[3]), main = "Turned points diagram, w = 7")
barplot(turn_points_value(kSmooth4[[2]], length(kSmooth4[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[4]), main = "Turned points diagram, w = 9")
barplot(turn_points_value(kSmooth5[[2]], length(kSmooth5[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[5]), main = "Turned points diagram, w = 11")
barplot(turn_points_value(kSmooth6[[2]], length(kSmooth6[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[6]), main = "Turned points diagram, w = 13")
barplot(turn_points_value(kSmooth7[[2]], length(kSmooth7[[2]])), col = "lightsalmon",</p>
        <p>names.arg = c(1:k_t[7]), main = "Turned points diagram, w = 15")
k_r_xy &lt;- vector()# correlation coefficients when smoothing according to formulas from Kendel
for (i in 1:7){
k_r_xy[i] &lt;- correlation_coeficient(dataset$NATO, kendallSmoothing[, i + 1], dim(kendallSmoothing)[1])
}
k &lt;- c(3, 5, 7, 9, 11, 13, 15) # displaying on the screen the number of turning points and correlation coefficients at different w
kndl_matrix &lt;- cbind(k, k_t, k_r_xy)
colnames(kndl_matrix) &lt;- c("Smoothing parameter", "Amount of turned points", "Correlation coefficient")
cat("\nKendall smoothing criteria:\n")
print(kndl_matrix)
cat("\n\t\tRepeated Kendall smoothing")# re-smoothing according to formulas from Kendel
kSmooth1 &lt;- ksmooth(time(dataset$NATO), dataset$NATO,'normal',bandwidth = 3)
kSmooth2 &lt;- ksmooth(time(kSmooth1[[2]]), kSmooth1[[2]], 'normal', bandwidth = 5)
kSmooth3 &lt;- ksmooth(time(kSmooth2[[2]]), kSmooth2[[2]], 'normal', bandwidth = 7)
kSmooth4 &lt;- ksmooth(time(kSmooth3[[2]]), kSmooth3[[2]], 'normal', bandwidth = 9)
kSmooth5 &lt;- ksmooth(time(kSmooth4[[2]]), kSmooth4[[2]], 'normal', bandwidth = 11)
kSmooth6 &lt;- ksmooth(time(kSmooth5[[2]]), kSmooth5[[2]], 'normal', bandwidth = 13)
kSmooth7 &lt;- ksmooth(time(kSmooth6[[2]]), kSmooth6[[2]], 'normal', bandwidth = 15)
plot.ts(dataset$NATO, main = "Repeated Kendall Smoothing", ylab = '#NATO') # image of re-smoothing using Kendel's formulas on a graph
lines(dataset$NATO,col='1')
lines(kSmooth1,type='l',col='2')
lines(kSmooth2,type='l',col='3')
lines(kSmooth3,type='l',col='4')
lines(kSmooth4,type='l',col='5')
lines(kSmooth5,type='l',col='6')
lines(kSmooth6,type='l',col='7')
lines(kSmooth7,type='l',col='8')
legend(180,150000, col = c('1','2','3','4','5','6','7','8'),</p>
        <p>legend = c('original', 'KSM3', 'KSM5','KSM7', 'KSM9','KSM11', 'KSM13','KSM15'), lty = 1, cex = 0.8)
kendallSmoothing &lt;- cbind(dataset$NATO,kSmooth1[[2]], kSmooth2[[2]], kSmooth3[[2]], # construction of a generalized correlation table
kSmooth4[[2]],kSmooth5[[2]], kSmooth6[[2]], kSmooth7[[2]])
colnames(kendallSmoothing) &lt;- c("Y", "Y'n, w = 3", "Y'n, w = 5", "Y'n, w = 7", "Y'n, w = 9", "Y'n, w = 11", "Y'n, w = 13", "Y'n, w = 15")
View(kendallSmoothing)
for (i in 1:7){
rep_k_t &lt;- vector()# the number of turning points during re-smoothing according to Kendel's formulas
rep_k_t[i] &lt;- turn_points(kendallSmoothing[, i + 1], dim(kendallSmoothing)[1])
} # construction of diagrams of turning points for re-smoothing according to formulas from Kendel
barplot(turn_points_value(kSmooth1[[2]], length(kSmooth1[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[1]), main = "Turned points diagram, w = 3")
barplot(turn_points_value(kSmooth2[[2]], length(kSmooth2[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[2]), main = "Turned points diagram, w = 5")
barplot(turn_points_value(kSmooth3[[2]], length(kSmooth3[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[3]), main = "Turned points diagram, w = 7")
barplot(turn_points_value(kSmooth4[[2]], length(kSmooth4[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[4]), main = "Turned points diagram, w = 9")
barplot(turn_points_value(kSmooth5[[2]], length(kSmooth5[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[5]), main = "Turned points diagram, w = 11")
barplot(turn_points_value(kSmooth6[[2]], length(kSmooth6[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[6]), main = "Turned points diagram, w = 13")
barplot(turn_points_value(kSmooth7[[2]], length(kSmooth7[[2]])), col = "lightblue",</p>
        <p>names.arg = c(1:rep_k_t[7]), main = "Turned points diagram, w = 15")
for (i in 1:7){
}
rep_k_r_xy &lt;- vector()# correlation coefficients with repeated smoothing according to Kendel's formulas
rep_k_r_xy[i] &lt;- correlation_coeficient(dataset$NATO, kendallSmoothing[, i + 1], dim(kendallSmoothing)[1])
k &lt;- c(3, 5, 7, 9, 11, 13, 15) # displaying on the screen the number of turning points and correlation coefficients at different w
rep_kndl_matrix &lt;- cbind(k, rep_k_t, rep_k_r_xy)
cat("\nRepeated Kendall smoothing criteria:\n")
print(rep_kndl_matrix)
colnames(rep_kndl_matrix) &lt;- c("Smoothing parameter", "Amount of turned points", "Correlation coefficient")</p>
        <p>Correlation analysis of time series is used when it is necessary to assess the presence and strength
of the relationship between certain indicators. In our case, the total number of tweets for each day since
the beginning of the war was taken as a series for analysis (we will present the data in the form of a
table). The correlation field is a graphical representation of the relationship between the two
investigated sequences (in our case, time and quantity). Correlation coefficient: an indicator of the
quantitative assessment of the tightness of the connection; in the range from -1 to 1; characterizes a
linear relationship, where when one value increases, the other increases (decreases)..</p>
        <p>The following formula is used for calculations:  
=</p>
        <p>∑ =1     − ∑ =1   ∑</p>
        <p>Since the correlation coefficient is negative, we can conclude that the dependence is decreasing, that
is, as the value of one value increases, the other decreases. In our case, over time, the number of hashtags
decreases (activity and interest in the problem, respectively). Correlation relation:
used in the study of nonlinear dependencies;
limits: [0,1], and if is strictly equal to one, then there is an unambiguous functional dependence
between the values.</p>
      </sec>
      <sec id="sec-3-9">
        <title>Algorithm for finding a correlation relation:</title>
        <p>Division of the correlation field by variable X into L grouping intervals of different lengths.
2. Search for "partial" mathematical expectations Y in each of the L selected groups  ̅  
=
 =1   ,  = ̅1̅̅,̅̅,  = ̅1̅̅,̅̅̅,   is number of sample elements in  th grouping intervals.</p>
        <p>Mathematical expectation search of partial groupings of reviews using  ̅  =</p>
        <p>2
4. Calculation of group variance for variable y:   ̅ 
=
 =1   ∗ ( ̅  
−  ̅  ) .
5. Calculation of variance obtained from ungrouped response  ̅ 2 =

1 ∑
 =1(  −  ̅  )2.
6. Construction of the correlation relation  ̅∗
=
̅̅̅̅ .
̅

correlation relation</p>
        <p>Properties of the correlation relation ̅∗</p>
        <p>≥ 0, 0 ≤  ̅∗ ≤ 1,   ∗ ≥ |  ∗ |. The obtained value
belongs to the space of permissible values for this indicator. 0.948 is a number close to 1, but not 1.
Conclusion: there is a strong relationship between the values, but not an unambiguous dependence.</p>
      </sec>
      <sec id="sec-3-10">
        <title>Correlation matrix:</title>
        <p>

a table containing correlation coefficients when analyzing a large number of observations;
a square table in which the correlation coefficient between the corresponding parameters is
located at the intersection of the corresponding row and column.
Multiple correlation coefficients  
= √
  2
+  2 +2</p>
        <p>1−  2
 .</p>
        <p>Calculation of autocorrelation is an indicator that characterizes the existence of dependence between
the previous and next levels of the time sequence and is calculated according to the formula:
1
 ( ) =  −
∑ =−1 (  −  ̅)∗(  + −  ̅)
1</p>
        <p>−1 ∑ =1(  − ̅)2
.</p>
      </sec>
      <sec id="sec-3-11">
        <title>Graphing the autocorrelation function</title>
        <p>Looking at the time series graph, we can conclude that the data has a downward trend. Therefore, it
is possible to assume the non-stationarity of the original time series.</p>
        <p>To more accurately determine the stationarity of the series, the correlogram of the autocorrelation
function is analyzed. In the case of a stationary time series, a rapid decline with increasing t will be
depicted already after the first few values. The constructed correlogram demonstrates that the studied
series is not stationary, but contains a trend component.
cat("\t\tCorrelation analysis\n")# correlational analysis
cor_field &lt;- function(x, y, a){ # construction of the correlation field
}
correlation_ratio &lt;- function(x, y, n){ # function for calculating the correlation ratio
m &lt;- 15 # number of separation intervals
field &lt;- plot(x, y, xlab = "Factor charasteristics X", ylab = "Resulting characteristics Y", main=a, pch=16, lwd=2, type="b", col="darkorchid4")
L &lt;- matrix(NA, m, 3) # division of the correlation field by variable X into 15 parts of different lengths
k &lt;- round(n/m)
for (i in 1:m){ # a random number from a range [10, 15]
if (a != x[n]){
if (i %% 2 == 0){
L[i, 1] &lt;- a
if(round(a + 0.7*k) &lt;= x[n])# checking not to go beyond the initial X interval
L[i, 2] &lt;- round(a + 0.7*k)
else
L[i, 2] &lt;- x[n]
a &lt;- L[i, 2] + 1
}
for(i in 1:m){ L[i, 3] &lt;- L[i, 2]-L[i, 1] + 1 }# determination of interval lengths
colnames(L) &lt;- c("The beginning", "The end", "Length")
cat("Created intervals:\n") # displaying the formed intervals on the screen
print(L)
m_y_j &lt;- vector() # finding partial mathematical expectations for each interval
for (i in 1:m){
S &lt;- 0
for(j in L[i, 1]:L[i, 2]){ S &lt;- S + y[j] }
m_y_j[i] &lt;- S/L[i, 3]
}
mult_cor_coefficient &lt;- function(r, z, x, y){ # the function of determining the multiple correlation coefficient
for (i in 1:3){ # z dependent variable, х, у - independent variables
R_zxy &lt;- sqrt((r[x, z]^2 + r[y, z]^2 - 2 * r[x, z] * r[y, z] * r[x, y])/(1 - r[x, y]^2))
}
return(R_zxy)
}
auto_cor_coeficient &lt;- function(y, t, n){ # autocorrelation coefficient determination function
S1 &lt;- 0
S2 &lt;- 0
for (i in 1:(n - t)){ S1 &lt;- S1 + (y[i] - mean(y)) * (y[i + t] - mean(y)) }
for (i in 1:n){ S2 &lt;- S2 + (y[i] - mean(y))^2 }
r &lt;- ( (1/(n - t)) * S1 ) / ( (1/(n - 1)) * S2 )
return(r)
}
table &lt;- cbind(c(1:n), dataset[1], dataset[2]) # table formation
colnames(table) &lt;- c("№", "Day of war", "Amount of tweets")
cor.test(table[, 1], table[, 3]) # checking for correlation
cor_field(table[, 1], table[, 3], "Correlation field")# construction of the correlation field
r &lt;- correlation_coeficient(table[, 1], table[, 3], n) # finding the correlation coefficient
}
rownames(auto_cor_value) &lt;- c("Lag", "Autocorrelation coefficient")
colnames(auto_cor_value) &lt;- c(1:round(n/4))
View(auto_cor_value)
cat("\n\nAutocorrelation coefficient for different lags:\n")
print(t(auto_cor_value))
barplot(auto_cor_value[2, ], space = 1, names.arg = auto_cor_value[1, ],
col = "orchid3",xlab = "Lag", ylab = "Autocorrelation coefficient",
main = "Correlogram of autocorrelation function")</p>
        <p>Construction of a correlation table for all smoothing intervals, including a number of original values:</p>
        <p>Returning to smoothing, we will modify the graph using Pollard's formulas, and also check their
effectiveness (similarly to Kendall's formulas, which were presented above). Series for analysis:
number of NATO hashtags. Smoothing according to the formulas from Pollard for different w (3, 5, 7,
9, 11, 13, 15):</p>
        <p>Number of turning points Correlation coefficient</p>
        <p>As you can see in the graphs above, Pollard's smoothing also reduces the number of turning points,
but unlike the previous smoothing results, it does not decrease as linearly. In contrast, the correlation
coefficient when applying these formulas decreases more rapidly, but still linearly. Re-smoothing
according to the formulas from Pollard for different w (3, 5, 7, 9, 11, 15):</p>
        <p>Construction of a correlation table for all re-smoothing intervals according to formulas from Pollard:</p>
        <p>Construction of turning point diagrams for all re-smoothing intervals using formulas from Pollard:</p>
        <p>Determination of efficiency criteria for repeated smoothings according to the formulas from Pollard:</p>
        <p>During repeated application of the Pollard method, the number of turning points began to drop
rapidly and linearly, and the same applies to the correlation coefficient. The formulas "smoothed out"
our graph so much that the relationship between the quantities changed its sign to the opposite. We ca n
conclude that Pollard's repeated method for a given function has too strong an effect, and it is better to
use previous methods (ordinary moving average or Kendall's method).
cat("\t\tPollard smoothing")# smoothing according to formulas from Pollard
pol_m &lt;- function(data, w, a){
if(w == length(a)){
for(i in 1:length(data)){
d &lt;- 0
if(i &lt; w){
data[i] &lt;- 0
}else{
for (k in 0:(w-1)){
data[i] &lt;- (a[k+1] * data[i-k]) + d
d &lt;- data[i]
}
}
}
}
return(data)</p>
        <p>Exponential smoothing - the smoothed value is determined by only two values - the current and the
last smoothed level and the ratio of their weights. Series for analysis: the share of #Ukraine hashtags
among the total number of tweets. Carrying out exponential smoothing for different weights - α (0.1,
0.15, 0.2, 0.25, 0.3):</p>
        <p>Therefore, the smaller the α indicator is, the more the levels in the analyzed series are smoothed. We
will conduct studies similar to those conducted for the Kendall and Pollard methods. Construction of a
correlation table for all smoothing intervals, including a number of original values:</p>
      </sec>
      <sec id="sec-3-12">
        <title>Plotting pivot points for all smoothing intervals:</title>
      </sec>
      <sec id="sec-3-13">
        <title>Determination of performance criteria for exponential smoothing:</title>
        <p>Parameter α 0,1 ≤ α ≤ 0,3</p>
        <p>This table differs from the ones we observed before, because now it is the opposite: the smaller the
parameter, the more the function is smoothed (the smaller the number of turning points and the
correlation coefficient).
cat("\t\tExponential smoothing")# exponential smoothing
exp_1 &lt;- HoltWinters(dataset$Ukraine/dataset$daily_tweets, alpha = 0.1, beta = FALSE, gamma = FALSE) # exponential smoothing at a = 0.1
t1 &lt;- turn_points(exp_1$fitted[, 1], n - 1) # number of turning points at а = 0.1
r1_xy &lt;- correlation_coeficient(dataset$Ukraine[1:(n-1)]/dataset$daily_tweets[1:(n-1)], exp_1$fitted[, 1], n - 1) # coefficient at a = 0.1
exp_2 &lt;- HoltWinters(dataset$Ukraine/dataset$daily_tweets, alpha = 0.15, beta = FALSE, gamma = FALSE) # smoothing at a = 0.15
t2 &lt;- turn_points(exp_2$fitted[, 1], n - 1) # the number of turning points at а = 0.15
r2_xy &lt;- correlation_coeficient(dataset$Ukraine[1:(n-1)]/dataset$daily_tweets[1:(n-1)], exp_2$fitted[, 1], n - 1) # coefficient at a = 0.15
exp_3 &lt;- HoltWinters(dataset$Ukraine/dataset$daily_tweets, alpha = 0.2, beta = FALSE, gamma = FALSE) # exponential smoothing at a = 0.2
t3 &lt;- turn_points(exp_3$fitted[, 1], n - 1) # the number of turning points at а = 0.2
r3_xy &lt;- correlation_coeficient(dataset$Ukraine[1:(n-1)]/dataset$daily_tweets[1:(n-1)], exp_3$fitted[, 1], n - 1) # coefficient at a = 0.2
exp_4 &lt;- HoltWinters(dataset$Ukraine/dataset$daily_tweets, alpha = 0.25, beta = FALSE, gamma = FALSE) # smoothing at a = 0.25
t4 &lt;- turn_points(exp_4$fitted[, 1], n - 1) # the number of turning points at а = 0.25
r4_xy &lt;- correlation_coeficient(dataset$Ukraine[1:(n-1)]/dataset$daily_tweets[1:(n-1)], exp_4$fitted[, 1], n - 1) # coefficient at a = 0.25
exp_5 &lt;- HoltWinters(dataset$Ukraine/dataset$daily_tweets, alpha = 0.3, beta = FALSE, gamma = FALSE) # exponential smoothing at a = 0.3
t5 &lt;- turn_points(exp_5$fitted[, 1], n - 1) # the number of turning points at а = 0.3
r5_xy &lt;- correlation_coeficient(dataset$Ukraine[1:(n-1)]/dataset$daily_tweets[1:(n-1)], exp_5$fitted[, 1], n - 1) # coefficient at a = 0.3
plot.ts(dataset$Ukraine/dataset$daily_tweets, xlab = "Day of war", # an image of exponential smoothing on a graph
ylab = "The share of #Ukraine hashtag in all tweets", main = "Exponential smoothing")
lines(exp_1$fitted[, 1], col = "lightskyblue", lwd = 2)
lines(exp_2$fitted[, 1], col = "deepskyblue1", lwd = 2)
lines(exp_3$fitted[, 1], col = "deepskyblue2", lwd = 2)
lines(exp_4$fitted[, 1], col = "deepskyblue3", lwd = 2)
lines(exp_5$fitted[, 1], col = "deepskyblue4", lwd = 2)
legend(170,0.43,col=c('black', 'lightskyblue','deepskyblue1','deepskyblue3','deepskyblue4'),</p>
        <p>legend=c('original', 'α = 0.1', 'α = 0.15','α = 0.2', 'α = 0.25','α = 0.3'),lty=1,cex=0.7)
exp_cor_table &lt;- cbind(dataset$Ukraine[1:(n-1)]/dataset$daily_tweets[1:(n-1)], exp_1$fitted[, 1], exp_2$fitted[, 1], exp_3$fitted[, 1],
exp_4$fitted[, 1], exp_5$fitted[, 1]) # construction of a generalized correlation table
colnames(exp_cor_table) &lt;- c("Y", "Yn, α = 0.1", "Yn, α = 0.15", "Yn, α = 0.2", "Yn, α = 0.25", "Yn, α = 0.3")
View(exp_cor_table) # plotting pivot points for exponential smoothing
barplot(turn_points_value(exp_cor_table[, 2], n - 1), col = "aquamarine3", names.arg = c(1:t1), main = "Turned points diagram, α = 0.1")
barplot(turn_points_value(exp_cor_table[, 3], n - 1), col = "aquamarine3", names.arg = c(1:t2), main = "Turned points diagram, α = 0.15")
barplot(turn_points_value(exp_cor_table[, 4], n - 1), col = "aquamarine3", names.arg = c(1:t3), main = "Turned points diagram, α = 0.2")
barplot(turn_points_value(exp_cor_table[, 5], n - 1), col = "aquamarine3", names.arg = c(1:t4), main = "Turned points diagram, α = 0.25")
barplot(turn_points_value(exp_cor_table[, 6], n - 1), col = "aquamarine3", names.arg = c(1:t5), main = "Turned points diagram, α = 0.3")
# виведення на екран кількості поворотних точок та коефіцієнтів кореляції при різних a
a &lt;- c(0.1, 0.15, 0.2, 0.25, 0.3)
exp_r_xy &lt;- c(r1_xy, r2_xy, r3_xy, r4_xy, r5_xy)
exp_t &lt;- c(t1, t2, t3, t4, t5)
exp_matrix &lt;- cbind(a, exp_t, exp_r_xy)
colnames(exp_matrix) &lt;- c("Smoothing parameter", "Amount of turned points", "Correlation coefficient")
cat("\n Exponential smoothing criteria:\n")
print(exp_matrix)</p>
        <p>Median smoothing is a type of smoothing based on the distributed average, which instead of the
arithmetic mean (depending on the indicator) takes the value corresponding to the median of the interval
(depends on the indicator). Performing median smoothing for different w (3, 5, 7, 9, 11, 13, 15), data
for analysis: number of hashtags #NATO:</p>
      </sec>
      <sec id="sec-3-14">
        <title>Determination of efficiency criteria for performed median smoothing:</title>
        <p>As a conclusion: the larger the parameter, the more the function is smoothed. Given the given data,
we can observe a slight deviation from the rule at parameter "15". We will consider this as an error of
this method given the given input data.</p>
      </sec>
      <sec id="sec-3-15">
        <title>Performing repeated median smoothing for different w (3, 5, 7, 9, 11, 15):</title>
      </sec>
      <sec id="sec-3-16">
        <title>Construction of a correlation table for all intervals of repeated median smoothing:</title>
      </sec>
      <sec id="sec-3-17">
        <title>Determination of efficiency criteria for repeated median smoothing:</title>
        <p>Correlation coefficient  
18
0
0
0
0
0
0</p>
        <p>The repeated median smoothing method, as we can see, like the Pollard method, was not very
effective against our data, as the number of turning points variously fell to zero and the correlation
coefficient became negative.
legend(180,150000,col = c('1','2','3','4','5','6','7'),legend =</p>
        <p>c('original', 'MM3', 'MM5','MM7', 'MM9','MM11', 'MM13','MM15'),lty = 1,cex = 0.8)
movingMedian&lt;- cbind(dataset$NATO,my_movingMedian1,my_movingMedian2, my_movingMedian3,my_movingMedian4
,my_movingMedian5,my_movingMedian6,my_movingMedian7) # generalized correlation table from sliding repeated medians
colnames(movingMedian) &lt;- c("Y", "Y'n, w = 3", "Y'n, w = 5", "Y'n, w = 7", "Y'n, w = 9", "Y'n, w = 11", "Y'n, w = 13", "Y'n, w = 15")
View(movingMedian)
rep_m_t &lt;- vector()# the number of turning points during repeated median smoothing
for (i in 1:7){
rep_m_t[i] &lt;- turn_points(movingMedian[, i + 1], dim(movingMedian)[1])
}
barplot(turn_points_value(my_movingMedian1, length(my_movingMedian1)), col = "lightsalmon",</p>
        <p>names.arg = c(1:rep_m_t[1]), main = "Turned points diagram, w = 3")# plotting pivot points for repeated median smoothing
rep_m_r_xy &lt;- vector()# коефіцієнти кореляції при повторному медіанному згладжуванні
for (i in 1:7){
rep_m_r_xy[i] &lt;- correlation_coeficient(dataset$NATO, movingMedian[, i + 1], dim(movingMedian)[1])
}
k &lt;- c(3, 5, 7, 9, 11, 13, 15) # displaying on the screen the number of turning points and correlation coefficients at different w
rep_mdn_matrix &lt;- cbind(k, rep_m_t, rep_m_r_xy)
colnames(rep_mdn_matrix) &lt;- c("Smoothing parameter", "Amount of turned points", "Correlation coefficient")
cat("\nRepeated median smoothing criteria:\n")
print(rep_mdn_matrix)</p>
        <p>Hierarchical agglomerative cluster analysis of multivariate data:
- solves the problem of group homogeneity of data, ensures the selection of compact, distant groups
of objects, that is, looks for a "natural" division of the population into areas of accumulation of
objects;
- allows dividing objects not by one parameter, but by a whole set of features;
- allows you to view fairly significant volumes of data, sharply shorten and compress them, make
them compact and clear.</p>
      </sec>
      <sec id="sec-3-18">
        <title>For the method is necessary:</title>
        <p>- normalize the data (so that the variance is equal to 1);
- make a matrix of closeness (relationship of the form "indicator-indicator");
- choose a strategy of unification.</p>
        <p>Let's build the "operator-individual indicators" table. Dimensions:    , where  = ̅1̅,̅9̅ is the
number of hashtags, and = ̅1̅̅,1̅̅1̅ is the number of descriptive statistics indicators used. For cluster
analysis: the set G , which includes m objects, each of which is characterized by n features.</p>
      </sec>
      <sec id="sec-3-19">
        <title>Normalization of the object-property table:</title>
      </sec>
      <sec id="sec-3-20">
        <title>Choosing a metric for building a proximity matrix at Euclidean metric:</title>
        <p>( ⃗1,  ⃗2) =
( 1 −  2 )2.</p>
      </sec>
      <sec id="sec-3-21">
        <title>Formation of the proximity table according to the defined metric:</title>
      </sec>
      <sec id="sec-3-22">
        <title>General view of any strategy: .</title>
        <p>Ukraine
putin
left and rows up.</p>
        <p>For our data, we chose the nearest neighbor strategy: distance between groups - the distance between
the two most distant elements of the groups. For her, the parameters acquire the following values :   =
= 0.5, 
= 0,</p>
        <p>= 0.5. Features of the strategy: monotonous, greatly stretches the space.</p>
        <p>Carrying out cluster analysis. Finding the smallest value in the proximity matrix and combining the
objects it corresponds to into one group.</p>
        <sec id="sec-3-22-1">
          <title>Finding the smallest element in the proximity matrix</title>
          <p>1
StopPutin
2.9686466</p>
          <p>Extracting columns belonging to these objects. Eliminate empty space by shifting all columns to the
5
UkraineRussiaWar</p>
          <p>StopRussia &amp;
russian</p>
          <p>9
StopPutin
2.9686466 Ukraine
1.5115275 Russia
0.5295263 StandWithUkraine
1.0453304 Putin
0.5600772 UkraineRussiaWar
0.9269777 NATO</p>
          <p>StopRussia &amp;</p>
          <p>Russian
0.0000000 StopPutin</p>
        </sec>
      </sec>
      <sec id="sec-3-23">
        <title>Enumeration of the value of the extracted columns according to the selected strategy:</title>
        <p>ℎ =    ℎ +    ℎ +    +  | ℎ −  ℎ |.</p>
        <p>The nearest neighbor strategy:   =   = 0.5,  = 0,  = 0.5.</p>
        <p>Insertion of the listed column in the place (empty) of the first removed column. Checking whether
its zero lies on the main diagonal.
russia
putin</p>
        <p>Copying the values of this column, transposing them into a ribbon and replacing it with the ribbon
of the first removed column.
russia
putin
Ukraine
russia
StandWithUkraine</p>
        <p>putin
UkraineRussiaWar</p>
        <p>NATO
StopRussia &amp;
russian
StopPutin</p>
        <p>Assignment to the new object formed as a result of merging the removed objects, next in order of
number.
Ukraine
russia
StandWithUkraine</p>
        <p>putin
UkraineRussiaWar</p>
        <p>NATO
StopRussia &amp;
russian
StopPutin</p>
        <p>1
Ukraine
0.000000
2.012974
2.446287
2.491046
2.888187
2.890969
2.9272550</p>
        <p>2
russia
2.0129736
0.0000000
1.1426592
0.5920176
1.7624964
1.0991498
1.3189270
putin</p>
      </sec>
      <sec id="sec-3-24">
        <title>Repeating the procedure until the matrix is reduced to size 2  2.</title>
        <p>russia
putin
russia
putin
StopRussia &amp; russian
&amp; StopPutin
2.9686466
1.5115275
0.5601943
1.0453304
0.8880631
0.9269777
0.0000000
Ukraine
russia
StandWithUkraine</p>
        <p>putin
UkraineRussiaWar</p>
        <p>NATO
StopRussia &amp; russian &amp;</p>
        <p>StopPutin</p>
        <p>Ukraine
russia
StandWithUkraine &amp; StopRussia &amp;
russian &amp; StopPutin</p>
        <p>putin
UkraineRussiaWar</p>
        <p>NATO</p>
        <p>№
russia
putin
2.4910462
0.5920176
1.0453304
0.0000000
1.4223876
0.6294907</p>
        <p>5
UkraineRussiaWar
2.8881866
1.7624964
0.8880631
1.4223876
0.0000000
1.4376388
1
Ukraine
0.000000
2.491046
2.968647
2.888187
2.890969</p>
        <p>13
russia &amp;
putin
2.491046
0.000000
1.511528
1.762496
1.099150</p>
        <p>12
StandWithUkraine &amp; StopRussia &amp;
russian &amp; StopPutin
2.9686466
1.511528
0.0000000
0.8880631
0.9269777</p>
        <p>5
UkraineRussiaWar
2.8881866
1.762496
0.8880631
0.0000000
1.4376388
1
Ukraine
2.8909688
1.099150
0.9269777
1.4376388
0.0000000</p>
        <p>6</p>
        <p>NATO
2.8909688
1.099150
1.437639
Step #7: proximity matrix by size 3  3</p>
        <p>Ukraine
russia &amp; putin &amp; NATO
StandWithUkraine &amp; StopRussia &amp; russian &amp;</p>
        <p>StopPutin &amp; UkraineRussiaWar
Step #8: proximity matrix by size 2  2</p>
        <p>Ukraine
russia &amp; putin &amp; NATO &amp; StandWithUkraine &amp; StopRussia &amp;</p>
        <p>russian &amp; StopPutin &amp; UkraineRussiaWar
1
2
3
4
5
6
7
8
7 + 8
10 + 9
2 + 4
3 + 11
13 + 5
12 + 6
15 + d14
1 + 16
d10
d11
d12
d13
d14
d15
d16
d17</p>
        <p>Metrics</p>
        <p>The dimensionality of the matrix 2  2 =&gt; STOP. Construction of a dendrogram:</p>
        <p>Interpretation of the result of cluster analysis at level 0.6, 6 clusters are formed: 1 cluster – object
Ukraine; 2nd cluster – object Russia; cluster 3 – objects StandWithUkraine, StopPutin, StopRussia,
russian; cluster 4 – putin facility; 5th cluster – object UkraineRussiaWar; cluster 6 is a NATO facility.</p>
        <p>At level 1.0, 4 clusters are formed: 1 cluster – object Ukraine; 2nd cluster – objects russia, putin;
cluster 3 – objects UkraineRussiaWar, StandWithUkraine, StopPutin, StopRussia, russian; cluster 4 is
a NATO facility. At level 1.4, 3 clusters are formed: 1 cluster – object Ukraine; cluster 2 – NATO,
russia, putin facilities; 3rd cluster - objects of UkraineRussiaWar, StandWithUkraine, StopPutin,</p>
      </sec>
      <sec id="sec-3-25">
        <title>StopRussia, russian.</title>
        <p>As a general conclusion of the results of the cluster analysis, the following can be defined: the largest
share of all hashtags was occupied by #Ukraine, followed by #NATO and #UkraineRussiaWar,
respectively. Detailed information about which indicators apply to which clusters at which level is
provided above.
cat("\tCluster analysis\n")# cluster analysis of multivariate data
operator_property &lt;- describe(dataset[3:11])[3:13] # formation of the "object-property" table
norm_operator_property &lt;- matrix(NA, dim(operator_property)[1], dim(operator_property)[2]) # object-property table normalization
View(operator_property)
for (j in 1:dim(operator_property)[2])
norm_operator_property[, j] &lt;- norm_sequence(operator_property[, j])
colnames(norm_operator_property) &lt;- colnames(operator_property)
row.names(norm_operator_property) &lt;- row.names(operator_property)
View(norm_operator_property)
norm_op_pr_1 &lt;- norm_operator_property # creation of "original table" and "copy table"
norm_op_pr_2 &lt;- norm_operator_property
View(norm_op_pr_1)
View(norm_op_pr_2)
proximity_matrix &lt;- matrix(NA, dim(operator_property)[1], dim(operator_property)[1]) # building the proximity matrix
colnames(proximity_matrix) &lt;- row.names(operator_property)
rownames(proximity_matrix) &lt;- row.names(operator_property)
for (i in 1:dim(operator_property)[1]){ # cycle through the rows of the original table
# цикл по рядках таблиці-копії
for(j in 1:dim(operator_property)[1]){
S &lt;- 0
for(k in 1:dim(operator_property)[2]){ # cycle through stacks of tables</p>
        <p>S &lt;- S + (norm_op_pr_1[i, k] - norm_op_pr_2[j, k])^2 # definition of the Euclidean metric
}
}
proximity_matrix[j, i] &lt;- sqrt(S)
}
View(proximity_matrix) # conducting cluster analysis
t &lt;- vector()# the position of the smallest element
l &lt;- vector()# table of future clusters
cluster_matrix &lt;- proximity_matrix
while(dim(cluster_matrix)[1] != 2){ # combining elements into clusters until the dimension of the matrix is 2 x 2
for(i in 1:dim(cluster_matrix)[1]){
for(j in 1:dim(cluster_matrix)[1])
if(cluster_matrix[i, j] == min(as.dist(cluster_matrix))){ # finding the smallest value in the proximity matrix
t &lt;- c(i, j)
break
}
}
a &lt;- t[1] # ensuring the shift of columns (rows) to the left (up)
t[1] &lt;- min(t)
t[2] &lt;- a # formation of a new column name of the proximity matrix (based on the names of the combined elements)
colnames(cluster_matrix)[t[1]] &lt;- paste(colnames(cluster_matrix)[t[1]], colnames(cluster_matrix)[t[2]], sep=" &amp; ")
row.names(cluster_matrix)[t[1]] &lt;- colnames(cluster_matrix)[t[1]] # forming a new row name of the proximity matrix
l &lt;- rbind(l, c(colnames(cluster_matrix)[t[1]], cluster_matrix[t[1], t[2]])) # adding generated nodes and corresponding metrics to the table
# enumeration of the values in the removed columns by the nearest neighbor strategy
cluster_matrix[, t[1]] &lt;- 0.5*cluster_matrix[, t[1]] + 0.5*cluster_matrix[, t[2]] +</p>
        <p>0.5*abs(cluster_matrix[, t[1]] - cluster_matrix[, t[2]])
cluster_matrix[t[1], ] &lt;- t(cluster_matrix[, t[1]]) # filling an empty row (transposing a calculated column)
cluster_matrix[t[1], t[1]] &lt;- 0 # replacing the minimum value with 0
cluster_matrix &lt;- cluster_matrix[, -t[2]] # reducing the dimensionality of the matrix (removing empty rows and columns)
cluster_matrix &lt;- cluster_matrix[-t[2], ]
}
View(cluster_matrix)
# creation of the "union - node - metric" table, adding the last combination of elements to the cluster table
l &lt;- rbind(l, c(paste(colnames(cluster_matrix)[1], colnames(cluster_matrix)[2], sep = "&amp;"), cluster_matrix[1, 2]))
n &lt;- vector()# creating a column of nodes
for (i in 1:dim(l)[1])
n[i] &lt;- paste("d", dim(proximity_matrix)[1] + i, sep = "")
l &lt;- cbind(c(1:dim(l)[1]), l[, 1], n, l[, 2])
colnames(l) &lt;- c("Step", "Unification", "Node", "Metric")
cluster_groups &lt;- as.matrix(l)
View(cluster_groups)
cat("Created clusters:\n")
print(as.character(cluster_groups[, 2]))
library(cluster) # construction of a dendrogram
clust &lt;- hclust(as.dist(proximity_matrix), method = "complete")
plot(clust, xlab = "Hashtags")
interpretation &lt;- vector() # interpretation of cluster analysis
cl &lt;- c(6, 4, 3)
cl_level &lt;- c(0.6, 1.0, 1.4)
for(i in 1:3){
rect.hclust(clust, cl[i], border = i + 1)
cluster &lt;- cutree(clust, cl[i])
interpretation &lt;- cbind(interpretation, colnames(as.dist(proximity_matrix)), cluster)
colnames(interpretation)[i] &lt;- paste("cluster division on level ", "'", cl_level[i], "'", sep = "")
}
cat("\nInterpretation of cluster analysis:\n")
print(interpretation)</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusions</title>
      <p>Summing up, it can be stated that the hypothesis put forward at the beginning is confirmed. Having
analyzed the appearance of graphs of functions, correlations between time-quantity values, we can
confirm the opinion that at the beginning of the war (the first week), information about the situation in
Ukraine became a kind of explosion in the media space of the world. The number of comments, posts,
"tweets" on this topic reached a record high of 4,000,000. After a week, this trend changed dramatically
- the number of discussions decreased significantly. This is due in particular to the fact that the crisis
stage has passed (there has been a lot of information, it is no longer so "interesting" for the world).
Although it is also worth saying that this decrease continued to a specific mark. For 50 days from the
start of the invasion, the number of publications decreased from 4,000 thousand to 500 thousand. At
the level of 500 thousand, the number is maintained until now (with insignificant deviations during
certain military events). We can assume that under constant circumstances, without significant changes,
this trend will take on a downward trend. At the same time, if the circumstances change, it is impossible
to make objective assumptions. Another interesting conclusion can be that during the cluster analysis,
it was found that the world focuses more on Ukraine, as a victim of the war, than on russia, as the
aggressor. Also, a large share of foreigners supports the initiative to help Ukraine (in particular, the
issue of joining NATO). On the contrary, it should not be forgotten that the assumptions "support"
"do not support" are not objective, since people tagging the NATO hashtag can also write negative
posts. A fair statement would be the following: foreigners are concerned about the situation in Ukraine
and its request for protection from NATO. In addition, the last conclusion is that, although interest in
Ukraine decreases over time, it still does not fall lower than before the war. It can be said that the war
drew attention to our country, foreigners heard about us and will not forget our name or location. The
conducted analysis demonstrated the dynamics of changes in the activity of discussing the war and
confirmed the previously put forward assumptions.
6. References
[1] O. Snopok, We watch, read, listen: how the media consumption of Ukrainians changed in the
conditions of a full-scale war. URL: https://www.pravda.com.ua/columns/2022/06/22/7353987/
[2] Media consumption of Ukrainians in conditions of full-scale war. URL:
https://www.oporaua.org/report/polit_ad/24068-mediaspozhivannia-ukrayintsiv-v-umovakhpovnomasshtabnoyi-viini-opituvannia-opori
[3] M. Ulyanovska, The Russian information war in Ukraine is failing - Nina Yankovich. Interview.</p>
      <p>URL: https://ukrainian.voanews.com/a/jankowicz-russian-disinformation-failing/6961072.html
[4] E. Musk:
https://twitter.com/elonmusk/status/1526834598935949312?cxt=HHwWgMDTxc3s7AqAAAA
[5] S. McSweeney. Our current approach to the war in Ukraine. URL:
https://help.twitter.com/uk/safety-and-security/our-ongoing-approach-to-the-war-in-ukraine
[6] MediaSapiens, Twitter will warn about disinformation about the war in Ukraine. URL:
https://ms.detector.media/sotsmerezhi/post/29536/2022-05-20-twitter-poperedzhatyme-prodezinformatsiyu-shchodo-viyny-v-ukraini/
[7] MediaSapiens, Musk announced how many bots he had on Twitter. URL:
https://ms.detector.media/it-kompanii/post/29532/2022-05-20-mask-povidomyv-skilky-botivnarakhuvav-u-twitter/</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>