<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A content analysis software system for eficient monitoring and detection of hate speech in online media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuliya Krylova-Grek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr Burov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute for Digitalisation of Education</institution>
          ,
          <addr-line>9 M. Berlinskoho Str., Kyiv, 04060</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National University of “Kyiv-Mohyla Academy”</institution>
          ,
          <addr-line>2 Hryhoriya Skovorody Str., Kyiv, 04655</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Vienna</institution>
          ,
          <addr-line>5 Liebiggasse, Vienna, 1010</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Uppsala University</institution>
          ,
          <addr-line>Gamla torget 3, 753 20 Uppsala</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <fpage>224</fpage>
      <lpage>233</lpage>
      <abstract>
        <p>This paper presents the results of interdisciplinary project that is a combination of computer program and psycholinguistic approach to media study. In the research we presented the programs that can be used for monitoring and analysis of media content to identify hate speech at its early stage. The aims of research were the following: 1) develop content analysis program for monitoring Russian media outlets; 2) apply the psycholinguistic approach for identifying hidden and manipulative hate speech. In the research there were used two types of content-analysis: quantitative and qualitative. Quantitative content analysis was conducted with computer program that was developed to select publication that could have contained hate speech. For qualitative content analysis the psycholinguistic method of text analysis was used. The method applies for identification methods and tolls that journalists use to incriminate hidden and manipulative hate speech. It is hypothesized that programs of content-analysis help to optimize work and makes it less time-consuming and more efective for analyst, journalists and other specialists who involved into media study. Methods. Quantitative content analysis, psycholinguistic method of qualitative content-analysis. Quantitative content analysis was developed with Python programming language. The publications were selected according to the key words, periods of search (month) and the name of outlet. The list of key words includes words that are used in media for discrimination, dehumanization, and marginalization of objects of hate. Implementation such a program helped to reduce time of monitoring of media outlets. The qualitative content-analysis was conducted with the authors' psycholinguistic method of text analysis that can be applied for analyzing media texts. The programs of content analysis were applied within the project “Hate Speech in Online Media Publicizing Events in Crimea”. The results were published in a data analysis report on spreading the hate speech in the Russian language media communicating the armed Ukraine - Russia conflict and events related to it in Crimea on a regular base (December 2020 - May 2021). The research showed that the content analysis programs used in the project are useful tools for systematizing and processing data in humanities research and can be used by a wide range of specialist who have deal with collection and processing of information (media, communication, human rights and so on).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;content-analysis</kwd>
        <kwd>media</kwd>
        <kwd>hate speech</kwd>
        <kwd>text</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Nowadays, people deal with large amounts of information from diferent online resources. According to
the Digital 2020 Global Overview, the average person spends 6 hours and 43 minutes online per day, or
approximately 40% of their time [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. To navigate a large amount of information, a person needs critical
thinking and analysis skills, which are considered to be among the priorities in the 21st century [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Examining large amounts of textual data is a time-consuming and labour-intensive process that can be
simplified by using computer programs that allow quicker and more eficient processing of information
and receive reliable quantitative results [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Content analysis is a qualitative and quantitative method of
analysing the content of documents in order to identify and measure various facts and trends reflected
in these documents. The peculiarity of programs are that the documents are studied in their social
context through content analysis. It can be used both as a primary research method and in combination
with other methods (e.g. in studies of media performance, in classifying responses to open-ended
questionnaires). Unlike other research methods, content analysis of a text is characterised by the fact
that its procedure involves counting the frequency and volume of references to certain semantic units
of the text under study. The quantitative characteristics of the text obtained in this way make it possible
to draw conclusions about the qualitative content of the documents.
      </p>
      <p>At the same time, if information analysis is part of a professional activity, a specialist needs additional
special skills for monitoring, processing, and analysing data. For example, a such specialist as media
analyst, PR, journalist etc. has not only to collect and analyse information, but also to monitor sources
to find a certain type of information. Therefore, it would be advisable to introduce special courses
in the curriculum to teach the basics of big data analysis for future humanities professionals whose
work involves processing and analysing information. In this context, teaching content analysis skills
allows reducing time and conduct research, proceed data and analysis information more efectively.
For example, for civil activists such skills allow efectively identify texts that contain manipulation,
disinformation and hate speech. In our study, we are talking about journalists, human right defenders,
and other specialties related to media analysis and processing.</p>
      <p>In the context of our work, we will consider an example of the content analysis software application
for conducting quantitative and qualitative research of the online media space. The study was conducted
as part of a project involved a technical specialist, a psycholinguist, the human right activists, and
journalists. The task of the project was to monitor Russian media to observe the situation before the
military aggression, because we consider media as an important tool for understanding the vectors
of strategic communication in the country. The results of the study (2014–2021) revealed that the
Russian media actively used covert and manipulative hate speech aimed at creating a dehumanised,
marginalised, and demonised image of Ukrainians as speakers of the Ukrainian language and culture.
Hidden and manipulative hate speech was a signal that should have drawn public attention to further
discrimination and genocide against the targets of hate.</p>
      <p>
        As we see, the importance of monitoring hate speech at its initial stage is stated in documents of
international organisations and research by scholars. For example, the organization “Anti-Defamation
League” use “Pyramid of hate” that illustrates the prevalence of bias, hate and oppression in society and
demonstrates the progressive intensification of hate: from supporting each next step to genocide. The
lowest level “bias attitude” includes stereotypes, insensitive remarks, microaggression and so on. The
upper levels “bias motivated violence” and “genocide” include act of violence and annihilate an entire
people [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The stages in the bottom act as a base for mass atrocities can be consider as an early stage
of hate speech that can be detected with the help of content analysis.
      </p>
      <p>
        In the 2017 annual report of the Council of Europe highlights the danger of the dissemination and
amplification of poor-quality information in media that is divided into three types: mis-, dis-, and
mal-information, which are difered based on the dimensions of harm and falseness: mis-information is
when false information is shared, but no harm is meant; dis-information is when false information is
knowingly shared to cause harm, and mal-information is when genuine information is shared to cause
harm. The report stressed that hate speech is used to cause harm and discriminate against a group of
people on religion, race, and other grounds: “. . . people are often targeted because of their personal
history or afiliations. While the information can sometimes be based on reality (for example targeting
someone based on their religion) the information is being used strategically to cause harm” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In recent years, there has been considerable interest in issues of the negative influence of mass media
in humanities (psychology, linguistics, sociology, etc.), IT, and Computational linguistics.</p>
      <p>
        In linguistics, much work on the potential of speech propaganda and manipulation has been carried
out by Bulyigina and Shmelev [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], Elswah and Howard [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Aronson and McGlone [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], Soloviova [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Numerous studies have been published on the distortion of the meaning of concepts in media texts
(McGlone et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Shmelev [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Burridge and Allan [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Vakaliuk et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Pilkevych et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]
and others).
      </p>
      <p>
        There is a vast amount of literature on specific features of media representations of war and military
conflicts (Pocheptsov [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], Pack [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], Kamalipour [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], Galtung [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], Dawes [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]). In particular,
Dawes [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] points out that since speech shapes our perception of reality, the age of information can be
called the age of manipulation.
      </p>
      <p>
        In computational linguistics, the study and detection of hate speech explore by using natural language
processing. For example, Schmidt and Wiegand [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] studied the ways of the automatic detection of
hate speech. All surveyed methods include common features that are usually used in the computer
program to identify hate speech, such as a set of negative words or expressions, using various complex
features using (“dependency parse information”, “features modeling specific linguistic constructs”,
“meta-information” and so on). At the same time the authors stressed that in most cases, computer
program results can’t be considered as full, because “they are only evaluated on individual data sets
most of which are not publicly available” [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Taking into consideration the weak features of computer
analysis of hate speech, we think that the best result researchers can receive if they combine computer
and human inspection.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Aims</title>
      <p>The aims: 1) develop content analysis program for monitoring Russian media outlets; 2) apply the
psycholinguistic approach for identifying hidden and manipulative hate speech.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Theoretical background</title>
      <p>
        Content analysis is a quantitative and qualitative method of analysing content (text, pictures etc.)
in order to identify and measure various facts and trends reflected in these materials. The purpose
of quantitative content-analysis is to fins valid resources and to collect certain data The purpose of
qualitative content analysis is to organize and elicit meaning from the collected data and to draw
realistic conclusions from it. The utilization of content analysis in Humanities and Social Sciences
are described by Dey [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], Lester et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], Bengtsson [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The authors pay attention to qualitative
content analysis and stress the advantages of utilizing capabilities of computer (applications, programs,
soft) for gathering and proceeding data. However, the authors didn’t consider the question of media
content-analysis and didn’t propose any special programs that can help to gather and measure hate
speech in online media.
      </p>
      <p>The peculiarity of content analysis of a media text is its direct correlation with external circumstances,
such as the socio-political situation. Content analysis is based on counting the occurrence of determined
components in the analysed information materials, supplemented by identifying statistical relationships
and analysing the structural links between them.</p>
      <p>The need to use content analysis for analysis of large text arrays was caused by the development
of mass communications in the late nineteenth and early twentieth centuries. The content-analysis
was used for analytical research of texts in the media outlets. This method can be used as the main
method, for example, to analyse the political orientation of a publication, or as an auxiliary or control
method with other methods, for example, to measure media efectiveness. Content-analysis can also be
applied alongside with other methods, for example linguistic and psycholinguistic, for evaluating the
efectiveness of a media outlet.</p>
      <p>
        One of the main founders of the research procedure is the sociologist Lasswell [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], who was the
ifrst to examine the impact of the media on the worldview of the population during the First World
War. He chose the texts of newspapers, bulletins, and other information messages as sources, identified
key themes, statements and social models on which propaganda was based. Based on the data obtained,
he drew conclusions about the strategic goals of the countries of a conflict [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. The method of content
analysis was implemented the during second world war for identifying if an American newspaper
consists pro-Nazi texts. As a result, the newspaper was closed, and the method was widely recognised
and began to be actively used to analyse media products [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>
        The classical model of content analysis proposed by Lasswell [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] is the following: 1) depending on
the purpose of the study, the text is divided into parts, each of which was subsequently subjected to
analysis, then the results were compared and summarised; 2) key units are selected for text analysis.
The keywords may vary depending on the research objective, for example, a word-symbol such as the
name of a leader, the name of a country, the name of an ideology, etc. As a result of the analysis of
information messages, we should get answers to the main 5 questions: who transmits the information,
what is the information about, how the information is disseminated, what audience the information is
intended for, what efect the message has [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        Lasswell’s research became the basis for the development of further procedures for analysis of media
texts and answering a variety of questions, ranging from the peculiarities of the worldview of the
average citizen to the ideological orientation of a particular media organ [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>Today, various types of content analysis are used both for analysing large data sets and for analysing
a single text. Because of a large demand of such programs. IT industry is actively developing software to
improve functionality and customise content-analysis programs to analyse big text data, which parses
text from unformatted content and unstructured data from social media, news reports, surveys, etc. to
provide practical information such as the mentioning frequency of a certain brand, personality event, or
counting particular words in the texts (e.g., programs for semantic content analysis such as Semantrum,
Nvivo, MAXQDA, Yoshikoder, Advego, SentiStrength, Nvivo, etc.). The brief characteristics of some of
these programmes are given below.</p>
      <p>
        ADVEGO is a free programme for semantic analysis of texts. It performs text statistics by the number
of words, determines the most frequently used words and their number, as well as words and phrases
that are part of the semantic core [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
      </p>
      <p>
        Yoshikoder is a free programme that allows you to find online texts by a given phrase or expression.
This programme allows you to examine keywords-in-context, and perform basic content analyses, in
any language. At the same time, in our research we need to use some specific keywords that are not
included in the dictionary [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
      </p>
      <p>
        Programme SentiStrength which refers to the programme for Sentiment Analysis or Opinion Mining.
The programme is set up to search for evaluative and emotive vocabulary in the text according to a
pre-compiled dictionary, which is divided into groups: “negative” and “positive” vocabulary, which is
placed on the scale of emotional intensity of the statement [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ].
      </p>
      <p>
        MaxQDA, NVivo software can be applied in a range of sectors: from social science and education to
healthcare and business. These programs can be used to analyse data from interviews, surveys, field
notes, web pages, and journal articles. These programmes are similar in terms of functionality and
operation [
        <xref ref-type="bibr" rid="ref31 ref32">31, 32</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <p>In the research the following methods were used: quantitative content analysis, psycholinguistic method
of qualitative content-analysis. Quantitative content analysis was developed with Python programming
language. The publications were selected according to the key words, periods of search (month) and
the name of outlet. The list of key words includes words that are used in media for discrimination,
dehumanization, and marginalisation of objects of hate. Implementation such a program helped to
reduce time of monitoring of media outlets. The qualitative content-analysis was conducted with the
psycholinguistic method of text analysis that is the authors’ method that can be applied for analysing
media texts.</p>
      <sec id="sec-4-1">
        <title>4.1. Quantitative content analysis</title>
        <p>
          The study outlines the approaches and opportunities for interdisciplinary interaction of humanities and
computer sciences. The research of hate speech was done with collaboration with the Crimean Human
Rights Group (Siedova and Krylova-Grek [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]). The objective of the research was to identify hidden and
manipulative hate speech in a number of Russian media outlets. A content analysis programme was
used for selecting media texts based on selected keywords. The study was conducted in 2014-2022 and
consisted of three stages: basic monitoring; quantitative content analysis; qualitative content-analysis;
conclusions.
I. Basic monitoring. At this stage, a list of media outlets has been made. The study investigated 11
online media outlets with audiences of over 1 million readers per month.
        </p>
        <p>II. Quantitative content analysis. This stage involves 1) monitoring of the selected media outlets; 2)
identification of groups in relation to which negative characteristics are applied; identification of
keywords used to create a negative image of the selected groups; 4) texts selection.
III. Qualitative content analysis. At this stage, we applied the author’s psycholinguistic analysis of the
selected texts.</p>
        <p>For conducting quantitative content-analysis, a special program was developed by analyst and
technical specialists of the Crimean Human Rights Group O. Sedov (now the program is under the
process of patent). The program allows to optimize the teamwork by surveying a large array of
information to find a pool of texts that can incite xenophobia and hatred.</p>
        <p>The program sets up the following parameters for hate speech content analysis:
1. Online resource name. We specified the names of news agencies that specialize in current news.</p>
        <p>The initial cohort includes the nine most popular online media, whose trafic ranges from 1
million to 15 million visitors monthly.
2. Keywords. We entered words that usually accompany texts with hate speech. To single out
keywords, we developed Hate dictionary based on the careful analysis of publications in selected
media. Monitoring and word selection took place in 2017-2018. The dictionary includes more
than 400 words. It should be noted the dictionary is constantly supplemented and changed due to
the emergence of new narratives, concepts, and words.
3. Search period. We set the time interval: date and year. In our research, the most appropriate
search period that can provide us with the required amount of data is one month. For example,
from 1 of August to 31 of August.</p>
        <p>The content analysis program searched the selected websites to find in which hate speech regarding
such ethnic and social groups might be used: Ukrainians, Crimean Tatars, Jews, Residents of Crimea
and Donbas Russians, Activists and Journalists, Euromaidan Participants, LGBTQ groups. Most of the
hate words referred to Ukrainians.</p>
        <p>Key words included new created lexicon, archetypes of the second world war and other words
that marginalized some ethnic and social group. Studies found that articles featuring hate speech
often use terms and concepts that are not recorded in the oficial dictionaries of the Russian language.
Some of these words are ‘made up’, created purposefully to incite hatred. In some cases, such words
are formed by combining parts of words denoting nationality and obscene lexical units, for example,
“kriptobanderovtsy” (a compound word made up of “crypto” and “Banderite”), “natsgady” (a compound
word made up of a shortened form of Nazi/Nationalists that sounds similar and is used by Russian media
as a play on words; and the word ‘gady’ that is similar to ‘foul people’ in English (skunk, despicable),
“khokhlodauny” (a compound word made up of ‘khokhly’ (see above) and the Russian word for Down’s
syndrome suferers ([douny] /n., pl.(ukr.)/).</p>
        <p>Apart from other things, we considered a well-known vocabulary available in glossaries, and with
negative connotations, that was used by journalists in the investigated media in relation to objects
of hatred, ridicule, marginalisation, and so on. Most such words go back to of WWII archetypes, and
common negative stereotypes. For example, “fascist”, “fascism”, “nazi”. Moreover, based on the phonetic
similarity of the words “Nazi” and “nationalism” in Ukrainian and Russian, journalists use the word
[natsist] instead of [natsionalist] (Nazi, nationalist).</p>
        <p>Among other words are those that humiliate and marginalise another people’s language by distorting
the phonetic sound of words for sarcastic or mocking reasons, and bracketing the phonetic spelling of
words (the transliteration of Ukrainian words into Russian in terms of our study), which contextually,
in the publications, is sarcastic and expresses contempt for the Ukrainian language and its speakers
(e.g., “svidomyye”, “nezalezna”).</p>
        <p>The study materials were monitored and sampled using a content analysis program: texts were
selected using set key words and word combinations. Such words and word combinations had been
collected by the project’s monitoring group during preliminary studies into the hate content of
Russianlanguage online media and public social networks. Nine Russian-language sites were searched. The
Russian language vocabulary used in relation to the main ethnic groups, as well as to the most vulnerable
social groups of the population, who currently live in the territory occupied by the Russian Federation,
was added to the list of key words and word combinations.</p>
        <p>As an example, we delve into the interface of the hate speech content analysis program in greater detail.
First and foremost, it’s important to highlight the distinctive features of the program’s terminology:
“query words” refer to the specific words we input into the search based on our Dictionary of Hate. On
the other hand, “keywords” are the words that most frequently appear in the text of a publication. For
instance, let’s take the online platform Politnavigator as an example. We selected the keywords for
September 2019 (figure 1).</p>
        <p>The program’s interface presented us with the following information:
• the total number of articles containing the selected keywords (both in their titles and within the
text);
• a list of these articles along with corresponding links;
• a list of the keywords identified within the articles;
• the frequency of each keyword’s appearance within the articles;
• the number of query words used.</p>
        <p>Upon hovering over a specific article, the interface provided the subsequent details:
• a direct link to the respective article;
• a hyperlink directing to the media outlet’s website where the article is published;
• the keywords featured within the article;
• the frequency and number of the keywords;
• the frequency and number of the query words.</p>
        <p>This interface design allows for a thorough examination of hate speech content, aiding in the analysis
and understanding of its prevalence and distribution across various publications.</p>
        <p>The content analysis program was used to electronically process the content of the selected websites
for the indicated period regarding abovementioned key words and word combinations. The sample
of publications produced by the content analysis program was subject to a psycholinguistic analysis
exercise. The content was divided into two groups based on the exercise results: publications featuring
hate speech, and publications with other types of manipulations.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. The psycholinguistic analysis</title>
        <p>The texts selected through content analysis underwent psycholinguistic analysis to distinguish between
those containing hate speech and those that did not. The content analysis program selected all texts that
contain key words, but not all the texts consists hate speech. The reasons are related to the algorithm for
configuring the content analysis programme which scans the page together with comments and other
information. Therefore, the reasons for the errors are justified by the following factors: 1) comments
under the text included hate vocabulary. As the research aimed to analyse the products of the media
specialists’ activities, the comments were not taken into account and such texts were attributed to
error. Among other things, comments can be a product distributed by bots or specifically hired people,
which requires additional technical methods for their analysis; 2) texts in which keywords have a direct
meaning; for example, the word “fascists” used in the text give a factual retrospective to the military
events of the Second World War. At the same time, there were only a few such texts (2%) that were
removed from the list for further analysis.</p>
        <p>
          Subsequently, the texts containing hate speech were scrutinized with psycholinguistic analysis to
identify the methods and techniques employed by journalists in these publications [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ]. The analysis is
a part of an innovative author’s methodology, which assists in identifying both direct and manipulative
hate speech that does not contain direct insults and calls for gender, racial or religious intolerance,
but forms a negative attitude towards certain groups and individuals. The psycholinguistic analysis
was performed manually, as only a professional can assess sarcasm, infer indirect meanings of words,
decode and elucidate the significance of newly created words that do not exist in the dictionary.
        </p>
        <p>After having conducted the textual analysis, hate speech was divided into three types, which are
characterized by the specific linguistic and graphic tools used in the publication:
Type 1: direct hate speech;
Type 2: indirect (hidden) hate speech;
Type 3: manipulative hate speech.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and discussion</title>
      <p>
        The monitoring period of online media are from December 1, 2020, to May 31, 2021. We think that
and the amount of content for this period is enough to validate the study outcomes. The project
monitoring group obtained 1,284 publications when keyword-based electronic content sampling had
been completed. The content was processed according to the psycholinguistic method of media text
analysis, and the result was that 560 publications featuring hate speech elements selected from the
entire content were received by the project analytical experts. These include 16 texts with Hate Speech
Type One; 341 texts with Hate Speech Type Two; and 203 texts with Hate Speech Type Three [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ].
Type 1. Direct use of hate speech is characterised by: the use of obscenities, direct insults and calls for
violence.
      </p>
      <p>Type 2. Indirect use of hate speech is characterised by: marginalisation of the other party with usage of
ofensive ethnonyms, polarisation, dividing a society into in-group and out-group,
generalisation (attributing one case to an entire group), negative sarcasm and irony, the use of archetypes
and stereotypes that develop a certain world view and attitude, the creation of new concepts
with negative connotations.
Type 3. Manipulative hate speech is characterised by: substituting the meanings of concepts that create
negative associations; using fake; use the opinion of biased or ‘pseudo-experts’ citing (people
who have little or no experience or knowledge about the problem on which they comment);
distorting and misinterpreting historical facts; justifying aggression or violence against a target
group;enhancing informative messages with non-linguistic means (photograph, pictures); using
manipulative titles that does not match, or distorts, the information presented in the text of
the article.</p>
      <p>
        The study investigated eleven Russian language online media, including: five news websites, with
news items on the situation in Crimea largely dominating the content, the audiences of over 1 million
readers per month, and a minimum 25% share of Ukrainian readers; three Russian news websites
regularly writing about Crimea and Donbas, with audiences of over 1 million readers per month, and a
minimum 25% share of Ukrainian readers; two news websites, with largely dominating content that
describes the situation in Crimea, financed out of the Russian Federation budget; and the oficial website
of the “government” of the Pravitel’stvo Kryma (figure 2). The analysis of the auditoriums at selected
sites was conducted using SimilarWeb, a platform that ofers insights into global digital trafic [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ].
      </p>
      <p>The reason for including texts that did not contain hate speech in the search results is as follows: a)
certain articles did contain words from the dictionary, but they corresponded to their direct meaning,
for example, the word “down” was used to refer to people with Down syndrome; b) the program counted
words and expressions contained in the comments under the publication. The second aspect was the
main reason for the large percentage of inappropriate texts.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>The efectiveness of using a quantitative content analysis programme to monitor hate speech is
determined by the following criteria:
1) the amount of time a specialist spends on selecting texts: a specialist spends up to 15 minutes on
selecting articles;
2) search criteria: in the software, you can set the search period, the source to be analysed, and a fairly
large number of keywords that can be entered, for example, in our study we identified about 400
words related to hate speech;</p>
      <p>As a result, reducing the time spent on quantitative analysis allows you to spend more time on
qualitative analysis of texts. Another advantage of the quantitative-qualitative approach is that, unlike
other search engines and programs, it has better functionality for selecting texts, reduces the time
for text selection, and when combined with psycholinguistic analysis of texts, allows us to go beyond
simple mathematical calculations and analyse the psycholinguistic methods and techniques used by
journalists to influence the minds of the audience.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We would like to express our special thanks of gratitude to Oleksandr Sedov for creating the quantitative
content analysis program and the civil organization “The Crimean Human Rights Group” for selecting
and providing research materials.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[1] Digital</source>
          <year>2020</year>
          ,
          <article-title>Global digital overview</article-title>
          ,
          <source>Technical Report, Hootsuite</source>
          ,
          <year>2020</year>
          . URL: https://media.rbcdn. ru/media/reports/Digital_2020.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[2] The Future of Jobs Report</source>
          <year>2020</year>
          ,
          <string-name>
            <given-names>Technical</given-names>
            <surname>Report</surname>
          </string-name>
          , The World Economic Forum,
          <year>2020</year>
          . URL: https: //www3.weforum.org/docs/WEF_Future_of_Jobs_
          <year>2020</year>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. J.-Y.</given-names>
            <surname>Guédé</surname>
          </string-name>
          ,
          <article-title>From Content Analysis to Content Analysis of Digital Social Networks</article-title>
          , in: 2022 International Conference on
          <source>International Studies in Social Sciences and Humanities (CISOC</source>
          <year>2022</year>
          ), Atlantis Press,
          <year>2022</year>
          , pp.
          <fpage>235</fpage>
          -
          <lpage>248</lpage>
          . doi:
          <volume>10</volume>
          .2991/978-2-
          <fpage>494069</fpage>
          -25-1_
          <fpage>23</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Anti-Defamation</surname>
            <given-names>League</given-names>
          </string-name>
          ,
          <source>Pyramid of Hate</source>
          ,
          <year>2021</year>
          . URL: https://www.adl.org/sites/default/files/ pyramid
          <article-title>-of-hate-web-english_1</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wardle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Derakhshan</surname>
          </string-name>
          , Information Disorder:
          <article-title>Toward an Interdisciplinary Framework for Research and Policymaking</article-title>
          ,
          <source>Technical Report, Council of Europe</source>
          ,
          <year>2017</year>
          . URL: https://rm.coe.
          <source>int/ information-disorder-report-november-2017/1680764666.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bulyigina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shmelev</surname>
          </string-name>
          ,
          <article-title>Linguistic conceptualization of the world (based on the material of Russian grammar</article-title>
          ), Moscow,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Elswah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Howard</surname>
          </string-name>
          , “
          <article-title>Anything that Causes Chaos”: The Organizational Behavior of Russia Today (RT)</article-title>
          ,
          <source>Journal of Communication</source>
          <volume>70</volume>
          (
          <year>2020</year>
          )
          <fpage>623</fpage>
          -
          <lpage>645</lpage>
          . doi:
          <volume>10</volume>
          .1093/joc/jqaa027.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Aronson</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. S. McGlone</surname>
          </string-name>
          ,
          <article-title>Stereotype and social identity threat</article-title>
          , in: T. D.
          <string-name>
            <surname>Nelson</surname>
          </string-name>
          (Ed.), Handbook of prejudice, stereotyping, and discrimination, Psychology Press, New York,
          <year>2009</year>
          , pp.
          <fpage>153</fpage>
          -
          <lpage>178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Soloviova</surname>
          </string-name>
          ,
          <article-title>The features of the learning of political linguistics at higher education institutions</article-title>
          ,
          <source>Educational Dimension</source>
          <volume>3</volume>
          (
          <year>2020</year>
          )
          <fpage>217</fpage>
          -
          <lpage>232</lpage>
          . doi:
          <volume>10</volume>
          .31812/educdim.v55i0.
          <fpage>3944</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M. S. McGlone</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Pfiester</surname>
          </string-name>
          , Contamination and Camouflage in Euphemisms,
          <source>Communication Monographs</source>
          <volume>73</volume>
          (
          <year>2006</year>
          )
          <fpage>261</fpage>
          -
          <lpage>282</lpage>
          . doi:
          <volume>10</volume>
          .1080/03637750600794296.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D. N.</given-names>
            <surname>Shmelev</surname>
          </string-name>
          , Euphemism [evfemizm], in: Y. N.
          <string-name>
            <surname>Karaulov</surname>
          </string-name>
          (Ed.),
          <article-title>Russkij yazyk</article-title>
          .
          <source>Enciklopediya [Russian language. Encyclopedia]</source>
          , 2 ed.,
          <source>Bolshaya ros. encikljpediya. Drofa</source>
          , Moscow,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Burridge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Allan</surname>
          </string-name>
          ,
          <article-title>Euphemism &amp; dysphemism: Language used as shield and weapon</article-title>
          , Oxford University Press,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Vakaliuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Pilkevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fedorchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Osadchyi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tokar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Naumchak</surname>
          </string-name>
          ,
          <article-title>Methodology of monitoring negative psychological influences in online media</article-title>
          ,
          <source>Educational Technology Quarterly</source>
          <year>2022</year>
          (
          <year>2022</year>
          )
          <fpage>143</fpage>
          -
          <lpage>151</lpage>
          . doi:
          <volume>10</volume>
          .55056/etq.1.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>I. A.</given-names>
            <surname>Pilkevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Fedorchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Romanchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. M.</given-names>
            <surname>Naumchak</surname>
          </string-name>
          ,
          <article-title>Approach to the fake news detection using the graph neural networks</article-title>
          ,
          <source>Journal of Edge Computing</source>
          <volume>2</volume>
          (
          <year>2023</year>
          )
          <fpage>24</fpage>
          -
          <lpage>36</lpage>
          . doi:
          <volume>10</volume>
          .55056/jec.592.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Pocheptsov</surname>
          </string-name>
          , Theory of communication, Refl-book, Moscow,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>G. G. Pocheptsov,</surname>
          </string-name>
          (Des)information, Palivoda, Kiev,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Pack</surname>
          </string-name>
          , Book Review:
          <source>The Language of War by Steve Thorne</source>
          ,
          <year>2006</year>
          . London and New York: Routledge, pp.
          <source>vii+ 104 ISBN: 0 415 35868 X (pbk)</source>
          ,
          <source>Language and Literature</source>
          <volume>18</volume>
          (
          <year>2009</year>
          )
          <fpage>408</fpage>
          -
          <lpage>410</lpage>
          . doi:
          <volume>10</volume>
          .1177/09639470090180041201.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Y. R.</given-names>
            <surname>Kamalipour</surname>
          </string-name>
          ,
          <article-title>Language, media and war: Manipulating public perceptions</article-title>
          ,
          <source>Media</source>
          <volume>3</volume>
          (
          <year>2010</year>
          )
          <fpage>87</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Galtung</surname>
          </string-name>
          ,
          <article-title>Language and War: Is There a Connection?</article-title>
          ,
          <source>Current Research on Peace and Violence</source>
          <volume>10</volume>
          (
          <year>1987</year>
          )
          <fpage>2</fpage>
          -
          <lpage>6</lpage>
          . URL: https://www.jstor.org/stable/40725052.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Dawes</surname>
          </string-name>
          ,
          <article-title>The language of war: literature and culture in the US from the Civil War through World War II</article-title>
          , Harvard University Press,
          <year>2002</year>
          . doi:
          <volume>10</volume>
          .4159/9780674030268-intro.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegand</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          <article-title>Survey on Hate Speech Detection using Natural Language Processing</article-title>
          , in: L.
          <string-name>
            <surname>-W. Ku</surname>
          </string-name>
          , C.-T. Li (Eds.),
          <source>Proceedings of the Fifth International Workshop on Natural Language Processing for Social Media</source>
          , Association for Computational Linguistics, Valencia, Spain,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>W17</fpage>
          -1101.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>I. Dey</surname>
          </string-name>
          ,
          <article-title>Qualitative data analysis: A user friendly guide for social scientists</article-title>
          ,
          <source>Routledge</source>
          ,
          <year>2003</year>
          . doi:
          <volume>10</volume>
          .4324/9780203412497.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J. N.</given-names>
            <surname>Lester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Lochmiller</surname>
          </string-name>
          ,
          <article-title>Learning to do qualitative data analysis: A starting point</article-title>
          ,
          <source>Human resource development review 19</source>
          (
          <year>2020</year>
          )
          <fpage>94</fpage>
          -
          <lpage>106</lpage>
          . doi:
          <volume>10</volume>
          .1177/1534484320903890.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bengtsson</surname>
          </string-name>
          ,
          <article-title>How to plan and perform a qualitative study using content analysis</article-title>
          ,
          <source>NursingPlus Open</source>
          <volume>2</volume>
          (
          <year>2016</year>
          )
          <fpage>8</fpage>
          -
          <lpage>14</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.npls.
          <year>2016</year>
          .
          <volume>01</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>H. D.</given-names>
            <surname>Lasswell</surname>
          </string-name>
          ,
          <article-title>The theory of political propaganda</article-title>
          ,
          <source>American political science review 21</source>
          (
          <year>1927</year>
          )
          <fpage>627</fpage>
          -
          <lpage>631</lpage>
          . doi:
          <volume>10</volume>
          .2307/1945515.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Kostenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. F.</given-names>
            <surname>Ivanov</surname>
          </string-name>
          ,
          <source>Experience of Content Analysis: Models and Practices</source>
          , Centre for Free Press,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>H. D.</given-names>
            <surname>Lasswell</surname>
          </string-name>
          ,
          <article-title>The structure</article-title>
          and
          <article-title>function of communication in society</article-title>
          ,
          <source>The communication of ideas 37</source>
          (
          <year>1948</year>
          )
          <fpage>136</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Advego</surname>
          </string-name>
          ,
          <source>Semantic Text Analysis for SEO</source>
          ,
          <year>2024</year>
          . URL: https://advego.com/text/seo/.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>W.</given-names>
            <surname>Lowe</surname>
          </string-name>
          , Yoshikoder:
          <article-title>Cross-platform multilingual content analysis</article-title>
          ,
          <year>2015</year>
          . URL: https://yoshikoder. org.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Y. Z.</given-names>
            <surname>Hung</surname>
          </string-name>
          , Python 3 wrapper
          <string-name>
            <surname>for</surname>
            <given-names>SentiStrength</given-names>
          </string-name>
          ,
          <year>2021</year>
          . URL: https://github.com/zhunhung/ Python-SentiStrength.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Nvivo</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://help-nv.
          <source>qsrinternational.com/12/win/v12.1</source>
          .
          <fpage>112</fpage>
          -
          <lpage>d3ea61</lpage>
          /Content/cases/ cases.htm.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <article-title>MAXQDA | Oficial Site | All-In-One Tool for Qualitative Data Analysis</article-title>
          ,
          <year>2024</year>
          . URL: https://www. maxqda.com/.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>I.</given-names>
            <surname>Siedova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Krylova-Grek</surname>
          </string-name>
          ,
          <article-title>Hate Speech in Online Media Publicizing Events in Crimea: a data analysis report on spreading the hate speech in the Russian language media communicating the armed Ukraine - Russia conflict and events related to it in Crimea on a regular base (December 2020</article-title>
          - May
          <year>2021</year>
          ),
          <source>Technical Report, Kyiv</source>
          ,
          <year>2022</year>
          . URL: https://crimeahrg.org/wp-content/uploads/ 2023/04/hate-spe.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Krylova-Grek</surname>
          </string-name>
          ,
          <article-title>Psycholinguistic approach to the analysis of manipulative and indirect hate speech in media</article-title>
          ,
          <source>East European Journal of Psycholinguistics</source>
          <volume>9</volume>
          (
          <year>2022</year>
          )
          <fpage>82</fpage>
          -
          <lpage>97</lpage>
          . doi:
          <volume>10</volume>
          .29038/ eejpl.
          <year>2022</year>
          .
          <volume>9</volume>
          .2.kry.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Similarweb</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://www.similarweb.com/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>