<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring the Relationship between News Reliability and Violent Comments in Digital Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Beatriz Botella-Gil</string-name>
          <email>beatriz.botella@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alba Bonet-Jover</string-name>
          <email>alba.bonet@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robiert Sepúlveda-Torres</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patricio Martínez-Barco</string-name>
          <email>patricio@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Estela Saquete</string-name>
          <email>stela@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Software and Computing Systems, University of Alicante</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Natural Language Processing (NLP) has become an essential tool for the automatic detection of violent language and unreliable information. The misuse of Information and Communication Technologies (ICTs) fosters the generation of disinformation and digital violence, thus polarising society. Furthermore, the lack of reliable, neutral and accurate language when presenting news can trigger an increase in negative and violent users' reactions. It is necessary to find a linkage between disinformation and violent discourse in order to moderate online content, facilitate early detection of both phenomena and ensure healthy online behaviour. This research uses NLP techniques to explore the reliability of news headlines and their correlation with violent language generated in comments by users. In addition, the generation of a novel resource annotated in Spanish is created to jointly address the automatic detection of violent language and reliability.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Natural Language Processing</kwd>
        <kwd>Violent language</kwd>
        <kwd>Reliability detection</kwd>
        <kwd>Data analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>come ideal venues for the propagation of violent and false main research carried out in the fields of news reliability
content that generates group-based divisions in society. and language employed in disinformation, on the one</p>
      <p>As stated by [1], hate speech is based on the divide hand, and violent language, on the other; in Section 3,
et impera concept, which intends to group society and the methodology employed for this exploratory study is
turn diferent groups against one another. When society outlined; Section 4 delves into the exploratory analysis to
becomes polarised, disinformation can proliferate more determine the relationship between disinformation and
easily. In particular, as reflected in [ 3]’s research, Spain violence; Section 5 discusses the findings of this research,
stands out as the most polarised country in Europe. and finally, in Section 6, conclusions and limitations for</p>
      <p>The intersection of violent language and news disin- future work are discussed.
formation in the digital world constitutes a crucial area
of today’s research. Addressing these issues will not only
aid in identifying the reliability of information but also 2. State of the Art
contribute to fostering a safer and more equitable online
society. This research aims to conduct an exploratory The scientific community has approached the issue of
study on the prevalence of violence in news comments, online violence from various fields of knowledge, driven
examining both the source and the subject matter. Ad- by the alarming prevalence of violent content in digital
ditionally, it seeks to determine whether a correlation media. Eforts have been made to address issues such as
exists between the violent content generated by readers hate speech, cyberbullying, and toxicity, aiming to early
and the reliability of that content, achieved through an detect these problems and propose solutions.
analysis of language usage. The objective is to determine Regarding the disinformation phenomenon, its
objecwhether the language used in a news article influences tive is to disseminate misleading or unreliable content,
users to employ a higher level of violent language in the which undermines societal trust by instilling insecurity
comments pertaining to the news item in question. and doubts regarding the authenticity of the information</p>
      <p>One of the limitations in both the task of detecting received.
violent discourse and the task of detecting disinformation This section aims to present relevant research related
is the scarcity of resources in Spanish to train models in to both the violent discourse and the disinformation tasks,
Natural Language Processing (NLP). For that reason, this as well as to define the main important concepts of these
research focuses on the analysis of a set of news items lines of research, specially related to violent and reliable
and comments in Spanish, in order to analyse violent and language.
unreliable language in the Spanish language and propose
a resource for both tasks. 2.1. Language in violence discourse</p>
      <p>The main novelty of our proposal consists of
examining the correlation between violent discourse and
disinformation in Spanish by means of NLP tools. To achieve
this objective, the following contributions to this research
area have been provided:</p>
      <sec id="sec-1-1">
        <title>Language, as a means of communication, should be neu</title>
        <p>tral. Still, its use can promote any ideology and incite
hatred and violence [1].</p>
        <p>An example of the power of language is evident in
contexts of warfare, as seen in historical scenarios like
• A new resource consisting of headlines and news the Cold War, the World War I, and the Vietnam and Iraq
comments, annotated with the reliability and vio- wars [4]. Similarly, in political discourses such as the
lence criteria. Reliability was manually annotated, 2016 US presidential elections or the Brexit referendum
while violence was automatically annotated by [5], language has played a significant role, particularly
means of an existing automatic violence classifier in addressing the issue of disinformation.
originally created for tweets annotation that is In the scientific field, research on linguistic violence
applied in this research to news comments anno- has been approached from various angles. It is important
tation. to acknowledge the challenge in defining what
consti• An exploratory analysis of how source profiles tutes violent content. Studies have aimed to delineate
and topics can influence the generation of digital diferent forms of violent language prevalent in ICTs,
violence and exacerbate hostility among users. including:
• An assessment of how language used in news
articles can impact the occurrence of hate speech,
and more specifically, how the language used in
news headlines may either incite readers to
exhibit varying levels of violence when expressing
their opinions on an issue.
• Ciberbullying: the deliberate and repetitive use
of specific ICTs like email, mobile phone
messages, instant messaging, and personal online
defamatory behaviour by an individual or group,
with the intention of hostilely causing harm to
another party [6].
• Hate speech: any form of communication that
discriminates against an individual or group
based on characteristics such as race, ethnicity,
gender, sexual orientation, nationality, religion,
or other identifying features [7].
• Toxicity: this term is defined by [ 8] and [9] as
those messages that include unacceptable, rude,
and disrespectful comments. They are messages
of contempt that are part of what is called toxic
discourse or toxic language, which either invite
other users to leave the discussion or use
language at the same level as the discussion.
• Ofensiveness : in [10] asserts that this term is
commonly defined as hurtful, derogatory, or
obscene remarks directed from one person to
another1.</p>
      </sec>
      <sec id="sec-1-2">
        <title>In addition, we also find other works that further spec</title>
        <p>ify the type of violence they study, such as in the case of
misogyny or racism [11]. Our research will use the terms
“violent language” or “violence” to refer to any type of
message containing violent content.</p>
        <p>Given the vast volume of data present in the virtual
environment, manual detection of violent messages by
humans is impractical. This underscores the significance
of NLP tools in our research. From a NLP standpoint,
detecting hate speech can be conceptualised as text
classification. As stated by [ 12], “the automatic detection of
this kind of speech is usually addressed as a
classification task, and it is related to a family of other tasks such
as detecting cyberbullying, ofensive language, abusive
language, toxic language, among others”.</p>
        <p>From this perspective, numerous research eforts
focus on detection methods, including the development of
resources aimed at aiding identification. These resources
primarily consist of lexicons containing lists of negative
words or expressions that may be present in messages,
serving as indicators of potentially violent content. These
word lists are employed in binary classification models to
ascertain the presence or absence of hate speech [13],
violent language [14], abusive language [15], and ofensive
language [16].</p>
        <p>Furthermore, corpus generation is utilised in
various violence detection strategies, thereby enhancing the
model’s capability to efectively discern both violent and
non-violent content [17, 18].</p>
        <p>Besides, eofrts have been made to identify the most
suitable detection tools that ofer improved and expedited
solutions to the problem. This includes exploring
strategies such as heuristic techniques [19], Machine Learning
(ML) [20], and Deep Learning (DL) [21].</p>
      </sec>
      <sec id="sec-1-3">
        <title>1https://thelawdictionary.org/</title>
        <sec id="sec-1-3-1">
          <title>2.2. Language in disinformation</title>
          <p>Language also plays a very important role in
disinformation detection and there are key linguistic indicators
that are characteristic of news items. As described by
[4], the essence of fake news or deceptive information
is the news text, and, consequently, the language: “the
news text is the basic communicative unit of journalism
and consequently the basic unit of analysis in most
research on fake news”. First of all, it is important to define
several key concepts in this field of research: fake news,
disinformation, misinformation, reliability and veracity.</p>
          <p>• Fake news: even if there is no universal
definition, it is commonly defined as “any news that is
suspected to be inaccurate, biased, misleading, or
fabricated” and is understood as a “product of a
range of practices that are related to the validity
of information being shared by the news media”
[4].
• Disinformation vs. Misinformation: in [1]
states that “the information shared with malicious
intent is recognised as disinformation, whereas
the same information shared by a poorly
informed party is considered as misinformation”.</p>
          <p>The main diference lies in the intention:
disinformation refers to that content which is
intentionally and deliberately created with the intention to
deceive while misinformation is inaccurate
information that can results from an honest mistake,
negligence or unconscious bias [22]. Briefly, as
stated by [23], “disinformation is also wrong
information but, unlike misinformation, it is a known
falsehood”.
• Reliability vs. Veracity: these concepts are
closely related, but from what can be observed in
the literature, the term veracity is usually used in
tasks in which the information is contrasted and
verified [ 24], while the concept of reliability is
most commonly used in methods where the
credibility of the source of the news is investigated, as
is the case of the proposed source-based method
[25].</p>
          <p>For this research, we will focus on the disinformation
and the reliability concepts. Reliability is an essential
metric to be considered when assessing the quality of
information. Some of the indicators that influence the
reliability of a news item are: the ambiguity of the
information, the lack of data and sources [26], the intention of
hiding information, the representativeness or opacity of
the headline, the external quotes from experts, the quotes
from studies and organisations, emotional-charged
expressions [27], or stylistic features such as the
punctuation, the extension, the use of capital letters or informal
or swear words [28].</p>
          <p>In addition, [29] suggested that there are ten indicators which presents a dataset based on user responses to posts
that are most likely to detect the reliability of a message from Argentinian digital newspapers on Twitter. Their
and these appear when the message is: complete, concise, research is specially focused on the hate speech detection
coherent, well presented, objective and representative; task but the authors work with replies to digital
newswhen it contains no spin, uses expert sources, is perceived papers posts, which is an interesting point of view for
to have an impact and is professional. our research because we also work with users’ replies, in</p>
          <p>As in Section 2.1, NLP plays an essential role in this our case with news comments, but instead from digital
line of research, as the continuous and easy access to newspapers posts, we focus on digital news.
the internet, the large volume of misleading content and These authors also describe in their article a work in
its rapid viralisation make it impossible to process and progress on hate speech that analyses Spanish tweets
treat data manually in the time required and before it linked to newspapers, which is also related to our work,
becomes pervasive in society. For that reason, disinfor- albeit with a diferent type of document.
mation detection needs to be automated by means of NLP In [32] aim to analyse the relationship between the
techniques. consumption of misinformation and the online hate and</p>
          <p>Disinformation detection is specially being addressed toxic language. To that end, they analyse a corpus of
as a classification task, by training supervised ML models comments on Italian Youtube videos, distinguishing
beto automatically distinguish between real and fake news tween four categories of hate speech and categorising
[4]. For that purpose, annotated corpora are needed to two types of speech channels: questionable (channels
mark patterns of language that will help in the training likely to disseminate unverified and false content) and
and classification processes. Most corpora generated reliable (the remainder of the channels).
in this domain is composed of news articles classified In [33] analyse a large Twitter dataset annotated with
following several techniques, ranging from stylistic and hate speech and counterhate speech. They state that,
linguistic annotations to binary classifications depending even if misinformation and hate speech have been
tackon fact-checking verification [ 28, 30, 26]. In addition, led in parallel, they are often interwoven because “biased
several methods are being applied to the disinformation people justify and defend their hate speech using
misinproblem based on knowledge, style, propagation, source formation”. That reason makes this linkage crucial for
or even hybrids methods, each of them approaching the efective online content moderation.
automatic detection from diferent approaches [25]. To the best of the authors’ knowledge, there is limited
research that conducts an analysis connecting these two
2.3. Violence and Reliability: a new line of lines of research (disinformation and violent discourse),
especially in the context of Spanish language.
Furtherresearch
more, our research proposes a comprehensive analysis
Polarisation emerges as a key factor linking the phe- of both news reliability and violent discourse, using a
nomena of disinformation (in particular, reliability) and newly created resource tailored specifically to address
violent online messages that initially appear disparate. these two research areas.</p>
          <p>Disinformation, by spreading erroneous or misleading
information, can intensify polarisation by encouraging 3. Methodology
the generation of radical opinions and the adoption of
hostile language. Polarisation, in turn, acts as a catalyst The main objective of this research is to ascertain the
for the spread of hate speech, creating an environment correlation between various attributes of news articles,
conducive to distrust towards those with divergent views. including topic, source credibility, and content reliability,
This vicious cycle continuously feeds back on itself, with and the emergence of violence and polarisation within
disinformation and hate speech fuelling polarisation, and readers’ comments on said articles. To achieve this goal,
with polarisation facilitating the spread of more disinfor- two principal analyses are conducted: i) an exploratory
mation and hate. analysis into the prevalence of violent language within</p>
          <p>Through a detailed analysis of relevant case studies digital media, aiming to assess the degree of violence
relaand research, it is assessed how this cycle contributes tive to topic and source; and ii) a study of the association
to the escalation of online violence and undermines the between the reliability of news headlines and the
occurfoundations of social cohesion and democracy [31]. It rence of violent language within readers’ comments on
also discusses possible strategies to address these prob- news articles. The methodology employed to accomplish
lems, highlighting the importance of media education, this objective entails a dual process involving data
collecthe promotion of critical thinking and the building of tion and the analysis of both the violent language within
more inclusive and respectful online communities. comments and the reliability of news headlines. This</p>
          <p>Among research linking hate speech and disinforma- is achieved through annotation using specific schemes
tion, it can be highlighted the work presenting by [12],
tailored for each aspect.</p>
        </sec>
        <sec id="sec-1-3-2">
          <title>3.1. Data collection</title>
        </sec>
      </sec>
      <sec id="sec-1-4">
        <title>News was gathered using a web crawler that extracted</title>
        <p>data from ten Spanish digital sources. A diverse sample
of widely-read national media digital newspapers was se- 3.3. Reliability analysis
lected, each representing various ideological orientations.</p>
        <p>This diverse selection enables us to assess the prevalence The objective of this second analysis is to identify a
corof violent language across diferent news sources. To relation between the usage of violent language in news
analyse the extent of violence by topic, ten news items comments and the reliability of news headlines. As
defrom each source were examined, resulting in a total tailed in (anonymous), the reliability of content is
asof 100 articles for this initial study. These articles en- sessed by considering the objectivity and accuracy of
compass a range of topics, ranging from politics, society, the language used. Through this analysis, a research is
economics, health, and sports. conducted to determine whether a relationship exists
be</p>
        <p>Given that news articles typically maintain a more tween disinformation (unreliable news) and the hatred
neutral tone and considering that many of the chosen expressed in the comments pertaining to the news item
sources consist of newspapers authored by professional in question.
journalists, the decision was made to evaluate violence
within the comments section rather than within the ar- 3.3.1. Data annotation
ticles themselves. The objective is to gauge the level
of violence stemming from readers’ comments, as these The annotation task for this second analysis was
concomments reflect individual users’ opinions and, for the ducted manually. Out of the 100 headlines, 31 were
catemost part, are not moderated, ofering insights into how gorised as Unreliable, while 69 were categorised as
Reliusers express themselves. able. All annotations were performed by an expert in NLP</p>
        <p>Moreover, concerning disinformation, the focus of this specialised in linguistics and disinformation detection.
study is placed on the content reliability, as unreliable The annotation of headlines considered the following
content is deemed potential disinformation. In this re- reliability criteria, as outlined in (anonymous):
gard, to investigate whether the reliability of the news
items influences the level of present violence, news
headlines were used to examine if there is a relationship
between news reliability and violence in comments.</p>
        <p>The chosen system classified a total of 1,757 comments
as Violent and 3,940 comments as Non-Violent. Following
the application of automatic classification, manual
supervision was conducted by an expert in NLP specialised in
criminology and violent language.</p>
        <p>• Data accuracy: data should avoid vagueness or
ambiguity. The employment of evasive or
indefinite expressions suggests concealment or an
inability to substantiate a fact. Additionally, the
absence of evidence, such as scientific studies or
verified oficial data, undermines the reliability of
a news item.
• Data objectivity: the information presented
should maintain neutrality. Information that
sways the reader either positively or negatively,
or that reeflcts the author’s standpoint through
personal remarks or experiences, indicates low
reliability. Subjective data make the reader more
vulnerable to believe unreliable news items.
• Headlines style: headlines should be
informative, concise and neutral. Alarmist, subjective,
opaque and striking headlines are characteristic
of unreliable news items, as well as clickbait
headlines, which are those “sensational headlines that
often exaggerate facts, usually to entice readers to
click on them” [34]. Other features, such as long
headlines or the presence of many exclamation
marks or words in capitals can also influence the
reliability of the content.</p>
        <sec id="sec-1-4-1">
          <title>3.2. Violent language analysis</title>
          <p>In this initial analysis, the primary objective is to examine
the prevalence of violent language within the digital
journalism landscape. To achieve this, news articles collected
from various digital newspapers presented diverse
political afiliations and levels of popularity. The selection
criteria included not only the newspapers’ popularity but
also their editorial style and user engagement.
3.2.1. Data annotation</p>
        </sec>
      </sec>
      <sec id="sec-1-5">
        <title>The annotation process of the downloaded comments was</title>
        <p>performed using an automatic violence classifier
(anonymous). This classifier was developed by fine-tuning a
RoBERTa model in Spanish, using a dataset of tweets
annotated with Violent and Non-Violent labels. Given
the similarity in the language and length between tweets
and news comments, the system was applied to this
research. Moreover, it yields significant results in
discerning whether a tweet demonstrates violence, achieving an
1 of 0.854 in a test set.
to Nazi ideology, confirming what is known as Godwin’s
law, which states that as an online discussion lengthens,
This section outlines three preliminary studies examin- the probability of someone comparing another person or
ing the use of violent language within the digital news group of people to Nazis or Adolf Hitler tends to increase,
landscape and its correlation with various external fac- regardless of the topic or position being debated [35]. For
tors, including topic, source, and the reliability of the example:
news content. These studies facilitate a comprehensive
understanding of violent discourse behaviour within this
context, as well as the impact of news language
reliability on the proliferation of violent language in news
comments.
• Comment in political news item: socialistas? =
autoritarios fascionazis!! [socialists? = fascionazi
authoritarians!!]
• Comment in sport news item: Mucho
machinazi llorando [A lot of machinazi crying]</p>
        <sec id="sec-1-5-1">
          <title>4.1. Level and type of violence by topic</title>
        </sec>
        <sec id="sec-1-5-2">
          <title>4.2. Level of violence by source</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Exploratory analysis</title>
      <p>From these results it can be concluded that the more
credibility the source has, the lower the percentage of
violence is generated in the comments associated with the
news items published. Still, since this was a preliminary
study to prove if there was a linkage between reliability
and violence, the analysis was carried out in a small
sample. As the results show that we can go even deeper</p>
      <sec id="sec-2-1">
        <title>After analysing violent discourse related to the topics of the news items, the results obtained regarding the level of violence across diferent news topics are presented in Table 1.</title>
      </sec>
      <sec id="sec-2-2">
        <title>One of the factors we aimed to analyse was the percent</title>
        <p>age of violence generated by the source, contingent upon
its credibility.</p>
        <p>As stated by [25],“one can detect fake news by
assessTable 1 ing the credibility of its source, where credibility is often
Violence percentage per news topic. defined in the sense of quality and believability”. This
Topic Violence research proposes to relate credibility to the content
analysed according to the proposed reliability criteria in
SecPolitics 38.5% tion 3.3.1. Since there is currently no standard measure
Economics 28.6% of media credibility, but several factors are taken into
Society 27.0% account, in this research we will consider as credible
SHpeoarltth 2243..65%% sources those presenting reliable content, while those
sources whose news items analysed did not meet the
reliability criteria of neutrality, objectivity, accuracy and
coherence were classified as sources with low
credibil</p>
        <p>As can be observed, political comments exhibit the ity. To accomplish this, a correlation will be established
highest level of violence (38.5%), followed by economics, between the number of headlines annotated as reliable
society, sport and health, with the latter having the lowest and unreliable, and the credibility level of the media in
percentage of violent language. which these news items are published.</p>
        <p>It is important to emphasise that while the news items In this regard, three levels of credibility will be
delinalign with these topics as categorised by digital media, eated: high credibility, indicated when the percentage of
comments that deviate from the associated topic of the reliable headlines exceeds 80%; low credibility, identified
news items have been encountered. For instance, eco- when over 80% of the analysed headlines were deemed
nomic news may also attract violence directed towards unreliable and failed to meet the established reliability
crithe political sphere due to the close relationship between teria; and medium credibility otherwise, denoting sources
these topics. Similarly, in sports news, sexist comments that present both reliable and unreliable headlines in
simimay arise, particularly when addressing women’s foot- lar proportions. In our study, upon analysing all the news
ball, as exemplified by: headlines, six sources were classified as highly credible,
three were identified as having low credibility, and one
source was deemed partially credible.</p>
        <p>Results obtained concerning the level of violence
according to the sources credibility are presented in Table
2.
• Violent comment: Descanse en paz la z
empoderada que quería hacer cosas de chicos, como jugar al
fútbol y conducir su propio coche, con el riesgo que
eso tiene. A trabajos y actividades de hombres,
riesgos de hombres, eso es todo. [Rest in peace to the
empowered z who wanted to do boys’ things, like
play football and drive her own car, with all the
risk that entails. Men’s jobs and men’s activities,
men’s risks, that’s all.]</p>
      </sec>
      <sec id="sec-2-3">
        <title>In addition, we found cases in diferent topics where the discussion drifts into radicalisation with comparisons</title>
        <p>to test our hypothesis, this analysis will be carry out in a
larger sample in future work.</p>
        <sec id="sec-2-3-1">
          <title>4.3. Degree of violence depending on the reliability of the headline</title>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>One of the main purposes of this research is to find if there</title>
        <p>is a relationship between the reliability concept (mostly
related to disinformation tasks) and the violent discourse.
This work aims to join these two lines of research to delve
into the misleading and malicious digital content and
propose a preliminary combined solution to automatic
detect both disinformation and violent discourse. This
analysis is carried out in the news scenario on the basis
of the following data:
• News: 69 Reliable news (69%) and 31 Unreliable
news (31%).
• Comments: 3,940 Non-Violent comments
(69.15%) and 1,757 Violent comments (30.85%).</p>
      </sec>
      <sec id="sec-2-5">
        <title>After conducting the analysis contrasting the violence</title>
        <p>generated in the comments of news articles with a reliable
headline versus those with an unreliable headline, as
depicted in Figure 1, it is evident that reliable headlines
generate less violence than unreliable ones.</p>
        <p>Following our analysis, we observed that of the 31
headlines classified as unreliable, 19 of them attracted a
higher number of violent comments. This implies that
more than 61.29% of the unreliable headlines elicited
more negative responses from users, evidencing a
relationship between the unreliability of a news item and
the propensity of users to express violence in their
comments.</p>
        <p>On the other hand, of the 69 news items that have
been classified as reliable, only 9 of these news items
had more violent comments. Therefore, 81.15% of these
news items recorded more non-violent comments and
only 5.79% generated an equal amount of violent and
non-violent comments. These findings support the
reliability of the headlines, suggesting that those news items
considered more reliable tend to generate a lower
proportion of violent comments compared to the unreliable
ones.</p>
        <p>The following examples show how the way a headline
is presented or written can incite users to generate more
violent content in comments:
• Unreliable headline: Pedro Sánchez en
Estrasburgo; sentí vergüenza ajena [Pedro Sánchez in
Strasbourg; I felt ashamed]
– Violent comment: Usted y todos. Nos dejó
a la altura del betún, como si fueramos una
dictadura bolivariana cualquiera. (Que, sin
duda, es lo que el quiere. . . ) [You and
everyone else. He made us feel small, as if we
were a Bolivarian dictatorship (which, no media sources plays a significant role in shaping the tone
doubt, is what he wants...)] and content of user-generated comments. A lack of trust
in the source can trigger negative and hostile reactions,
• Unreliable headline: Detengamos a la izquierda underscoring the pivotal role of integrity and reliability
ya [Let’s stop the left now] in digital journalism.</p>
        <p>– Violent comment: La izquierda no existe Reliability: the reliability of the language used in
es todo extrema derecha los que se dicen de a news item can indeed impact both the quantity and
izquierdas en realidad no lo son, es una falsa tone of the violent comments it provokes. Trustworthy
izquierda que traicionando al pueblo se ha media outlets typically uphold ethical and professional
aliado con la banca usurera satanica que es standards in information presentation, thus lowering the
la que hay que detener el verdadero enemigo probability of users responding in such an aggressive
de la humanidad [The left does not exist it manner. Conversely, media with a lower credibility level
is all extreme right those who say they are may disseminate misleading, biased, or sensationalist
leftists actually they are not, it is a false left content, prompting negative and violent reactions from
that betraying the people has allied itself readers.
with the satanic usurious banking institu- Additionally, it is crucial to highlight the significance
tion which is the real enemy of humanity of social media moderators used by certain media
outthat must be stopped] lets. These moderators play a vital role in curbing the
dissemination of violent comments by moderating and</p>
        <p>As can be observed in the previous examples, the lan- censoring those that contravene website usage policies.
guage employed in news headlines is inherently biased,
characterised by a subjective style that mirrors the
author’s polarised political stance. The manner in which 6. Conclusions and future work
the headline is formulated and presented allows little
space for readers to formulate their own conclusions, but
rather prompts them to generate negative and violent
content on the given topic.</p>
        <p>In this study, we have investigated the correlation
between violence and disinformation. By integrating these
two research approaches, our aim is to ascertain whether
a relationship exists between violent language and the
reliability language in news.
5. Discussion Our findings show that addressing both violence and
reliability concurrently may facilitate the identification
The present research has yielded a series of significant and mitigation of malicious digital content. For instance,
results, which are detailed below. the level of violence in comments may correlate with the</p>
        <p>Topics: our analysis indicates that politics emerges reliability of news headlines and the credibility of the
as the subject with the highest incidence of violence media outlets.
in comments across various news topics. This finding For future work, we aim to expand the initial resource
suggests that the polarising and impassioned nature of created to propose a novel corpus annotated with
reliapolitics encourages users to express themselves more bility and violent language that facilitates the automatic
freely, potentially leading to more confrontational and detection of both tasks. The objective is to ensure a
balhostile interactions. Additionally, it is noteworthy that anced distribution of the news articles across diferent
health garners the lowest frequency of violent comments, topics in order to ofer a more comprehensive and
reprepossibly due to its relatively lower prominence in our sentative understanding of the diverse types of violence
selection of news articles. prevalent in the digital sphere.</p>
        <p>We have also found that there is not always a relation- In addition to expanding and balancing the corpus,
ship between the topic of the news item and the type of a manual and meticulous revision process of the
autoviolence that appears in users’ comments. As explained matic violence annotation made by the classifier will
in Section 4.1, we found cases of news items classified be carry out, in order to ensure the correct
classificaas economics or sports containing comments of sexist or tion of the comments. This revision task, which will be
political violence. accomplished by two NLP experts in linguistics and
vi</p>
        <p>Credibility: the analysis also uncovers a direct asso- olent discourse, will enable the generation of a quality
ciation between the credibility of media sources and the annotated resource for both the violence discourse and
prevalence of violence in online comments. Users are in- disinformation tasks.
clined to express a higher frequency of violent comments Finally, it is proposed to conduct an assessment of a
on news sourced from outlets that are perceived as less set of comments categorised as Non-Violent to determine
reliable. This observation suggests that the credibility of whether they present concealed violence, that is, whether
language employed subtly hides violence patterns, such [5] R. Greifeneder, M. E. Jafe, E. J. Newman,
in the case of irony or sarcasm. Alongside further exami- N. Schwarz, The Psychology of Fake News:
Acnation of the extent of violence, a future hypothesis will cepting, Sharing, and Correcting Misinformation,
focus on studying whether reliable news correlates with Routledge, New York, NY, 2020.
levels of mild and more subtler forms of violence, and [6] B. Belsey, Cyber-bullying: An emerging
whether unreliable news demonstrates more pronounced threat to the ‘always on’generation, 2006,
violence expressed in a more aggressive manner. This Retirado de: http://www. cyberbullying.
step will contribute to refining detection methods and ca/pdf/Cyberbullying_Article_by_Bill_Belsey.
enhancing the accuracy of evaluations. pdf (2014).</p>
        <p>In summary, this study represents a significant step [7] W. Warner, J. Hirschberg, Detecting hate speech on
towards a deeper understanding of the linkage between the world wide web, in: Proceedings of the second
violence and disinformation in the NLP context, laying workshop on language in social media, 2012, pp.
the groundwork for future research eforts. This research 19–26.
aims at fostering a safer and healthier digital environment [8] R. Nielsen, N. de Domenico, Volume and patterns
by creating a resource that combines both news reliability of toxicity in social media conversations during the
and violent language of comments and thus proposing a covid-19 pandemic (2020).
common strategy to address these two research lines. [9] E. Wulczyn, N. Thain, L. Dixon, Ex machina:
Personal attacks seen at scale, in: Proceedings of the
26th international conference on world wide web,
Acknowledgments 2017, pp. 1391–1399.
[10] M. Wiegand, J. Ruppenhofer, T. Kleinbauer,
DeThe research work is part of the R&amp;D&amp;I projects: tection of abusive language: the problem of biased
CLEAR.TEXT: Enhancing the modernization public datasets, in: Proceedings of the 2019 conference of
sector organizations by deploying Natural Language the North American Chapter of the Association for
Processing to make their digital content CLEARER to Computational Linguistics: human language
techthose with cognitive disabilities” (TED2021-130707B- nologies, volume 1 (long and short papers), 2019,
I00), funded by MCIN/AEI/10.13039/501100011033 pp. 602–608.
and “European Union NextGenerationEU/PRTR”; [11] F.-M. Plaza-Del-Arco, M. D. Molina-González, L. A.
COOLANG.TRIVIAL: Technological Resources for Ureña-López, M. T. Martín-Valdivia, Detecting
Intelligent VIral AnaLysis (PID2021-122263OB-C22) misogyny and xenophobia in spanish tweets
usfunded by MCIN/AEI/10.13039/501100011033/ and ing language technologies, ACM Transactions on
by "ERDF A way of making Europe"; SOCIALFAIR- Internet Technology (TOIT) 20 (2020) 1–19.
NESS.SOCIALTRUST: Assessing trustworthiness [12] J. M. Pérez, F. M. Luque, D. Zayat, M. Kondratzky,
in digital media (PDC2022-133146-C22) funded by A. Moro, P. S. Serrati, J. Zajac, P. Miguel, N.
DeMCIN/AEI/10.13039/501100011033/ and by the "Euro- bandi, A. Gravano, et al., Assessing the impact of
pean Union NextGenerationEU/PRTR". At regional level, contextual information in hate speech detection,
this research has been funded by the project NL4DISMIS: IEEE Access 11 (2023) 30575–30590.
Natural Language Technologies for dealing with dis- and [13] G. Xiang, B. Fan, L. Wang, J. Hong, C. Rose,
Detectmisinformation with grant reference (CIPROM/2021/21) ing ofensive tweets via topical feature discovery
by the Generalitat Valenciana. over a large scale twitter corpus, in: Proceedings
of the 21st ACM International Conference on
InReferences formation and Knowledge Management, 2012, pp.
1980–1984.
[1] M. Konieczny, Ignorance, disinformation, manipu- [14] P. Burnap, M. L. Williams, Cyber hate speech on
lation and hate speech as efective tools of political twitter: An application of machine classification
power, Policija i sigurnost 32 (2023) 123–134. and statistical modeling for policy and decision
[2] A. J. Stewart, N. McCarty, J. J. Bryson, Polariza- making, Policy &amp; internet 7 (2015) 223–242.
tion under rising inequality and economic decline, [15] C. Nobata, J. Tetreault, A. Thomas, Y. Mehdad,
Science advances 6 (2020) eabd4201. Y. Chang, Abusive language detection in online
[3] N. Gidron, J. Adams, W. Horne, How ideology, eco- user content, in: Proceedings of the 25th
internomics and institutions shape afective polarization national conference on world wide web, 2016, pp.
in democratic polities, in: Annual conference of 145–153.</p>
        <p>the American political science association, 2018. [16] F. M. Plaza-del Arco, A. B. P. Portillo, P. L. Úda,
[4] J. Grieve, H. Woodfield, The language of fake news, B. Gil, M.-T. Martín-Valdivia, SHARE: A lexicon
Cambridge University Press, 2023. of harmful expressions by Spanish speakers, in:
Proceedings of the Thirteenth Language Resources 766.</p>
        <p>and Evaluation Conference, 2022, pp. 1307–1316. [29] A. Appelman, S. S. Sundar, Measuring message
[17] M. Corazza, S. Menini, E. Cabrio, S. Tonelli, S. Vil- credibility: Construction and validation of an
exlata, A multilingual evaluation for online hate clusive scale, Journalism &amp; Mass Communication
speech detection, ACM Transactions on Internet Quarterly 93 (2016) 59–79.</p>
        <p>Technology (TOIT) 20 (2020) 1–22. [30] H. Rashkin, E. Choi, J. Y. Jang, S. Volkova, Y. Choi,
[18] V. Kolhatkar, H. Wu, L. Cavasso, E. Francis, Truth of varying shades: Analyzing language in
K. Shukla, M. Taboada, The SFU opinion and com- fake news and political fact-checking, in:
Proceedments corpus: A corpus for the analysis of online ings of the 2017 conference on empirical methods in
news comments, Corpus Pragmatics 4 (2020) 155– natural language processing, 2017, pp. 2931–2937.
190. [31] G. Americans, Medición del impacto de la
informa[19] F. Huang, H. Kwak, J. An, Chain of explanation: ción falsa, la desinformación y la propaganda en
New prompting method to generate quality natural américa latina, Global Americans (2021).
language explanation for implicit hate speech, in: [32] M. Cinelli, A. Pelicon, I. Mozetič, W. Quattrociocchi,
Companion Proceedings of the ACM Web Confer- P. K. Novak, F. Zollo, Dynamics of online hate and
ence 2023, 2023, pp. 90–93. misinformation, Scientific reports 11 (2021) 22083.
[20] S. Rosenthal, P. Atanasova, G. Karadzhov, [33] J. Y. Kim, A. Kesari, Misinformation and hate
M. Zampieri, P. Nakov, A large-scale semi- speech: The case of anti-asian hate speech during
supervised dataset for ofensive language the covid-19 pandemic, Journal of Online Trust and
identification, arXiv preprint arXiv:2004.14454 Safety 1 (2021).</p>
        <p>(2020). [34] S. Chawda, A. Patil, A. Singh, A. Save, A novel
[21] C. Arcila-Calderón, J. J. Amores, P. Sánchez- approach for clickbait detection, in: 2019 3rd
InHolgado, D. Blanco-Herrero, Using shallow and ternational conference on trends in electronics and
deep learning to automatically detect hate moti- informatics (ICOEI), IEEE, 2019, pp. 1318–1321.
vated by gender and sexual orientation on twitter [35] F. A. Wilson, Enough Already!: A Socialist
Femiin Spanish, Multimodal technologies and interac- nist Response to the Re-emergence of Right Wing
tion 5 (2021) 63. Populism and Fascism in Media, Brill, 2020.
[22] D. Fallis, The varieties of disinformation, The</p>
        <p>philosophy of information quality (2014) 135–161.
[23] B. C. Stahl, On the diference or equality of
information, misinformation, and disinformation: A critical
research perspective, Informing Science 9 (2006)
83.
[24] S. Vosoughi, D. Roy, S. Aral, The spread of true and</p>
        <p>false news online, science 359 (2018) 1146–1151.
[25] X. Zhou, R. Zafarani, A survey of fake news:
Fundamental theories, detection methods, and
opportunities, ACM Computing Surveys (CSUR) 53 (2020)
1–40.
[26] S. Mottola, Las fake news como fenómeno social.</p>
        <p>análisis lingüístico y poder persuasivo de bulos en
italiano y español, Discurso &amp; Sociedad (2020) 683–
706.
[27] A. X. Zhang, A. Ranganathan, S. E. Metz, S. Appling,</p>
        <p>C. M. Sehat, N. Gilmore, N. B. Adams, E. Vincent,
J. Lee, M. Robbins, et al., A structured response
to misinformation: Defining and annotating
credibility indicators in news articles, in: Companion
Proceedings of the The Web Conference 2018, 2018,
pp. 603–612.
[28] B. Horne, S. Adali, This just in: Fake news packs a
lot in title, uses simpler, repetitive content in text
body, more similar to satire than real news, in:
Proceedings of the international AAAI conference
on web and social media, volume 11, 2017, pp. 759–</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>