<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Spanish FatPhoCorpus 2023: Combating Fatphobia in Social Media in Spanish using Transformers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>José Antonio García-Díaz</string-name>
          <email>joseantonio.garcia8@um.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ronghao Pan</string-name>
          <email>ronghao.pan@um.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Salud María Jiménez-Zafra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael Valencia-García</string-name>
          <email>valencia@um.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department, SINAI, CEATIC, Universidad de Jaén</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Departamento de Informática y Sistemas, Facultad de Informática, Universidad de Murcia</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Social media can aggravate body dissatisfaction and weight discrimination. Fat-shaming content is widespread on social media and targets people of diferent weights. Fatphobia and its harmful efects afect not only the general public but also children, causing lasting psychological damage. In this work, we make two contributions to the detection of fatphobia in social networks. On the one hand, we compile the Spanish FatPhoCorpus 2023, a multiclass dataset that allows the identification of hate speech, ofensive content, and hopeful messages related to body image. On the other hand, we evaluate this novel dataset with a baseline of several pre-trained models based on Transformers. The best result is obtained with multilingual TwHIN, which achieves a macro average F1-score of 59.158%.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Fatphobia detection</kwd>
        <kwd>Hate speech detection</kwd>
        <kwd>Hope speech detection</kwd>
        <kwd>Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>This section describes the state of the art concerning</title>
        <p>
          fatphobia (see Section 2.1) and the diferent labels
considered in our study, namely ofensiveness (see Section
2.2), hate speech (see Section 2.3), and hope speech (see language is usually divided into three classes: (i)
“profanSection 2.4). ity”, i.e. the of use of swear words (e.g. fuck), but without
the intention to insult someone; (ii) “insult”, when there
2.1. Fatphobia is a clear intention to ofend someone with disrespect
and contempt; and (iii) “abuse”, when a person is insulted
As far as we know, there is not much work on fatphobia by using a special type of degradation that is considered
from an NLP perspective. One of the papers identified representative of a group by negatively attributing to
was conducted by [
          <xref ref-type="bibr" rid="ref5">8</xref>
          ], in which they study the efect him/her a quality related to a universal, pervasive or
of shutting down two Reddit channels that incite hate: immutable characteristic [14].
(i) r/fatpeoplehate, a subreddit dedicated to posting Ofensive language detection has become one of the
pictures of overweight people to ridicule them and (ii) most popular research areas in the field of NLP, as it
r/CoonTown, a racist subreddit with violent hate speech is increasingly observed in social media. It is often
foragainst African Americans. To do this, they compiled a mulated as a binary (ofensive/non-ofensive) or
multidataset of all posting activity on Reddit in 2015 and used categorization task, considering the types of
ofensivethe texts from the two aforementioned subreddits to build ness and the targets of the ofensive speech [15].
a lexicon of hate words, which they used to examine the Several shared tasks have been organized to promote
user-level efects of the ban. the detection of this type of discourse, such as GermEval
        </p>
        <p>
          Other existing work deals with fatphobia as a cate- 2018 [16] and GermEval 2019 [17] on German tweets,
gory of hate speech or prejudice. [
          <xref ref-type="bibr" rid="ref6">9</xref>
          ] provided HateBR, OfensEval 2019 [ 14] on English texts, OfensEval 2020
a corpus of Brazilian Portuguese Instagram comments [18] on texts written in Arabic, Danish, English, Greek
for the detection of ofensive language and hate speech. and Turkish, and MeOfendES 2021 [ 15] on texts written
The corpus consists of 7000 annotated documents. Later, in Spanish variants.
the authors extended the work by generating a
specialized lexicon manually extracted by a linguist from the 2.3. Hate speech
HateBR corpus [
          <xref ref-type="bibr" rid="ref7">10</xref>
          ]. This Brazilian Portuguese lexicon
was annotated with contextual information, and trans- Hate speech is the language that targets a person or group
lated and adapted by native speakers into Turkish, Ger- with the intention of causing harm or social disruption
man, French, English and Spanish. This context-aware [19]. This targeting is usually done on the basis of some
lexicon is called “MOL-Multilingual Ofensive Lexicon”. characteristic such as race, gender, sexual orientation,
In addition, the authors conducted experiments using nationality or religion [20]. It is a class of ofensive
lanthe generated corpus and lexicon for the identification of guage, specifically “abuse”, and is sometimes referred to
hate speech in Brazilian Portuguese. [11] organized the as abusive language [21].
shared task HUHU in the IberLEF 2023 workshop [12], The negative impact of the spread of hate content has
on the Detection of Humor Spreading Prejudice in Twit- led to an increasing number of researchers focusing on
ter. They proposed a framework to study how humor is this issue. Most studies focus on the detection of hate
used to discriminate against minorities and analyze its content in general [13], the identification of racism [ 22],
interaction with the level of prejudice expressed against the detection of misogyny [23], and the identification of
specific groups. For this, it is provided a corpus of preju- xenophobia [24]. In fact, there are many shared tasks on
diced tweets in Spanish annotated with the presence of the identification of hate speech, such as HASOC 2019
humor, their degree of prejudice and the targeted groups, [25] and HASOC 2020 [26], and shared tasks on specific
being overweight people (fatphobia) one of them. topics such as sexism, toxicity or racism. For example,
        </p>
        <p>It is worth noting that previous work has treated fat- the AMI shared task on the automatic identification of
phobia as a category within hate speech. Our work difers misogyny at IberEval 2018 [27] and Evalita 2018 [28],
from previous work in that we treat fatphobia as a major the HatEval shared task on the detection of hate speech
theme to identify hateful, ofensive or hopeful messages. against immigrants and women [29], and some shared
In addition, we focus on the Spanish language. task organized in the framework of the IberLEF
workshop [30, 31, 12] on (i) the identification of sexism in
2.2. Ofensiveness social networks, as in EXIST 2021 [32] and 2022 [33],
(ii) the detection of toxicity in DETOXIS 2021 [34]), (iii)
the detection and classification of racial stereotypes in
DETESTS 2022 [35], and (iv) the identification of hate
speech towards the LGBTQ+ population in HOMO-MEX
2023 [36].</p>
      </sec>
      <sec id="sec-1-2">
        <title>Ofensiveness is the fact of being rude in a way that makes</title>
        <p>someone to feel upset or angry because it shows a lack
of respect. A text is considered ofensive if it contains
any form of unacceptable language, that is, whether it
contains insults, threats, or bad language [13]. Ofensive</p>
        <sec id="sec-1-2-1">
          <title>2.4. Hope speech</title>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>Hope speech is the type of speech that is able to relax</title>
        <p>a hostile environment [37] and that helps, gives
suggestions and inspires people for good when they are in times
of illness, stress, loneliness or depression [38].</p>
        <p>In the current digital age, social media provides a
place for people to freely express their views and
opinions. They have become critical for people from minority
groups seeking help and support online [39, 12]. People
have turned to these media to satisfy their informational,
emotional, and social needs as they seek to connect with
others, experience a sense of social inclusion, and
cultivate a sense of belonging through active participation
in online communities. The presence of these factors
has a profound impact on physical and psychological
well-being as well as mental health [40, 41]. This has
led to recent research exploring positive content, such
as messages of hope, and the promotion of constructive
activities in pursuit of equality, diversity, and inclusion.</p>
        <p>There has been a remarkable increase in attention to
this topic since 2021 and several workshops have been
organized to address the challenge of detecting hope
speech in English, Tamil, Malayalam, Bulgarian, Hindi
and Spanish, such as LT-EDI-EACL2021 [42],
LT-EDIACL2022 [43], LT-EDI-RANLP2023 [44] and the shared
task HOPE in IberLEF 2023 [12].</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Methodology</title>
      <sec id="sec-2-1">
        <title>In the following section, the compilation and annotation process of the Spanish FatPhoCorpus 2023 is described in Section 3.1 and the pipeline for the evaluation of the dataset is presented in Section 3.2.</title>
        <sec id="sec-2-1-1">
          <title>3.1. Spanish FatPhoCorpus 2023 compilation and annotation</title>
          <p>The UMUCorpusClassifier tool [ 45] was used to compile
the dataset. We started by querying for specific keywords
on X (formerly Twitter). The keywords are gordofobia
(fatphobia), sobrepeso (overweight), gordo/a (fat), obeso
(obese), obesidad (obesity), foca (seal), glotón (glutton),
glotonería (gluttony), “talla grande” (plus size), anorexia,
lfaco (skinny), “desorden alimenticio” (‘eating disorder),
delgado (thin), delgadez (thinnes), and calorías (calories).</p>
          <p>The dataset is compiled from all countries where Spanish
is spoken. This can be challenging due to diferences
in cultural factors between Spanish-speaking countries,
resulting in diferent interpretations of adjectives such
as “gordo” (fat) or “flaco” (thin). In the first iteration,
39,787 tweets were collected from June 2020 to July 2023,
focusing on the specified keywords.</p>
          <p>The annotation phase was performed by the research
team and an external collaborator. We considered the
following labels: (i) Hate, if the text contains hate speech
towards people because of their body, (ii) Hope, if the
text tends to help, relax, and inspire people for their
body weight perception, (iii) Ofensive , if the text contains
insults and profanity but no fatphobia, and (iv) None, if
none of the rest.</p>
          <p>Each tweet was annotated by four members, being two
men and two women, members of our research group.</p>
          <p>The annotators do not have specific knowledge of the
area, but rather linguistics and computer science. The
inter-annotator agreement based on Krippendorf’s alpha
is 0.712. The final label is assigned based on the mode of
the annotations. Ties were resolved in internal meetings.</p>
          <p>The analysis of the dataset revealed that some of the
offensive tweets contain homophobia, racism or
transphobia in addition to fatphobia. Furthermore, some tweets
contain irony and figurative language, while others are
self-deprecating. Fourth, we observed a lot of
internalized fatphobia, as several users wrote that they are very
afraid of being fat because they associate being fat with
being ugly. However, the main finding is that several
tweets used "thin" or "fat" in a familiar way to refer to
partners and pets in an afective way and these tweets
were labeled as None.</p>
          <p>To select the number of tweets that are not related
0%
25%
50%
75%
100%
to ofensive, hate, or hope, we use agglomerative clus- Figure 1 shows four examples from the dataset, one for
tering with the tweets labeled None. The idea is to each label. The top left example is for a text labeled as
get a large but representative number of tweets of dif- hate-speech, which contains fatphobia but also misogyny.
ferent types. To do this, we extract their contextual The top right example promotes hope. It is a statement
sentence embeddings, obtained by hackathon-pln-es/ from a person who refuses to give its opinion about other
paraphrase-spanish-distilroberta, and create five clusters. people’s bodies. The bottom left example is labeled as
We then randomly selected one tweet for each cluster None, and it is about a person who regrets eating junk
until the cluster with the fewest instances was empty. food and finds it hard to stop. The bottom right example
This process reduced the number of tweets marked as is labeled as Ofensive . It is about a person who curses
None to 4,570. and insults another man and threatens him with physical</p>
          <p>In the final version of the dataset, a total of 6,145 tweets violence.
were collected, annotated, and verified. The remaining To analyze the corpus, we used the UMUTextStats tool
tweets were removed because they were too short or it [46] to obtain linguistic features to calculate the
informawas not clear how to annotate them. These tweets are tion gain for each label. As Figure 2 shows, soft ofensive
divided into training, validation, and test sets using a language and swear words, psychological processes
re70-15-15 ratio, with stratification to maintain the balance lated to anger, and lexis related to sex show a correlation
of labels between the splits. Table 1 shows the statistics with hate speech, and to a lesser extent, with ofensive
for each label and split. Note that hope language is the language. The only exception is the food-related
termilabel with the lowest representation, while the remaining nology, which is less common in documents labeled as
labels have comparable proportions. Hate but is more common in hope speech and fatphobia
unrelated tweets. In addition, we observe that analytical
Table 1 thinking is used almost equally in all labels.
Dataset distribution for label and split. Finally, the dataset is now available to the scientific
community1. However, according to the X’s guidelines2,
Label Train Val Test Total we will only publish the tweet IDs to protect the users’
right to be forgotten.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>1https://pln.inf.um.es/corpora/hate-speech/hate-fatphobia</title>
        <p>2024.zip
2https://developer.twitter.com/en/developer-terms/policy</p>
        <sec id="sec-2-2-1">
          <title>3.2. System architecture</title>
          <p>is usually small, with no steps for ALBETO, BERTIN,
DistilBETO and TwHIN.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>To evaluate the dataset we rely on diferent pre-trained</title>
        <p>models for this classification task. For this purpose, the
system shown in Figure 3 is developed. In short, the Table 2
Hyperparameter tuning of the seven evaluated LLMs. The
pipeline can be described as follows. First, the data pre- parameters evaluated are the learning rate (lr), number of
processing stage is performed on the dataset to produce training epochs (e), the batch size (bs), the warm-up steps
a proper and clean format, since the quality of the input (ws), and the weight decay (wd).
data directly afects the performance and generalizability
of the model. In this case, we perform a data cleaning LLM lr e bs ws wd
stage where acronyms are expanded, elongations, and ALBETO 2.2e-05 5 8 0 0.21
digits are removed along with hyperlinks, hashtags, quo- BERTIN 1.1e-05 5 16 0 0.11
tation marks, and other punctuation symbols. Finally, we BETO 4e-05 3 8 500 0.23
evaluated diferent pre-trained models based on Trans- DistilBETO 3e-05 4 16 0 0.26
formers and with diferent architectures, such as BERT, MarIA 4.4e-05 4 8 250 0.18
RoBERTa, ALBERT and DistilBERT to classify diferent TwHIN 2.6e-05 3 8 0 0.17
categories using the fine-tuning approach. In this case, RoBERTuito 2.1e-05 5 16 500 0.16
since all the evaluated models are of the encoder type,
we have added a sequence classification layer to perform
the fine-tuning process.</p>
        <p>
          The fine-tuning process is performed using hyper- 4. Results and discussion
parameter optimization stage to obtain the learning rate
(uniformly sampled between 1e-5 and 5e-5), the num- In the experiments, a simple process is performed by
ber of epochs (between 1 and 5), the batch size ([
          <xref ref-type="bibr" rid="ref5">8, 16</xref>
          ]), ifne-tuning several pre-trained models based on
Transthe warm-up steps ([0, 250, 500, 1000]), and the weight formers. The models evaluated are of the encoder type,
decay (uniformly sampled between 0.0 and 0.3). A total i.e. means that they use only the encoder component of
of 10 runs are evaluated for each LLM. Table 2 shows a Transformers model. These models are often
characthe best hyperparameters for each LLM. The lower num- terized by having “bidirectional” attention, which is why
ber of training epochs is achieved by BETO and TwHIN. they are also called “auto-encoding” models. The
preBoth without no warm-up steps. Other models, includ- training of these models typically involves corrupting a
ing lightweight models such as ALBETO and DistilBETO, given sentence in some way, for example, by masking
require a larger number of epochs (5 and 4 respectively). random words in it (Masked Language Modeling), and
Another finding is that the number of warm-up steps then instructing the model to find or reconstruct the
original sentences. Therefore, depending on the language of training time.
the pre-training corpus, the models can be monolingual Next, we report the classification report of the best
or multilingual. model, TwHIN. The Table 4 shows the precision, recall,
        </p>
        <p>The monolingual models evaluated are: (i) BETO [47], and F1-score with the test set. It can be seen that TwHIN
a Spanish BERT model trained on the Spanish Unanno- performs quite well in predicting hate speech with an
tated Corpora; (ii) MarIA [48], which is a pre-trained F1-score of 64.945%, but the model is less reliable in
premodel based on RoBERTa and trained exclusively on dicting hope and ofensive content from a fatphobic
perSpanish texts collected from web crawling of the Na- spective. Furthermore, all labels behave similarly in terms
tional Library of Spain; (iii) BERTIN [49], another model of precision and recall.
based on RoBERTa and trained on the Spanish part of We used TwHIN to perform the error analysis because
the mC4 dataset; (iv) ALBETO [50], a pre-trained model it is the fine-tuned model that produced the highest macro
based on ALBERT (a lightweight version of BERT), pre- F1-score over the test split. Figure 4 shows the confusion
trained only on Spanish documents; (v) DistilBETO [50], matrix for this model, which allows us to identify cases
is a model trained using distillation techniques to trans- where the model makes incorrect predictions. We can see
fer the weights of BETO to a new model with fewer that the model misclassified 27 instances of hate speech
layers and reduced complexity; (vi) RoBERTuito [51], a as ofensive speech. Analyzing these cases, it was found
pre-trained model based on RoBERTa for the analysis of that when the level of hate is very high, our model tends
social media text in Spanish, and trained according to to identify it as ofensive. It was also observed that the
RoBERTa guidelines on 500 million tweets. In terms of analyzed model has dificulties to handle hope speech as
multilingual models, TWHIN [52] was evaluated, which it confused 5 instances with hate and 4 with none. The
is a multilingual model trained on X covering over 100 case of ofensive speech is also confused with hate (29
languages. cases) and none (21 cases).</p>
        <p>In this case, since these are encoder-type models, we It is worth noting that in an earlier version of the
anhave added a sequence classification layer to perform notation processes, we identify a large number of posts
ifne tuning for a classification task. This layer consists coming from diferent Spanish-speaking countries,
inof dense structure with as many neurons as there are cluding Spain, Argentina, Mexico, and so on. In this
output classes to classify. To evaluate the classifiers, we sense, there are adjectives with ambiguous meanings,
use the macro F1-score as the main benchmark metric, such as “flaco”, which is an informal way for Argentines
which weighs the precision and recall of each class and to refer to a man or a woman in an afectionate way, even
combines the results without considering the class imbal- if they are not thin. Other examples are the word “gorda”
ance. In this way, we select the best model that performs in text 3 and the word “flaco”. In the first version of the
equally well for each label. Another reference metric is dataset, these tweets were classified as Hope (leading to
the weighted F1-score (W-F1), which is an evaluation false positives), when in fact, they should have been
clasmeasure used in classification problems, especially when sified as None. In addition to the cultural background
dealing with unbalanced datasets, where some classes of diferent Spanish-speaking countries, we identified
may have many more examples than others. Unlike the another challenge related to the fact that people use
inmacro F1-score, this metric takes into account the class formal language on social networks, in addition to not
imbalance by assigning diferent weights to each class using punctuation or accents correctly. This is
particubased on their frequency in the dataset. larly serious in the case of Spanish, as Spanish does not</p>
        <p>Table 3 shows the results of all the pre-trained models. change the order of words when moving from declarative
In general, all the evaluated models performed equally to interrogative sentences. Often, this meaning can be
well in terms of macro and weighted F1-score. The re- taken out of context with considerable human efort, but
sults using the weighted precision, recall and F1-score are it makes it very dificult for NLP tools to understand the
higher due to the class imbalance and the reliability of all language.
models with the None label. Among the evaluated mod- Table 5 shows some examples of misclassifications
els, the multilingual TwHIN is the best performing model, made by TwHIN that summarize the challenge of
idenwith a weighted F1-score of 83.234% and a macro F1-score tifying hate, hope, and aggressive fatphobic comments.
of 59.158%. The monolingual models, BETO and MarIA, The first example contains a message suggesting that
based on the BERT and RoBERTa architectures respec- normalizing the message of being okay with one’s body
tively, also performed similarly, with a macro F1 score of is counterproductive. This example is dificult to identify
0.603% improvement of MarIA over BETO. Furthermore, because it uses polite language. The second example is
it is also worth noting that ALBETO (the lightweight considered ofensive by the model because it contains
version of BETO) outperformed distilled BETO (macro many derogatory terms, but it also contains terms related
F1-score of 52.159% vs 48.071). In this sense, ALBETO to fatphobia as well as homophobia. The third example,
gave similar results to BETO, with a significantly reduced annotated as hope speech, is incorrectly identified by
e
t
a
h</p>
      </sec>
      <sec id="sec-2-4">
        <title>TwHIN as hate speech. This example contains dificult</title>
        <p>words such as because it contains a slang from Chile and Predicted
Peru (weon) that refers to stupid and lazy people and
because it contains negation cues. The fourth example
was identified as None. This is correct because it uses Figure 4: Confusion matrix of the TwHIN model
the word flaco as colloquial way. However, it uses the
adjective Mongolian, which is used as an insult in Spain.</p>
        <p>However, this word contains a typo so it is possible that dataset for the detection of fatphobic comments in
soTwHIN does it not consider an ofensive word. The last cial networks in Spanish. This work has considered four
example is hate speech, which was misclassified classi- types of comments, according to whether they contain
ifed as hope speech. The author is complaining about the hate speech, hope speech, ofensiveness, or none of the
so-called crystal generation regarding issues of racism, above. The resulting dataset has been evaluated with
fatness and homophobia. In this case, the problem is that diferent pre-trained models based on Transformers,
obthe author it does only states a premise but the conclusion taining the best result TwHIN, a multilingual general
puris not clear. pose model trained with tweets, with a macro F1 score</p>
        <p>Finally, to get some insight into the errors made by of 59.158%.</p>
        <p>TwHIN, we get the subset of the test field that was mis- As future work, we will expand the dataset. One
limitaclassified and obtain the linguistic features (see Figure tion we found during the evaluation was the dificulty in
5). We observed that the errors contain morphological properly identifying hope and ofensive fatphobic tweets.
features, such as interjections, the use of verbs in the im- This is particularly relevant in the case of hope speech,
perative and in the infinitive; besides, we also observed where the dataset contains only 84 texts. Furthermore,
features related to negativity and that the number of we found that background and cultural diferences are
words and syllables is also relevant. very important for detecting fatphobia in social networks.
Therefore, we propose the annotation of the
Spanishspeaking country of the authors and evaluate the cultural
5. Conclusions diferences when it comes to finding similarities and
difIn this paper we have described the compilation and an- ferences between diferent Spanish-speaking countries.
notation process of the Spanish FatPho Corpus 2023, a In this sense, we will extend the annotation of the dataset
No es por ser gordofobico nada pero amigo en usa tratan de normalizar el "se tu mismo" oh ese estilo de cosas?.
Recomiendo ver el video de tri-line hacerca de ese tema pero en mi sincera opinion,no esta mal ser gordo pero
es mejor bajar de peso. (It is not because I am fatphobic at all, but my friend in the USA they try to normalize "be
yourself" oh that style of things? I recommend watching the Tri-Line video on this topic, but in my honest opinion, it
is not bad to be fat, but it is better to lose weight.)
Nunca nos olvidaremos a la derecha corrupta llevo a la ruina al país en cabeza del gordo marica q el uribismo
puso ahí unas ratas ( We will never forget that the corrupt right led the country to ruin at the head of the fat faggot
that Uribismo put some rats there)
Esta vez no estoy de acuerdo con su comentario, y debo decir que es bastante weon el comentario, la obesidad
se da no solo por comer mucho o una mala alimentacion, deberia informarse antes de decir esa estupidez, y se
gano un seguidor menos. (This time I do not agree with his comment, and I must say that the comment is quite bad,
obesity is not only caused by eating a lot or bad diet, you should inform yourself before saying that stupid thing,
and you gained one less follower.!)
Pero flaco ustedes solamente pagan cuota de socio, con eso estás adentro, nosotros abono para NUESTRA
cancha, y si no tenés abono y sos socio, tenés q pagar la entrada para ir al kempes, mogolico (But "flaco", you
only pay the membership fee, with that you are in, we have a subscription for OUR field, and if you don’t have a
subscription and you are a member, you have to pay the entrance fee to go to the Kempes, you idiot.)
La generación de cristal es bastante rara. Se molestan cuando uno le dice negro, flaco o gordo a alguien, o
cuando se les dice que solo existe hombre o mujer en términos biológicos... (The crystal generation is quite rare.
They get upset if you call someone black, thin, or fat, or if you tell them that there is only one biological male or
female...)
T
Ha
Ha
Ho</p>
        <p>P
No
Of</p>
        <p>Ha
Of</p>
        <p>No
Ha</p>
        <p>Ho
by including volunteers from diferent Spanish-speaking gional Support Program for the Transfer and
Valorizacountries. Finally, regarding the evaluation of the mod- tion of Knowledge and Scientific Entrepreneurship of the
els, we will evaluate the feature integration to combine Seneca Foundation, Science and Technology Agency of
the strengths of diferent pre-trained models based on the Region of Murcia. The research work conducted by
Transformers. Salud María Jiménez-Zafra has been supported by Action
7 from Universidad de Jaén under the Operational Plan
for Research Support 2023-2024. Mr. Ronghao Pan is
Acknowledgments supported by the Programa Investigo grant, funded by
the Region of Murcia, the Spanish Ministry of Labour and
This work has been partially supported by projects
Social Economy and the European Union -
NextGeneraLaTe4PoliticES (PID2022-138099OB-I00) funded by
MItionEU under the "Plan de Recuperación, Transformación
CIU/AEI/10.13039/501100011033 and the European
Rey Resiliencia (PRTR)".
gional Development Fund (ERDF)-a way of making
Europe, LT-SWM (TED2021-131167B-I00) funded by
MICIU/AEI/10.13039/ 501100011033 and by the European References
Union NextGenerationEU/PRTR, CONSENSO
(PID2021122263OB-C21), MODERATES (TED2021-130145B-I00), [1] T. F. Cash, Body image: Past, present, and future,
and SocialTOX (PDC2022-133146-C21) funded by Plan 2004.</p>
        <p>Nacional I+D+i from the Spanish Government, PRE- [2] B. B. E. Robinson, L. C. Bacon, J. O’reilly, Fat phobia:
COM (SUBV-00016) funded by the Ministry of Consumer Measuring, understanding, and changing anti-fat
Afairs of the Spanish Government, FedDAP (PID2020- attitudes, International Journal of Eating Disorders
116118GA-I00) and Trust-ReDaS (PID2020-119478GB-I00) 14 (1993) 467–480.
supported by MICINN/AEI/10.13039/501100011033, and [3] C. N. Wagner, E. Aguirre Alfaro, E. M. Bryant, The
"Services based on language technologies for political mi- relationship between instagram selfies and body
crotargeting" (22252/PDC/23) funded by the Autonomous image in young adult women, First Monday 21
Community of the Region of Murcia through the Re- (2016).</p>
        <p>(MOR) morphology-interjections
(MOR)
morphology-verbs</p>
        <p>imperative
(MOR)
morphology-verbsindicative-simple-imperative
0%
25%
50%
75%
100%
samiento del Lenguaje Natural 67 (2021) 183–194. ifre 2019: Hate speech and ofensive content
identi[16] M. Wiegand, M. Siegel, J. Ruppenhofer, Overview ifcation in indo-european languages, in:
Proceedof the germeval 2018 shared task on the identifi- ings of the 11th forum for information retrieval
cation of ofensive language, in: 14th Conference evaluation, 2019, pp. 14–17.
on Natural Language Processing KONVENS 2018, [26] T. Mandl, S. Modha, A. Kumar M, B. R. Chakravarthi,
2018, pp. 1–10. Overview of the hasoc track at fire 2020: Hate
[17] J. M. Struß, M. Siegel, J. Ruppenhofer, M. Wiegand, speech and ofensive language identification in
M. Klenner, Overview of germeval task 2, 2019 tamil, malayalam, hindi, english and german, in:
shared task on the identification of ofensive lan- Forum for Information Retrieval Evaluation, 2020,
guage, in: Proceedings of the 15th Conference pp. 29–32.
on Natural Language Processing (KONVENS 2019), [27] E. Fersini, P. Rosso, M. Anzovino, Overview of
German Society for Computational Linguistics &amp; the task on automatic misogyny identification at
Language Technology und . . . , 2019, pp. 352–363. ibereval 2018., IberEval SEPLN 2150 (2018) 214–228.
[18] M. Zampieri, P. Nakov, S. Rosenthal, P. Atanasova, [28] C. Bosco, D. Felice, F. Poletto, M. Sanguinetti,
G. Karadzhov, H. Mubarak, L. Derczynski, Z. Pitenis, T. Maurizio, Overview of the evalita 2018 hate
Ç. Çöltekin, Semeval-2020 task 12: Multilingual speech detection task, in: EVALITA 2018-Sixth
Evalofensive language identification in social media uation Campaign of Natural Language Processing
(ofenseval 2020), in: Proceedings of the Fourteenth and Speech Tools for Italian, volume 2263, CEUR,
Workshop on Semantic Evaluation, 2020, pp. 1425– 2018, pp. 1–9.</p>
        <p>1447. [29] V. Basile, C. Bosco, E. Fersini, N. Debora, V. Patti,
[19] J. A. García-Díaz, S. M. Jiménez-Zafra, M. A. García- F. M. R. Pardo, P. Rosso, M. Sanguinetti, et al.,
Cumbreras, R. Valencia-García, Evaluating feature Semeval-2019 task 5: Multilingual detection of hate
combination strategies for hate-speech detection in speech against immigrants and women in twitter,
spanish using linguistic features and transformers, in: 13th International Workshop on Semantic
EvalComplex &amp; Intelligent Systems 9 (2023) 2893–2914. uation, Association for Computational Linguistics,
[20] A. Schmidt, M. Wiegand, A survey on hate speech 2019, pp. 54–63.</p>
        <p>detection using natural language processing, in: [30] J. Gonzalo, M. Montes-y Gómez, P. Rosso, Iberlef
Proceedings of the fifth international workshop on 2021 overview: Natural language processing for
natural language processing for social media, 2017, iberian languages, in: Proceedings of the Iberian
pp. 1–10. Languages Evaluation Forum (IberLEF 2021), CEUR
[21] H. Sohn, H. Lee, Mc-bert4hate: Hate speech de- Workshop, 2021, pp. 1–15.</p>
        <p>tection using multi-channel bert for diferent lan- [31] J. Gonzalo, M. Montes-y Gómez, F. Rangel,
guages and translations, in: 2019 International Overview of iberlef 2022: Natural language
proConference on Data Mining Workshops (ICDMW), cessing challenges for spanish and other iberian
IEEE, 2019, pp. 551–559. languages, in: Proceedings of the Iberian Languages
[22] M. Sap, D. Card, S. Gabriel, Y. Choi, N. A. Smith, Evaluation Forum (IberLEF 2022), CEUR, 2022, pp.</p>
        <p>The risk of racial bias in hate speech detection, in: 1–12.</p>
        <p>Proceedings of the 57th annual meeting of the as- [32] F. Rodríguez-Sánchez, J. Carrillo-de Albornoz,
sociation for computational linguistics, 2019, pp. L. Plaza, J. Gonzalo, P. Rosso, M. Comet, T. Donoso,
1668–1678. Overview of exist 2021: sexism identification in
so[23] J. A. García-Díaz, M. Cánovas-García, R. C. Pala- cial networks, Procesamiento del Lenguaje Natural
cios, R. Valencia-García, Detecting misogyny in 67 (2021) 195–207.
spanish tweets. an approach based on linguistics [33] F. Rodríguez-Sánchez, J. Carrillo-de Albornoz,
features and word embeddings, Future Gener. L. Plaza, A. Mendieta-Aragón, G. Marco-Remón,
Comput. Syst. 114 (2021) 506–518. URL: https:// M. Makeienko, M. Plaza, J. Gonzalo, D. Spina,
doi.org/10.1016/j.future.2020.08.032. doi:10.1016/ P. Rosso, Overview of exist 2022: sexism
idenj.future.2020.08.032. tification in social networks, Procesamiento del
[24] F.-M. Plaza-Del-Arco, M. D. Molina-González, L. A. Lenguaje Natural 69 (2022) 229–240.</p>
        <p>Ureña López, M. T. Martín-Valdivia, Detecting [34] M. Taulé, A. Ariza, M. Nofre, E. Amigó, P. Rosso,
misogyny and xenophobia in spanish tweets using Overview of detoxis at iberlef 2021: detection of
language technologies, ACM Trans. Internet Tech- toxicity in comments in spanish, Procesamiento
nol. 20 (2020). URL: https://doi.org/10.1145/3369869. del lenguaje natural 67 (2021) 209–221.
doi:10.1145/3369869. [35] A. Ariza-Casabona, W. S. Schmeisser-Nieto,
[25] T. Mandl, S. Modha, P. Majumder, D. Patel, M. Dave, M. Nofre, M. Taulé, E. Amigó, B. Chulvi, P. Rosso,
C. Mandlia, A. Patel, Overview of the hasoc track at Overview of detests at iberlef 2022: Detection
and classification of racial stereotypes in spanish, //aclanthology.org/2022.ltedi-1.58. doi:10.18653/
Procesamiento del lenguaje natural 69 (2022) v1/2022.ltedi-1.58.</p>
        <p>217–228. [44] P. Kumaresan, B. R. Chakravarthi, S. Cn, M. Á.
[36] G. Bel-Enguix, H. Gómez-Adorno, G. Sierra, García, S. M. Jiménez-Zafra, J. A. García-Díaz,
J. Vásquez, S. T. Andersen, S. Ojeda-Trueba, R. Valencia-García, M. Hardalov, I. Koychev,
Overview of homo-mex at iberlef 2023: Hate speech P. Nakov, D. García-Baena, K. Kumar Ponnusamy,
detection in online messages directed towards the B. Preston, Overview of the shared task on hope
mexican spanish speaking lgbtq+ population, Proce- speech detection for equality, diversity, and
insamiento del lenguaje natural 71 (2023) 361–370. clusion, in: Proceedings of the Third
Work[37] S. Palakodety, A. R. KhudaBukhsh, J. G. Carbonell, shop on Language Technology for Equality,
DiHope speech detection: A computational analysis of versity and Inclusion, Association for
Computathe voice of peace, arXiv preprint arXiv:1909.12940 tional Linguistics, 2023, pp. 47–53. doi:10.26615/
(2019). 978-954-452-084-7\_007.
[38] B. R. Chakravarthi, HopeEDI: A multilingual [45] J. A. García-Díaz, Á. Almela, G. Alcaraz-Mármol,
hope speech detection dataset for equality, diver- R. Valencia-García, Umucorpusclassifier:
Compisity, and inclusion, in: Proceedings of the Third lation and evaluation of linguistic corpus for
natuWorkshop on Computational Modeling of People’s ral language processing tasks, Procesamiento del
Opinions, Personality, and Emotion’s in Social Lenguaje Natural 65 (2020) 139–142.
Media, Association for Computational Linguistics, [46] J. A. García-Díaz, P. J. Vivancos-Vicente, A. Almela,
Barcelona, Spain (Online), 2020, pp. 41–53. URL: R. Valencia-García, Umutextstats: A linguistic
feahttps://aclanthology.org/2020.peoples-1.5. ture extraction tool for spanish, in: Proceedings of
[39] D. García-Baena, M. Á. García-Cumbreras, S. M. the Thirteenth Language Resources and Evaluation
Jiménez-Zafra, J. A. García-Díaz, R. Valencia-García, Conference, 2022, pp. 6035–6044.</p>
        <p>Hope speech detection in spanish: The lgbt case, [47] J. Cañete, G. Chaperon, R. Fuentes, J.-H. Ho,
Language Resources and Evaluation (2023) 1–28. H. Kang, J. Pérez, Spanish pre-trained bert model
[40] D. N. Milne, G. Pink, B. Hachey, R. A. Calvo, Clpsych and evaluation data, in: PML4DC at ICLR 2020,
2016 shared task: Triaging content in online peer- 2020.
support forums, in: Proceedings of the third work- [48] A. G. F. no, J. A. Estapé, M. Pàmies, J. L. Palao, J. S.
shop on computational linguistics and clinical psy- Ocampo, C. P. Carrino, C. A. Oller, C. R. Penagos,
chology, 2016, pp. 118–127. A. G. Agirre, M. Villegas, Maria: Spanish language
[41] B. R. Chakravarthi, V. Muralidaran, R. Priyad- models, Procesamiento del Lenguaje Natural
harshini, S. Cn, J. P. McCrae, M. Á. García, S. M. 68 (2022). URL: https://upcommons.upc.edu/
Jiménez-Zafra, R. Valencia-García, P. Kumaresan, handle/2117/367156#.YyMTB4X9A-0.mendeley.
R. Ponnusamy, et al., Overview of the shared task doi:10.26342/2022-68-3.
on hope speech detection for equality, diversity, and [49] J. D. la Rosa, E. G. Ponferrada, M. Romero,
inclusion, in: Proceedings of the second workshop P. Villegas, P. G. de Prado Salas, M. Grandury,
on language technology for equality, diversity and Bertin: Eficient pre-training of a spanish
laninclusion, 2022, pp. 378–388. guage model using perplexity sampling,
Proce[42] B. R. Chakravarthi, V. Muralidaran, Findings of samiento del Lenguaje Natural 68 (2022) 13–23.
the shared task on hope speech detection for equal- URL: http://journal.sepln.org/sepln/ojs/ojs/index.
ity, diversity, and inclusion, in: Proceedings of php/pln/article/view/6403.
the First Workshop on Language Technology for [50] J. Cañete, S. Donoso, F. Bravo-Marquez, A. Carvallo,
Equality, Diversity and Inclusion, Association for V. Araujo, Albeto and distilbeto: Lightweight
spanComputational Linguistics, Kyiv, 2021, pp. 61–72. ish language models, in: Proceedings of the 13th
URL: https://aclanthology.org/2021.ltedi-1.8. Language Resources and Evaluation Conference,
[43] B. R. Chakravarthi, V. Muralidaran, R. Priyad- European Language Resources Association,
Marharshini, S. Cn, J. McCrae, M. Á. García, S. M. seille, France, 2022.</p>
        <p>Jiménez-Zafra, R. Valencia-García, P. Kumaresan, [51] J. M. Pérez, D. A. Furman, L. Alonso Alemany, F. M.
R. Ponnusamy, D. García-Baena, J. García-Díaz, Luque, RoBERTuito: a pre-trained language model
Overview of the shared task on hope speech de- for social media text in Spanish, in: Proceedings of
tection for equality, diversity, and inclusion, in: the Thirteenth Language Resources and Evaluation
Proceedings of the Second Workshop on Lan- Conference, European Language Resources
Associguage Technology for Equality, Diversity and In- ation, Marseille, France, 2022, pp. 7235–7243. URL:
clusion, Association for Computational Linguis- https://aclanthology.org/2022.lrec-1.785.
tics, Dublin, Ireland, 2022, pp. 378–388. URL: https: [52] X. Zhang, Y. Malkov, O. Florez, S. Park,
B. McWilliams, J. Han, A. El-Kishky,
Twhinbert: A socially-enriched pre-trained language
model for multilingual tweet representations at
twitter, arXiv preprint arXiv:2209.07562 (2022).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. L.</given-names>
            <surname>Gonzales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Hancock</surname>
          </string-name>
          , Mirror, mirror
          <article-title>on my Contextual-aware and expert data resources for facebook wall: Efects of exposure to facebook on brazilian portuguese hate speech detection, Reself-esteem, Cyberpsychology, behavior, and social search Square (</article-title>
          <year>2023</year>
          ). URL: https://doi.org/10.21203/ networking 14 (
          <year>2011</year>
          )
          <fpage>79</fpage>
          -
          <lpage>83</lpage>
          . rs.3.rs-
          <volume>2050376</volume>
          /v2.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Stoll</surname>
          </string-name>
          ,
          <article-title>Fat is a social justice issue, too</article-title>
          , Humanity [11]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Tamayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , Everybody hurts,
          <source>&amp; Society</source>
          <volume>43</volume>
          (
          <year>2019</year>
          )
          <fpage>421</fpage>
          -
          <lpage>441</lpage>
          . URL: https://doi.org/ sometimes overview of hurtful
          <source>humour at iberlef 10.1177/0160597619832051</source>
          . 2023:
          <article-title>Detection of humour spreading prejudice in</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Aziz</surname>
          </string-name>
          ,
          <article-title>Social media and body issues in young twitter, Procesamiento del Lenguaje Natural 71 adults: an empirical study on the influence of In- (</article-title>
          <year>2023</year>
          )
          <fpage>383</fpage>
          -
          <lpage>395</lpage>
          .
          <article-title>stagram use on body image and</article-title>
          fatphobia in cata- [12]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Á.</surname>
          </string-name>
          Garcia-Cumbreras, lan university students,
          <source>Master's thesis</source>
          , Universitat D.
          <string-name>
            <surname>García-Baena</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          <string-name>
            <surname>Garcia-Díaz</surname>
            ,
            <given-names>B. R.</given-names>
          </string-name>
          <string-name>
            <surname>Pompeu</surname>
          </string-name>
          <article-title>Fabra (UPF</article-title>
          ),
          <year>2017</year>
          . Chakravarthi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. A</surname>
          </string-name>
          . Ureña-
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Puhl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Suh</surname>
          </string-name>
          ,
          <article-title>Health consequences of weight López, Overview of hope at iberlef 2023: stigma: implications for obesity prevention and Multilingual hope speech detection, Procesamiento treatment</article-title>
          ,
          <source>Current obesity reports 4</source>
          (
          <year>2015</year>
          )
          <fpage>182</fpage>
          - del
          <source>Lenguaje Natural</source>
          <volume>71</volume>
          (
          <year>2023</year>
          )
          <fpage>371</fpage>
          -
          <lpage>381</lpage>
          .
          <fpage>190</fpage>
          . [13]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Plaza-del Arco</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. D.</surname>
            Molina-González,
            <given-names>L. A.</given-names>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Chandrasekharan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Pavalanathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Srini-</surname>
          </string-name>
          Urena-López,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Martín-Valdivia</surname>
          </string-name>
          , Comparvasan,
          <string-name>
            <given-names>A.</given-names>
            <surname>Glynn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Eisenstein</surname>
          </string-name>
          , E. Gilbert,
          <article-title>You can't ing pre-trained language models for spanish hate stay here: The eficacy of reddit's 2015 ban exam- speech detection, Expert Systems with Applications ined through hate speech</article-title>
          ,
          <source>Proceedings of the ACM</source>
          <volume>166</volume>
          (
          <year>2021</year>
          )
          <article-title>114120</article-title>
          .
          <article-title>on human-computer interaction 1 (</article-title>
          <year>2017</year>
          )
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          . [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zampieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Malmasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          , S. Rosenthal,
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>F.</given-names>
            <surname>Vargas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          , F. Rodrigues de Góes,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , Semeval-2019 task 6:
          <string-name>
            <surname>IdentifyT. Pardo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Benevenuto</surname>
          </string-name>
          ,
          <article-title>HateBR: A large expert ing and categorizing ofensive language in social annotated corpus of Brazilian Instagram comments media (ofenseval), in: Proceedings of the 13th for ofensive language and hate speech detection</article-title>
          ,
          <source>International Workshop on Semantic Evaluation, in: Proceedings of the Thirteenth Language Re- 2019</source>
          , pp.
          <fpage>75</fpage>
          -
          <lpage>86</lpage>
          . sources and Evaluation Conference, European Lan- [15]
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Plaza-del Arco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Casavantes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Escalante</surname>
          </string-name>
          , guage Resources Association, Marseille, France, M. T.
          <string-name>
            <surname>Martín-Valdivia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Montejo-Ráez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Montes</surname>
          </string-name>
          ,
          <year>2022</year>
          , pp.
          <fpage>7174</fpage>
          -
          <lpage>7183</lpage>
          . URL: https://aclanthology.org/ H.
          <string-name>
            <surname>Jarquín-Vásquez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Villaseñor-Pineda</surname>
          </string-name>
          , et al.,
          <year>2022</year>
          .lrec-
          <volume>1</volume>
          .
          <fpage>777</fpage>
          . Overview of meofendes at iberlef 2021: Ofen-
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Vargas</surname>
          </string-name>
          , I. Carvalho,
          <string-name>
            <given-names>T. A. S.</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>Benevenuto, sive language detection in spanish variants</article-title>
          , Proce-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>