<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>and Quantification of H Urtful H U mor (H UH U) on Twitter Using Classical Models, Ensemble Models, and Transformers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hugo Albert Bonet</string-name>
          <email>hugoalberthlu@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aina Magraner Rincón</string-name>
          <email>magraneraina@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alba Martínez López</string-name>
          <email>albamartinez584@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Politècnica de València</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Identifying hurtful comments in social media posts has high relevance in order to improve the common welfare. Nevertheless, sometimes hurtful messages are masked by humor, which may increase the dificulty of detecting and identifying this type of content. When making use of HUrtful HUmor (HUHU), the author feels free to spread prejudices without limits [1]. Because of the aforementioned reasons, the objective of this work is to propose a methodology to fasten the detection of harmful texts posted on social media by exploring the diferent machine and deep learning models for three diferent tasks in Spanish: HUrtful HUmor detection, target group identification, and prediction of the degree of prejudice. Diferent text representation together with classical models, ensemble models, and the Spanish transformers BETO</p>
      </abstract>
      <kwd-group>
        <kwd>Classical</kwd>
        <kwd>Transformers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>H</p>
      <p>U</p>
      <p>
        UH
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and RoBERTa [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] were evaluated on the dataset provided by the competition called “HUrtful
HUmour (HUHU) Detection of humour spreading prejudice in Twitter” [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It was observed that: (i)
transformers architectures highly outperform classical and ensemble models when it comes to detecting
degree of prejudice and the target group, but have a serious problem with overfitting for the last one; (ii)
oversampling was a key solution when dealing with imbalanced classes in a small data set; (iii) including
extra features regarding the written style or the underlying intentions of the writer are of great utility
when it comes to natural language tasks.
      </p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Natural language processing (NLP) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is currently one of the biggest and most promising
ifelds regarding machine learning and deep learning. The complexity of language makes NLP a
complicated and intriguing task. Some of the challenges faced when dealing with NLP tasks
are that language has strict rules when it comes to structure, it has multiple significations
depending on the context, or that minimal variations in some words completely change the
meaning or comprehension of the message. In a nutshell, NLP includes a group of machine and
nEvelop-O
LGOBE
†These authors contributed equally.
https://www.linkedin.com/in/hugoalbert/ (H. A. Bonet); https://www.linkedin.com/in/ainamagranerrincon/
      </p>
      <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
deep learning techniques which deal with text input to perform diferent tasks like classification
or regression.</p>
      <p>In this work we mainly focus on applying NLP techniques together with the state-of-the-art
transformer models for three tasks: (i) hurtful humor detection, which consists of distinguishing
common harmful tweets from those which are masked with humor; (ii) target identification,
where we classified if the tweet intends to spread sexism, prejudices against the LGBTIQ
community, racism, or fatphobia; and (iii) degree of prejudice prediction, which consists of
estimating how hurtful the tweets are in a scale from one to five.</p>
      <p>It is necessary to highlight that several aspects of the process are going to be taken into
account: (i) diferent representations of the text are going to be compared, such as bag of words,
cleaning the text, or word embeddings; (ii) classical models –like SVM or Logistic Regression–
are going to be compared with ensemble models –such as RandomForests or Stacking– and
stateof-the-art transformers for Spanish –BETO or RoBERTa–; and (iii) the change of performance
of the models when including extra features –like irony or emotions– is going to be considered.</p>
      <p>The main research question is: Which combination of text treatment, extra features, and
machine or deep learning model is better suited for each task? The above question is going
to be answered based on the results of the dataset given by the “HUrtful HUmour (HUHU)
Detection of humour spreading prejudice in Twitter” competition.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Methodology</title>
      <p>The methodology followed in this work did not contain a data collection phase as the dataset
was provided by the competition organizers. It was composed of six main steps:</p>
      <sec id="sec-3-1">
        <title>2.1. Data processing</title>
        <p>The data processing step corresponds to treat the text in order to feed the models properly. This
section explains all the diferent techniques employed in this process, although later on we
explain which of them where applied to each model.</p>
        <p>The first preprocessing employed consisted of a deep cleaning of the text. First, we got rid of
all URLs, HTML tags, and punctuation symbols. After, all words were turned to lowercase, stop
words were removed, the text was lemmatized, and words were stemmed using Porter Stemmer.
Said type of cleaning significantly reduced the complexity of the text, getting rid of noise and
other aspects that may or may not be important for the tasks.</p>
        <p>
          With the deeply cleaned text, we created a representation of the tweets making use of Bag of
Words (BoW) in both ways, just counting the appearences of each word and by applying the
weighting scheme TF-IDF. However, we realized that, TF-IDF was not providing us with any
significant advantage. The deeply cleaned text was also used to create a representation based
on word embeddings, by using Fast Text [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] [9] [10] and also All-Mini-LM [11], a multilingual
transformer which is known for being fast.
        </p>
        <p>The second data processing method applied was really simple. The text was not treated and
just used to extract the embeddings through All-Mini-LM.</p>
        <p>The third and last method consisted of tokenizing the text with the corresponding Tokenizers
for RoBERTa and BETO. The tokenized text was the input of the transformers when fine-tuning
[12] them. However, we also extracted the embeddings of RoBERTa to change the top model
–the classifier– for other classical ones.</p>
        <p>The last applied processing consisted of making use of pre-trained transformers [13] to
extract extra features about the toxicity, hate speech (hate), context hate speech (context), irony,
emotions, and sentiments, explained in detail in Appendix A.</p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Exploratory analysis</title>
        <p>An exploratory analysis of both the tweets and the ground truths was conducted to picture a
clear image of how the next steps needed to be developed.</p>
        <p>When it comes to the first subtask, the detection of hurtful humor, the amount of not humorous
tweets was more than twice the number of humorous ones, as seen in Appendix B. Besides,
the analysis aimed also to show whether the written style should be used to tackle the task.
Some aspects that were taken into account were the number of dashes (-) written, as a common
structure of a joke includes a dialogue, the number of exclamation marks (!), or the number of
uppercase letters, the last two because they represent emphasis. We discovered that humorous
tweets had six times more dashes per tweet than not humorous ones. When counting the
number of exclamation marks, the plots did not show a diference in quantity for humorous
and not humorous tweets. However, when considering the number of exclamation marks per
tweet, the plots showed that they are three times more frequent in tweets using humor. Last
but not least, non-humorous tweets apparently used more uppercase than humorous ones, but
again by normalizing the values we realized that there was no significant diference between
both classes.</p>
        <p>Regarding the second subtask, where the target of the comment had to be identified among
the four groups mentioned in Section 1, we conducted a similar analysis. The vast majority of
the tweets where sexist, while there was only a minority of tweets targeting fat people. The
number of LGBTIQ and racist tweets where balanced. Whereas the number of uppercases per
tweet was almost the same for every category, the number of dashes per tweet and the number
of exclamation marks per tweet where clearly superior in tweets spreading fatphobia. Said
discoveries regarding the punctuation of the text, led us to the idea of considering representations
of the text that took that into account, like the embeddings without cleaning the texts.</p>
        <p>For both classification tasks, an XGBoost classifier was applied using the BoW representation,
as a first step to extract the subset of most important words in order to classify the samples. In
both subtasks, the subsets where almost the same. Those most important words (referred as
”most” in the report) were also used as extra features.</p>
        <p>The last subtask, the prediction of the degree of prejudice of the tweet, needed a diferent
approach of the exploratory analysis, as the variable was numerical. The distribution of values
followed a Gaussian Distribution, although it presented some negative asymmetry, being the
mean approximately 3.</p>
      </sec>
      <sec id="sec-3-3">
        <title>2.3. Model implementation and hyper-parameter analysis</title>
        <p>For all three subtasks, several models were implemented, as well as a hyper-parameter analysis.</p>
        <p>In the first subtask, which consisted of a binary classification, we implemented a series of
models, most of which were based on a single-language model for Spanish: RoBERTa-base
transformer. This was fine-tuned using its corresponding tokenized text. To this embeddings
we added diferent task related features in order to observe if these were of any help in the
given tasks. Also, a hyper-parameters analysis was conducted for all RoBERTa-base models,
in which we studied the charts reflecting the loss for each epoch relating the training set and
validation one. We finally considered the following as the best parameters: diferent learning
rates (0.00001 and 0.000001 depending on the subtask) , 10 epochs as well as 16 or 32 batches.
To all of these models a batch normalization was performed as well as a drop out of 0.3 after
the concatenation of the embeddings and the extra features. Regarding the classical models
implemented, a hyper-parameter analysis was conducted using a GridSearch strategy with
5-fold cross validation. Then, an ensemble method was used in which the estimators are the
models with the best parameters obtained. The metrics used to evaluate them, in both Subtasks
1 and 2a, were the same as the ones used on the HUHU shared task: F1-Macro. The main models
implemented were:
• RoBERTa-base transformer plus toxicity features.
• RoBERTa-base transformer with an addition of several task related features: irony, toxicity,
hate, emotions, and sentiment.
• RoBERTa-base transformer plus the most important words extracted from the BoW
analysis mentioned above and toxicity features.
• RandomForest with a number of estimators of 200 and a maximum depth of 30. Plus the
most important words extracted from the BoW analysis mentioned above and toxicity
features.
• Bagging Classifier with a Support Vector Classifier as the base estimator with the following
parameters: C = 10, gamma = 0.1, kernel = rbf, and 50 estimators. This was trained with
an embedding matrix obtained from the uncleaned data set through All-Mini-LM.</p>
        <p>For the multi-label classification subtask we changed the number of batches for the
RoBERTabase transformer models to 32. The main trained models were:
• RoBERTa-base transformer plus toxicity features.
• RoBERTa-base transformer plus the most important words extracted from the BoW
analysis mentioned above and context of hate speech features.
• RoBERTa-base transformer with an addition of several task related features: toxicity, hate,
emotions and Sentiment.
• RoBERTa-base transformer with an addition of several task related features: toxicity, hate,
irony, emotions, context and sentiments.
• Voting Classifier between two models: Random Forest with 50 estimators and a MLP
Classifier using the identity activation, alpha: 0.001, 500 as the maximum iterations, two
hidden layers with sizes 128 and 32, a constant learning rate, and a lbfgs solver. This was
put through a Multi-Output Classifier.</p>
        <p>Finally, for the regression subtask these were the most important trained models:
• RoBERTa-base transformer plus sentiments, emotions, hate, irony and toxicity features,
as well as the most important words. The used hyper-parameters were: 10 epochs, 32
batches and a learning rate of 0.000001. This transformer is the only one without batch
normalization.
• ExtraTreesRegressor with a number of estimators of 250 and a maximum depth of 30, to
which we added sentiments, emotions, hate, irony and toxicity features.</p>
        <p>For this last task, the metric used to evaluate the models was the RMSE.</p>
      </sec>
      <sec id="sec-3-4">
        <title>2.4. Models comparison</title>
        <p>The next step is to compare the results of the diferent models in order to rank them according
to their performance. Due to the reduced amount of samples in the training set, a three-fold
cross validation strategy was applied to calculate the F1-Macro score for each machine learning
model and ensemble model. In the case of the transformers, which need high amounts of data
to be trained, the strategy diverged to a mix of bagging and cross validation. We repeated
three times the measurement of the model, randomly splitting each time into train, validation,
and test, and randomly reordering the samples. The measure obtained was the average of the
F1-Macro score obtained in the test set for the three runs.</p>
        <p>The following tables show the results of the most relevant models according to their F1-Macro
score from the whole amount of models tested, as mentioned in the previous section. Table 1
illustrates the results for the models for Subtask 1, Table 2 illustrates the results for the models
for Subtask 2a, and Table 3 illustrates the results for the models for Subtask 2b.</p>
        <p>F1-Macro
RoBERTa + toxicity
RoBERTa + context + most
RoBERTa + toxicity + hate + emotions + sentiment
RoBERTa + toxicity + hate + emotions + context + sentiment
Voting + RF + MLP</p>
      </sec>
      <sec id="sec-3-5">
        <title>2.5. Final models selection and submission</title>
        <p>Once we have studied the performance of all the possibilities explained in Section 2.5, the final
models selected to submit to the competition are shown in Table 4, so they could be tested on
unknown test set to prove their capability of generalization. The criterion was choosing the
models which presented a higher F1-Macro score in the case of the first two subtasks, and a
lower RMSE in the last one.
1st
2nd</p>
        <p>RoBERTa + toxicity
RoBERTa + toxicity + context + most</p>
        <p>RoBERTa + context + most</p>
        <p>RoBERTa + toxicity
RoBERTa + toxicity + emotions + hate
+ sentiments + irony
ExtraTreesRegressor</p>
      </sec>
      <sec id="sec-3-6">
        <title>2.6. Post-competition improvements</title>
        <p>We are aware that our findings could have been enriched by the inclusion of other approaches
or techniques. Once we had submitted our models, having received the labeled test set, we
applied some of those ideas which we had came up with but that we could not include in the
submissions, and we plan to continue thoroughly examining other methodologies, as we explain
later in Section 5.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Results</title>
      <p>In this section, we present the final results obtained in the HUHU competition and compare
them with the results obtained during our experiments, shown in Table 5.</p>
      <p>As we can see, the models submitted for the classification tasks show a significant diference
in performance compared to our previous results.</p>
      <p>The results show that the solutions we had implemented to address a possible overfitting
problem have not worked as we would have wanted. We used batch normalization for the
Subtasks 1 and 2a plus a dropout technique for all subtasks, even though this last approach has
successfully worked for Subtask 2b, it seems that the batch normalization has not influenced
positively to avoid overfitting. Moreover, it has fuelled this problem.</p>
      <p>Regarding the last subtask, where a regression model was presented, we have achieved
outstanding results in the first submission, using a RoBERTa model and adding as extra features:
toxicity, emotions, hate, sentiments, and irony. We have obtained a slight diference between
the test RMSE value and the training one, accomplishing a fourth position in the HUHU
ranking results. The ExtraTreesClassifier difered more when it comes to the performance in
the submission, although its performance was not too far from the first submission. In spite
of such small diference, said model is 28 positions below the transformer, representing the
competitiveness of the models sent to the competition.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusions</title>
      <p>In this research, we have explored diferent methodologies to detect and identify harmful
comments in social media posts, particularly on platforms like Twitter.</p>
      <p>The findings of the study have yielded valuable insights and practical implications for selecting
the optimal models. Additionally, the research outcomes have enhanced our comprehension of
the topic, empowering us to make informed decisions, which have subsequently guided our
development of additional models. These will be explained in Section 5.</p>
      <p>Regarding the aspects we discovered that need to be taken into account in tasks related to
HUHU, the following list remarks the most important ones:
• Implementing measures to mitigate the issue of imbalanced classes play an important
role.
• State-of-the-art transformers usually outperform classical and ensembled models,
although addressing the problem of overfitting is a serious issue when dealing with them.
We observed a significant disparity between the results we obtained during our model
evaluation and the actual results provided by the organizers, which suggests that more
robust measures should be applied. As transformers are a complex type of deep learning
models, a deeper investigation must be carried on before starting using them.
• Ensuring thorough data preparation and a comprehensive understanding of the variables
that may be related to the study and the semantic meaning of the text, the optimal
performance of the models.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Future work</title>
      <p>Continuous improvement is a crucial aspect when working with natural language and artificial
intelligence. Therefore, these working notes do not have an end and will be always open to
new improvements. This section has the aim to reflect the additional aspects investigated after
the submission and their influence in the result obtained once we were given the labeled test
set. Moreover, this section also aims to state the ideas we came up with for future ideas not yet
implemented.</p>
      <p>When it comes to aspects already added to the project, we created extra features with the
frequency of dashes and exclamation marks for each tweet, instead of just including them in
the tokenized input for the models. This change leaded to a better performance of almost all
the models. For example, the Random Forest went from an F1-Macro score of 0.6 to 0.69 by
including those variables.</p>
      <p>However, the greatest improvement was caused by tackling the problem of unbalanced
classes. The first approach consisted of assigning weights to the categories while training the
transformers, which did not solve the problem. The second try was based on oversampling [14]
–as undersamplig, which was the idea proposed by other teams, did not seem appropriate with
such an small database–. We proposed two options:
• Duplicating the rows corresponding with the minority class: This method was used both
in Subtasks 1 and 2a, improving the performance from 0.399 and 0.475 to 0.413 and 0.492
respectively. By discarding the batch normalization in Subtask 2a, we reached 0.725.
This method was also applied to a Random Forest Classifier for Subtask 1, obtaining an
F1-Macro score of 0.826 in the test set, superior to the 0.820 obtained by the winner of
the HUHU competition.
• Using SVMSMOTE method for oversampling [15]: The second method was just applied
to the first subtask. It provides the new set with more variability, which reflects in the
models as an improve in performance. The Random Forests Classifier obtained in this
case a value of 0.831 in the test set.</p>
      <p>For the future, we propose diferent changes in order to seek for the best model. The first
idea is to change the way of training the transformers by adding automatic functions included
in Python libraries to ensure the correct learning of the neural network. By adding this, we
can focus our eforts on changing the architecture in order to avoid overfitting, as well as
improving the model by adding decay to the learning rate or more complex ways to combine
the embeddings of the diferent words such as convolutional layers.</p>
      <p>The second proposal is to adapt the SVMSMOTE oversampling strategy to the multi-label
task, in order to avoid the problem of imbalanced classes. For this task, we also propose to solve
the problem with chain classifiers, which would allow the models to extract relations between
the predicted categories.</p>
    </sec>
    <sec id="sec-7">
      <title>A. Extra features</title>
      <p>The aim of this appendix is to ofer an explanation of the extra features obtained with pre-trained
models:
• Toxicity: Classifies the sentence according to diferent levels and ways of expressing toxic
statements. The features returned are the following:
• Hate Speech: Describes how hateful a sentence is, according to the following features:
• Context Hate Speech: Focuses on the target of the hateful statement. As our data set only
showed hateful tweets, it seemed perfect for these extra features. The features where:
– Toxicity
– Severe toxicity
– Obscene
– Identity attack
– Insult
– Threat
– Sexual explicit
– Hateful
– Targeted
– Aggressive
– CALLS
– WOMEN
– LGBTI
– RACISM
– CLASS
– POLITICS
– DISABLED
– APPEARANCE
– CRIMINAL
• Irony: Tells whether a tweet is expressing irony or not.
• Emotions: States if the sentence is expressing diferent emotions, which are the following:
– Joy
– Sadness
– Anger
– Surprise
– Disgust
– Fear
– Others
• Sentiments: Shows the degree of positivity, negativity, or neutrality of the text.</p>
    </sec>
    <sec id="sec-8">
      <title>B. Exploratory analysis. Graphics</title>
      <p>This appendix aims to show the plots created during the exploratory analysis.</p>
    </sec>
    <sec id="sec-9">
      <title>C. Extra models</title>
      <p>* indicates the data has been processed before the text representation, if the technique chosen
is embedding representation.
ing text classification models, 2016. a r X i v : 1 6 1 2 . 0 3 6 5 1 .
[9] A. Joulin, E. Grave, P. Bojanowski, T. Mikolov, Bag of tricks for eficient text classification,
2016. a r X i v : 1 6 0 7 . 0 1 7 5 9 .
[10] P. Bojanowski, E. Grave, A. Joulin, T. Mikolov, Enriching word vectors with subword
information, 2017. a r X i v : 1 6 0 7 . 0 4 6 0 6 .
[11] N. Reimers, I. Gurevych, Making monolingual sentence embeddings multilingual using
knowledge distillation, in: Proceedings of the 2020 Conference on Empirical Methods
in Natural Language Processing, Association for Computational Linguistics, 2020. URL:
https://arxiv.org/abs/2004.09813.
[12] I. Goyal, P. Bhandia, S. Dulam, Finetuning for sarcasm detection with a pruned dataset,
2022. a r X i v : 2 2 1 2 . 1 2 2 1 3 .
[13] J. M. Pérez, J. C. Giudici, F. Luque, pysentimiento: A python toolkit for sentiment analysis
and socialnlp tasks, 2021. a r X i v : 2 1 0 6 . 0 9 4 6 2 .
[14] A. T. Handoyo, H. rahman, C. J. Setiadi, D. Suhartono, Sarcasm detection in twitter
performance impact while using data augmentation: Word embeddings, INTERNATIONAL
JOURNAL of FUZZY LOGIC and INTELLIGENT SYSTEMS 22 (2022) 401–413. URL: https:
//doi.org/10.5391%2Fijfis.2022.22.4.401. doi:1 0 . 5 3 9 1 / i j f i s . 2 0 2 2 . 2 2 . 4 . 4 0 1 .
[15] Q. Wang, Z. Luo, J. Huang, Y. Feng, Z. Liu, A novel ensemble method for imbalanced
data learning: Bagging of extrapolation-smote svm, Computational Intelligence and
Neuroscience 2017 (2017) Article ID 1827016. URL: https://doi.org/10.1155/2017/1827016.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>L. I. Merlo</surname>
          </string-name>
          ,
          <article-title>When humour Hurts: A Computational Linguistic Approach, Bachelor's thesis</article-title>
          , Universitat Politècnica de València,
          <year>2022</year>
          . URL: http://hdl.handle.net/10251/188166.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Donoso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Bravo-Marquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Carvallo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Araujo</surname>
          </string-name>
          , Albeto and distilbeto:
          <source>Lightweight spanish language models</source>
          ,
          <year>2023</year>
          .
          <article-title>a r X i v : 2 2 0 4 . 0 9 1 4 5</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          ,
          <year>2019</year>
          .
          <article-title>a r X i v : 1 8 1 0 . 0 4 8 0 5</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          ,
          <year>2019</year>
          .
          <article-title>a r X i v : 1 9 0 7 . 1 1 6 9 2</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Furman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Alemany</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Luque</surname>
          </string-name>
          ,
          <article-title>Robertuito: a pre-trained language model for social media text in spanish</article-title>
          ,
          <year>2022</year>
          .
          <article-title>a r X i v : 2 1 1 1 . 0 9 4 5 3</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Labadie-Tamayo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <article-title>Everybody hurts, sometimes. overview of hurtful humour at iberlef 2023: Detection of humour spreading prejudice in twitter</article-title>
          ,
          <source>in: Procesamiento del Lenguaje Natural (SEPLN)</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Liddy</surname>
          </string-name>
          ,
          <article-title>Natural language processing</article-title>
          ,
          <source>in: Encyclopedia of Library and Information Science</source>
          , 2nd ed.,
          <string-name>
            <surname>Marcel</surname>
            <given-names>Decker</given-names>
          </string-name>
          , Inc., New York,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , E. Grave,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Douze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jégou</surname>
          </string-name>
          , T. Mikolov, Fasttext.zip: Compress-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>