<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>T)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Evaluation of Linguistic Features Separately or Combined with Transformers for Solving Automatic Text Classification Tasks in Spanish</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>José Antonio García-Díaz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Informática, Universidad de Murcia, Campus de Espinardo</institution>
          ,
          <addr-line>30100 Murcia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>1</volume>
      <fpage>9</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>In this paper we describe the evaluation stage and analysis of the UMUTextStats tool for extracting linguistic features applied to text classification in several domains, including the identification of sexist or ofensive comments, emotions, and a fine-grained analysis regarding what texts are funny and what mechanisms are involved to make them funny. These subtasks were organised by IberLEF and IberEval 2021 workshops. During the participation on these subtasks, the linguistic features were evaluated separately and combined with state-of-the-art transformers by means of ensembles and knowledge integration strategies, with the objective of achieve competitive results in all tasks. At the same time, we seek to improve our methods to obtain some interpretability of the results. In summary, our results suggest than the combination of diferent feature sets improves text classification tasks, especially when they are input in the same neural network.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Text classification</kwd>
        <kwd>Feature engineering</kwd>
        <kwd>Natural Language Processing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In the past edition of the doctoral symposium organised by the thematic network PLN.net,
we described the main objectives related to this doctoral thesis [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These objectives consist
in the development of a set of linguistic features for Spanish and their inclusion in Natural
Language Processing (NLP) tasks, such as forensics linguistic, author profiling, infodemiolgy
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], or misogyny identification [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] among others. We also described our participation in TASS
2020 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and MEX-A3T [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] shared tasks, and we described two NLPs tools developed, one for
compiling and annotation corpora [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and the other, inspired in LIWC [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], for extracting the
linguistic features [
        <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
        ].
      </p>
      <p>In summary, our main hypothesis is that linguistic features are somehow high-order features
than statistical features based on words and their relationships. Examples of these feature sets
are n-grams or contextual and non-contextual embeddings. Moreover, we state that applying
the linguistic features results in more reliable and interpretable models.</p>
      <p>
        Our previous participation in the symposium was very positive for us, as we received valuable
feedback from our mentors. Specifically, they recommend us to focus on the interpretability
of the machine-learning models. In addition, they note that the results achieved in the shared
tasks [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], in which we combined the linguistic features with non-contextual embeddings,
were limited compared to the results of other participants. Therefore, this year we have focus
on improving our pipeline by (1) evaluating other feature sets to combine and compare with
the linguistic features; (2) evaluating techniques for combining the features, such as knowledge
transfer, and ensemble learning; and (3) evaluating explainable deep-learning techniques. We
have measured our progress by participating in four shared tasks of IberLEF 2021 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and one
shared task from IberEval 2021. In addition, during this year, we have applied our methods for
conducting author profiling and hate-speech detection, with two publications that are under
review in scientific journals.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Validation</title>
      <p>After summarising the main hypotheses of this research, here we describe the validation process
carried out, that consisted in the participation in the following shared tasks: EXIST-2021 (see
Section 2.1), EmoEvalEs 2021 (see Section 2.2), Hahackathon 2021 and HaHa 2021 (see Section
2.3), and MeOfendEs 2021 (see Section 2.4). For each subtasks we include a summary of the
main objectives, the methods evaluated as well as the main insights extracted for each one.</p>
      <sec id="sec-2-1">
        <title>2.1. EXIST-2021. Sexist language identification</title>
        <p>
          The shared task EXIST-2021 [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] focuses on the identification and categorisation of sexist
language written in Spanish and English. The organisers of this shared task compiled and
annotated documents from several micro-blogging platforms. This shared task was divided into
two subtasks: (1) a binary classification of sexism utterances, and (2) a multi-class identification
of sexist traits, namely, ideological and inequality, stereotyping and dominance, objectification,
sexual violence, and misogyny and non-sexual violence.
        </p>
        <p>We participated in both subtasks with a combination of the linguistic features and
transformers. For this, we tackle each dataset independently and combine the results at the end.
During our research, we evaluated diferent types of embeddings, including (1) pre-trained non
contextual word embeddings from fastText, word2vec, and gloVe; (2) sentence non-contextual
embeddings from fastText; and (3) contextual word embeddings based on transformers. In
addition, we evaluate diferent neural network architectures, including multi-layer perceptrons,
convolutional neural networks, and bidirectional recurrent neural networks. We combined each
feature set with the functional API of Keras1, in a knowledge integration fashion, entering each
feature into separate hidden layers and combining them before predicting the result.</p>
        <p>Three runs were sent. One with the linguistic features, another combining the linguistic
features with transformers, and another based in an ensemble of neural networks of linguistic
features and contextual and non-contextual word and sentence embeddings. We achieved our
best result in task 1, with an accuracy of 75.14% using the ensemble learning approach, and an
accuracy of 61.70% for task 2, with the combination of the linguistic features and transformers.
1https://keras.io/guides/functional_api/ (Last accessed: 2021-07-17)
These results were not far from the best results achieved in the oficial leader board, achieved by
the team AI-UPV_1, with an accuracy of 78.04% in task 1, and an accuracy of 65.77% for task 2.</p>
        <p>
          We observed that our baseline, consisted in the linguistic features, achieved limited results.
This result was the expected with the English dataset, but unexpected with the Spanish dataset,
especially as the linguistic features provided promising results regarding misogyny identification
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The main diferences between both datasets (the Spanish MisoCorpus 2020 and the
EXIST2021) are the number annotators (3 for the MisoCorpus, 5 for EXIST-2021) and the fact that the
annotators from EXIST-2021 followed the guidelines from two experts in gender issues. Besides,
the Spanish MisoCorpus 2021 contains a large number of tweets from news sites that were
labelled as neutral.
        </p>
        <p>To gain some interpretability, we extracted the information gain from the linguistic features.
We observed that the linguistic features related to sexual issues and related to female social
groups were discriminatory features for the identification of sexism. However, both features
appeared less frequently in documents labelled as stereotyping and dominance.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. EmoEvalEs 2021. Emotion Detection</title>
        <p>
          The EmoEvalES shared task [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] is focused on extracting the emotions expressed by users on
social media, which it is challenging mainly due to the absence of prosodic features and facial
expressions. This shared task consists in a multi-class classification for determining if a text
contains one of the following classes: Anger, Disgust, Fear, Joy, Sadness, Surprise or Others.
        </p>
        <p>
          Like the other shared tasks in which we participated, we based our proposal in the combination
of the linguistic features and transformers [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. We achieved 6th position in the oficial leader
board with an accuracy of 68.5990%, falling only 4.1667% below the best result.
        </p>
        <p>
          In this shared task, we achieve a significant improvement in our pipeline, as we were able
to extract the contextual sentence embeddings from BETO [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. For this, we extracted a fixed
representation of 768-length vector from the [CLS] token, after fine-tuning the model with
the EmoEvalEs dataset [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. We observe than the performance of this approach was similar
to the one achieved using HuggingFace’s Trainer2. However, the fixed representation of the
BERT embeddings provided to us two important benefits: they are easier to combine with other
feature sets within the same neural network and the required time for training and performing
inference is reduced.
        </p>
        <p>
          As we expected, regarding the interpretability of our results, we observed a strong correlation
between lexicons related emotions with the labels. Lexicons containing sad expressions were
strong related to documents annotated as sadness and disgust. anger with the psycho-linguistic
process anger. Negative processes were also related to anger, disgust, fear, sadness, and surprise.
Regarding humour, we have participated in two tasks regarding its identicfiation,
categorisation, and evaluation. On the one hand, the HaHackathon 2021 shared task [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], proposed
in IberEval’2021, focused on texts written in English, and HaHa 2021 [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], focused on
Spanish. Both shared tasks were divided into four subtasks each one. HaHackathon focused on
2https://huggingface.co/transformers/main_classes/trainer.html(lastaccessed:2021-07-17)
determining if a text is funny or not (binary classification), how humorous it is (regression),
and if its humour is controversial or not and how much (binary classification and regression,
respectively). HaHa shared with HaHackathon the first two subtasks, but they included two
new subtasks for determining what are the mechanisms to make a text funny and what are the
targets of the joke, that were, respectively, a multi-classification and a multi-label tasks.
        </p>
        <p>In HaHackathon 2021 we achieved position 45, with a F1-score of 91.60% in the subtask 1a. A
RMSE of 0.8847 for subtask 1b, achieving position 47. Position 14 in subtask 1c, with a F1-score
of 57.22%. Finally, we achieved position 46 for subtask 2a, with a RMSE score of 0.8740. It is
worth mentioning that, as HaHackathon 2021 was focused in English, we only use the subset of
the linguistic features based on corpus statistics, such as the type/token ratio (TTR). In HaHa
2021, we achieved the 1st position in Funniness Score Prediction, the 8th position for humor
classification subtask, and the 7th and the 3rd position for the subtasks of humour mechanism
and target classification, respectively.</p>
        <p>
          For subtask 2 of HaHa 2021, in which we achieved the best result, we observed that stylometry
is a relevant linguistic category. We also observe that interjections, verbs in third person, adverbs,
augmentative sufixes, and proper nouns were also relevant features. It also caught our attention
to find features related to the number of orthographic errors, as they can be committed on
purpose as a humoristic device.
2.4. MeOfendES 2021
Finally, we participated in the MeOfendEs 2021 shared task [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], focused on the identification
and categorisation of ofensiveness, with datasets in European and Mexican Spanish extracted
from diferent social media platforms. This shared task was divided into two subtasks (two
subtasks per language variation). On the one hand, the European Spanish subtasks were based on
multi-classification, discerning among (1) ofensive texts whose target is a person; (2) ofensive
texts whose target to groups; (3) texts with inadequate language, but not necessary ofensive;
and (4) non ofensive texts. The Mexican Spanish, on the other hand, were binary classification
problems. Each linguistic variant included a subtask in which contextual features from the
documents could be considered.
        </p>
        <p>
          In this case, apart from the linguistic features and transformers, we evaluate fine-grained
negation features [
          <xref ref-type="bibr" rid="ref18 ref19 ref20">18, 19, 20, 21</xref>
          ] as a result of a collaboration with the Universidad de Jaén. All
these features were combined with ensemble learning. Specifically, we evaluated ensembles
based on the mode of the predictions, ensembles based on averaging the predictions of each
neural network, ensembles based on the highest probability, and ensembles based on training
regression machine learning model from the probabilities of the training split. We observed
that the ensembles based on linear regression provided the best results whereas the ones based
on the highest probability the best precision over the ofensive class.
        </p>
        <p>Our oficial results were promising, as we ranked in the 2nd place in subtask 1 (F1-score of
87.8289%), 1st in subtask 2 (F1-score 87.8289%), 5th in subtask 3 (F1-score of 67.0588%), and 1st
in subtask 4 (F1-score of 66.9449%). However, there were less participants in the subtasks that
included the contextual features. Regarding the interpretability of the models, we observed in
the Spanish dataset that negative psycho-linguistic processes were strong features to discern
from non-ofensive documents from the others, but that they were not good indicators to discern
among if the target is a person, a group or simply the use of inadequate language.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusions and further work</title>
      <p>Since I am in the last year of my doctorate, and having previously participated in the previous
version of this symposium, we have focus this study on the validation tasks. Specifically, we
described our participation in five shared tasks regarding text classification in which we have
achieved promising results. We have tried to follow the indications given by our mentors
and we feel that their advises have helped us in a great extent. There is, however, a still a
lot of room for improvement. For example, we are still focusing on the interpretability based
on the linguistic features in isolation, but not in the context of the neural network. To solve
this, we will evaluate the ensembles to analyse which features have the documents that are
successfully classified correctly by the transformers and not by the linguistic features and vice
versa. Moreover, we are adapting tools such as SHAP and LIME [22]. We are also focusing
on improving the detection of figurative language [ 23] to apply to specific domains such as
sarcasm, irony, and satire identification [24].</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This work was supported by the Spanish National Research Agency (AEI) through project
LaTe4PSP (PID2019-107652RB-I00/ AEI / 10.13039/501100011033). In addition, José Antonio
García-Díaz was supported by Banco Santander and the University of Murcia through the
Doctorado industrial programme.
Sfu review sp-neg: a spanish corpus annotated with negation for sentiment analysis. a
typology of negation patterns, Language Resources and Evaluation 52 (2018) 533–569.
[21] S. M. Jiménez-Zafra, R. Morante, E. Blanco, M. T. M. Valdivia, L. A. U. Lopez, Detecting
negation cues and scopes in spanish, in: Proceedings of The 12th Language Resources and
Evaluation Conference, 2020, pp. 6902–6911.
[22] Y. Rychener, X. Renard, D. Seddah, P. Frossard, M. Detyniecki, Sentence-based model
agnostic NLP interpretability, CoRR abs/2012.13189 (2020). URL: https://arxiv.org/abs/2012.
13189. arXiv:2012.13189.
[23] M. del Pilar Salas-Zárate, G. Alor-Hernández, J. L. Sánchez-Cervantes, M. A.
ParedesValverde, J. L. García-Alcaraz, R. Valencia-García, Review of english literature on figurative
language applied to social networks, Knowl. Inf. Syst. 62 (2020) 2105–2137. URL: https:
//doi.org/10.1007/s10115-019-01425-3. doi:10.1007/s10115-019-01425-3.
[24] M. del Pilar Salas-Zárate, M. A. Paredes-Valverde, M. Á. Rodríguez-García, R.
ValenciaGarcía, G. Alor-Hernández, Automatic detection of satire in twitter: A
psycholinguisticbased approach, Knowl. Based Syst. 128 (2017) 20–33. URL: https://doi.org/10.1016/j.knosys.
2017.04.009. doi:10.1016/j.knosys.2017.04.009.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Garcıa-Dıaz</surname>
          </string-name>
          ,
          <article-title>Using linguistic features for improving automatic text classification tasks in spanish 2802 (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cánovas-García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Ontology-driven aspect-based sentiment analysis classification: An infodemiological case study regarding infectious diseases in latin america</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>112</volume>
          (
          <year>2020</year>
          )
          <fpage>641</fpage>
          -
          <lpage>657</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cánovas-García</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Colomo-Palacios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Detecting misogyny in spanish tweets. an approach based on linguistics features and word embeddings</article-title>
          ,
          <source>Future Generation Computer Systems</source>
          <volume>114</volume>
          (
          <year>2020</year>
          )
          <fpage>506</fpage>
          -
          <lpage>518</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          , Á. Almela,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          , Umuteam at tass 2020:
          <article-title>Combining linguistic features and machine-learning models for sentiment classification</article-title>
          ,
          <source>in: Notebook Papers of 2nd SEPLN Workshop on Iberian Languages Evaluation Forum (IberLEF)</source>
          , Malaga, Spain,
          <year>2020</year>
          , pp.
          <fpage>187</fpage>
          -
          <lpage>196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          , Umuteam at mex-a3t'
          <year>2020</year>
          :
          <article-title>Detecting aggressiveness with linguistic features and word embeddings</article-title>
          ,
          <source>in: Notebook Papers of 2nd SEPLN Workshop on Iberian Languages Evaluation Forum (IberLEF)</source>
          , Malaga, Spain,
          <year>2020</year>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>292</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          , Á. Almela,
          <string-name>
            <given-names>G.</given-names>
            <surname>Alcaraz-Mármol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Umucorpusclassifier: Compilation and evaluation of linguistic corpus for natural language processing tasks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>65</volume>
          (
          <year>2020</year>
          )
          <fpage>139</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y. R.</given-names>
            <surname>Tausczik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Pennebaker</surname>
          </string-name>
          ,
          <article-title>The psychological meaning of words: Liwc and computerized text analysis methods</article-title>
          ,
          <source>Journal of language and social psychology 29</source>
          (
          <year>2010</year>
          )
          <fpage>24</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>García-Vega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Díaz-Galiano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>García-Cumbreras</surname>
            ,
            <given-names>F. M. P.</given-names>
          </string-name>
          <string-name>
            <surname>del Arco</surname>
            , A. MontejoRáez,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Zafra</surname>
            ,
            <given-names>E. M.</given-names>
          </string-name>
          <string-name>
            <surname>Cámara</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          <string-name>
            <surname>Aguilar</surname>
            , M. Antonio,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Cabezudo</surname>
          </string-name>
          , et al.,
          <source>Overview of tass</source>
          <year>2020</year>
          <article-title>: introducing emotion detection (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Aragón</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Jarquín-Vásquez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes-Y-Gómez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J.</given-names>
            <surname>Escalante</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. V.</given-names>
            <surname>Pineda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Gómez-Adorno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Posadas-Durán</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Bel-Enguix, Overview of mex-a3t at iberlef 2020: Fake news and aggressiveness analysis in mexican spanish</article-title>
          ., in: IberLEF@ SEPLN,
          <year>2020</year>
          , pp.
          <fpage>222</fpage>
          -
          <lpage>235</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Montes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , E. Aragón,
          <string-name>
            <given-names>R.</given-names>
            <surname>Agerri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <string-name>
            <surname>Álvarez-Carmona</surname>
            ,
            <given-names>E. Álvarez</given-names>
          </string-name>
          <string-name>
            <surname>Mellado</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Carrillo-de Albornoz</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Chiruzzo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Freitas</surname>
            ,
            <given-names>H. Gómez</given-names>
          </string-name>
          <string-name>
            <surname>Adorno</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Gutiérrez</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez Zafra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Lima</surname>
            ,
            <given-names>F. M.</given-names>
          </string-name>
          <string-name>
            <surname>Plaza-de Arco</surname>
          </string-name>
          , M. Taulé,
          <article-title>Proceedings of the iberian languages evaluation forum (iberlef 2021)</article-title>
          , in: CEUR workshop,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Rodríguez-Sánchez</surname>
          </string-name>
          , J. C. de Albornoz, L. Plaza,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Comet</surname>
          </string-name>
          , T. Donoso, Overview of exist 2021:
          <article-title>sexism identification in social networks</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>F. M.</surname>
          </string-name>
          <article-title>Plaza-del-</article-title>
          <string-name>
            <surname>Arco</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Zafra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Montejo-Ráez</surname>
            ,
            <given-names>M. D.</given-names>
          </string-name>
          <string-name>
            <surname>Molina-González</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureña-López</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          <string-name>
            <surname>Martín-Valdivia</surname>
          </string-name>
          ,
          <article-title>Overview of the EmoEvalEs task on emotion detection for Spanish at IberLEF 2021</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          , G. Chaperon,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          ,
          <source>PML4DC at ICLR</source>
          <year>2020</year>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Reimers</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurevych</surname>
          </string-name>
          ,
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          , arXiv preprint arXiv:
          <year>1908</year>
          .
          <volume>10084</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Meaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Wilson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chiruzzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lopez</surname>
          </string-name>
          , W. Magdy,
          <article-title>Semeval 2021 task 7, hahackathon, detecting and rating humor and ofense</article-title>
          ,
          <source>in: Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chiruzzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Góngora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rosá</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Meaney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          , Overview of HAHA at IberLEF 2021: Detecting, Rating and Analyzing Humor in Spanish,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>F. M.</surname>
          </string-name>
          <article-title>Plaza-del-</article-title>
          <string-name>
            <surname>Arco</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Casavantes</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Escalante</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          <string-name>
            <surname>Martín-Valdivia</surname>
          </string-name>
          , A. MontejoRáez, M.
          <article-title>Montes-y-</article-title>
          <string-name>
            <surname>Gómez</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Jarquín-Vásquez</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Villaseñor-Pineda</surname>
          </string-name>
          ,
          <article-title>Overview of the MeOfendEs task on ofensive text detection at IberLEF 2021</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>67</volume>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <article-title>Negation processing in spanish and its application to sentiment analysis</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>66</volume>
          (
          <year>2021</year>
          )
          <fpage>193</fpage>
          -
          <lpage>196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. P.</given-names>
            <surname>Cruz-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taboada</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. T.</surname>
          </string-name>
          Martín-Valdivia,
          <article-title>Negation detection for sentiment analysis: A case study in spanish</article-title>
          ,
          <source>Natural Language Engineering</source>
          <volume>27</volume>
          (
          <year>2021</year>
          )
          <fpage>225</fpage>
          -
          <lpage>248</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taulé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Martín-Valdivia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Urena-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Martí</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>