<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>R. Pan);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>UMUTeam at MentalRiskES@IberLEF 2024: Using the Fine-Tuning Approach of Transformer-Based Models with Sentiment Feature for Early Detection of Mental Disorders</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ronghao Pan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>José Antonio García-Díaz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rafael Valencia-García</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Facultad de Informática, Universidad de Murcia, Campus de Espinardo</institution>
          ,
          <addr-line>30100</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The alarming rise in mental disorders has sparked interest in early detection through social networking. The relationship between excessive use of social media and mental health problems, especially among adolescents, has led to a growing interest in early detection of these problems through social media comments. The MentalRiskES task in IberLEF 2024 focuses on this early detection of risks of mental disorders through comments in Spanish. This paper presents UMUTeam's contribution, focusing on disease and context detection. We use pre-trained linguistic models with outputs from emotion and sentiment models. In Task 1, we ranked 15th in decision and latency based classification; in Task 2, we ranked 10th in decision and latency based classification.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Mental disorders</kwd>
        <kwd>Deep learning</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Fine-tuning</kwd>
        <kwd>Transformers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The increase in mental illness in recent years is an alarming phenomenon that has captured the attention
of public health oficials, experts, researchers, and governments around the world. According to a recent
report by the World Health Organization (WHO), one in eight people in the world sufers from a mental
illness. There is no single cause for this increase, but rather a complex interplay of environmental, social
and biological factors. For example, in the COVID-19 era, the prevalence of anxiety and depression
increased by more than 26% in just one year. Suicide has become the fourth leading cause of death
among young people aged 15-29. This situation underscores the urgency of addressing the factors
contributing to the increase in these diseases and implementing efective strategies to improve the
mental and physical health of the global population [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Much evidence shows or suggests that there is a relationship between excessive use of social
networking sites by young people and their negative mental health outcomes, particularly an increase
in symptoms of depression and anxiety, as well as levels of stress. This relationship highlights the
importance of early identification of these symptoms in order to efectively intervene and prevent
these problems before they worsen. In other words, early identification of signs of deteriorating mental
health may enable parents, educators, and health professionals to take appropriate action to mitigate
the negative efects [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        For this reason, in recent years there has been a growing interest in detecting and identifying mental
disorders in social network, due to the increasing prevalence of mental health problems and their
relationship with the use of digital platforms. Various mental health related tasks have also emerged,
such as eRisk [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in the Cross-Lingual Evaluation Forum (CLEF) assessment campaign. However, these
campaigns have mainly focused on English, leaving aside other languages such as Spanish. Therefore,
the MentalRiskES [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] task in IberLEF 2024 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is the second edition that aims at the early detection of
risks of mental disorders through comments in Spanish from social network sources. In this edition,
the organizers have mainly proposed three tasks that focus on the identification of diferent mental
illnesses from diferent perspectives with a new corpus [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. These tasks include disease detection,
context detection, and suicidal ideation detection.
      </p>
      <p>This paper presents the UMUTeam’s contribution to the first two tasks, focusing on disease and
context detection, based on the fine-tuning of diferent pre-trained Transformer-based linguistic models
mixed with the outputs (logits) of the emotion and sentiment identification models with diferent early
detection methods. The rest of the paper is organized as follows. Section 2 presents the task and the
provided dataset. Section 3 describes the methodology of our proposed system to address subtask 1 and
subtask 2. Section 4 presents the obtained results. Finally, section 5 concludes the paper with some
conclusions and possible future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task description</title>
      <p>MetalRisk focuses on the early detection of mental illness through comments posted by diferent users
on Telegram. Thus, given a history of a user’s messages, the goal is to identify whether the user is
sufering from a disorder and the context that influences the mental health problem. The organizers
have considered the mental health detection from three perspectives: i) disorder detection, which is
a multi-class classification problem whose goal is to detect whether the user sufers from depression,
anxiety, or no detected disorder; ii) context detection, which is similar to the previous approach, but
with the addition of identifying the context from which the mental health problem appears to originate,
if a disorder is detected, and iii) suicide ideation detection, which is a binary detection problem whose
goal is to determine whether the user is exhibiting symptoms of potential suicidal ideation. Therefore,
one task has been proposed for each approach.</p>
      <p>Task 1, which is a multi-class classification problem, has 3 possible labels: depression, which is
characterized by persistent sadness, low mood, and lack of interest or pleasure in previously rewarding
and enjoyable activities; anxiety; and none (no disorder detected). In contrast, Task 2 has the same
objective as Task 1, but in this case, a multi-class classification problem is added, which consists of
identifying the context from which the detected mental health problem appears to originate. In this
case, the available contexts are: addiction context as “addiction”, emergency context as “emergency”,
family context as “family”, work context as “work”, social context as “social” and other contexts as
“other”. If no context is detected, “none” is assigned. Contexts are necessary only required if the subject
is predicted to have depression or anxiety.</p>
      <p>Table 1 shows the distribution of the dataset at the user level and at the message level after
preprocessing. At the user level, we can see that there are a total of 465 message histories from diferent
users, of which 213 have no mental illness, 164 have depression, and 99 have anxiety. To build a model
for identifying mental illness, we preprocessed each user’s history and retained only the negative
comments from users sufering from any mental illness. For those labeled “none”, we removed the
negative comments, leaving only the positive and neutral comments. In this way, we achieved noise
reduction, clarity in mental illness patterns, and simplicity in classification. However, by eliminating
negative comments from users labeled as “none”, we run the risk that the model will not learn to
properly distinguish between negative comments that are normal and those that are indicative of a
mental disorder. In Table 1, we can see the distribution of the message-level dataset after preprocessing.
There are about 9504 messages in total, of which 5931 are of the type “none”, 2322 are of the type
“depression”, and 1251 are of the type “anxiety”. The Table 2 shows the distribution of the data sets for
Task 2. In this case, the problem is of the multi-label type, which means that each user sufering from a
disease can be associated with more than one context.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        • Dataset: The process starts with a filtered dataset mentioned in Section 2.
• Preprocessing: Textual data is preprocessed to clean and prepare the text for analysis. In this
case, all emoticons, hashtags, links and special characters are removed.
• Split: Once the dataset is preprocessed, it is split into two subsets: one for training and one for
validation. This allows the efectiveness of the model to be evaluated on data not seen during
training.
• Pretrained language models: Pre-trained language models (e.g. BERT, RoBERTa, etc.) that
have already been trained on large amounts of text are used. These models generate vector
representations (embeddings) of the input text.
• Last Hidden State and Logits: The last hidden state of the pre-trained language model is
extracted, providing a deep representation of the text. Logits are the output of the Pysentimiento
model [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which are used to classify the text into diferent emotional categories, such as negative,
neutral, or positive.
• Sum of last hidden state and logits: The last hidden state representations and the logits are
combined to consolidate the information extracted from the text.
• Classification Head : Finally, a classification head is added, which is a combination of several
layers such as LayerNorm, Dropout and Linear, and Tanh as the activation function.
      </p>
      <p>Finally, the model is trained and the final output of the model is a prediction about the mental state
of the text, such as depression, anxiety, or none. This architecture makes it possible to detect early
signs of mental illness through text analysis, which can be crucial for early intervention and support of
afected individuals.</p>
      <p>For Task 2, which aims to identify the context of mental illness, we used the same approach as in
Task 1, as shown in Figure 2. However, in this case, the logits obtained from Model 1 are also added
in order to improve the performance of the model. Note that since this is a multi-label classification
problem, each user may have more than one context. Therefore, in the preprocessing, we converted the
labels into a one-hot representation.</p>
      <p>
        The models evaluated for both tasks are: MarIA, BETO and BERTIN. All three models are of the
monolingual type, i.e. they are pre-trained with a large dataset in Spanish. MarIA [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] is a
transformerbased language model for Spanish. It is based on the RoBERTa base model and has been pre-trained
on the largest Spanish corpus known to date, with a total of 570 GB of clean and deduplicated text.
BETO [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is a model based on BERT and pre-trained on a Spanish corpus. It is similar in size to a BERT
base and has been trained using the Whole Word Masking technique. BERTIN [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is a RoBERTa-based
model and was trained from scratch on the Spanish part of mC4 using Flax 1. The hyperparameters
used to the train the model for both tasks are: 16 training batch size, 15 epochs, 0.01 weight decay and
2e-5 learning rate.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>This section describes the systems submitted by our team in each run and shows the results obtained in
each task.
4.1. Task 1
For Task 1, diferent pre-trained models such as BETO, MarIA and BERTIN were evaluated from two
perspectives: normal fine-tuning for mental illness classification and fine-tuning of the pre-trained
models by adding sentiment features. In Table 3, we can observe the results where MarIA is the
most robust model in both configurations, with and without sentiment features, showing significant
improvements in accuracy and F1-score when sentiment features are added. MarIA obtained an
MF1 of 76.19 in the configuration without sentiment features and 81.12 with sentiment features. In
contrast, BETO and BERTIN also improve with sentiment features, but to a lesser extent than MarIA. In
particular, BERTIN shows a remarkable performance recovery when sentiment features are considered.
Therefore, based on the results obtained, we can conclude that the inclusion of sentiment can improve
the performance of models in the classification of mental disorders.</p>
      <p>Figure 3 shows the confusion matrix of the MarIA in the validation set. The confusion matrix shows
the percentages of correct and incorrect classifications for three categories: anxiety, depression, and
none. This analysis is important to evaluate the performance of a classification model. Overall, our
model shows high accuracy in the none category, with 98.15% of the instances correctly classified.
This indicates that the model is very efective in correctly identifying cases where neither anxiety
nor depression is present. For the depression category, the model also shows good accuracy, with</p>
      <p>MarIA
BETO
BERTIN
76.77% of instances correctly identified. However, there is significant confusion between depression
and anxiety, with 9.46% of depression cases incorrectly classified as anxiety and 13.76% classified as
none. The category anxiety has the lowest performance, with 54.00% of the instances correctly classified.
This suggests that the model has significant dificulty in accurately identifying anxiety. A significant
proportion of anxiety cases (36.40%) are misclassified as depression.</p>
      <p>In summary, the model performs well in identifying cases where there is neither anxiety nor depression,
and performs well in identifying depression. However, it shows significant dificulties in distinguishing
between anxiety and depression, as well as in correctly identifying cases of anxiety. This problem is
partly due to the smaller number of anxiety examples in the training set. The lack of anxiety examples in
the training data may lead to a bias in the model, making it less competent in recognizing this condition
compared to the other categories.</p>
      <p>However, the models obtained operate at the sentence level, so predicting whether a user has a mental
disorder requires taking into account predictions from the user’s previous comments. A user’s single
comment about a mental disorder does not necessarily mean that they have that disorder. Therefore,
to determine whether a user is sufering from depression or anxiety, our system processes the set of
user messages and uses the most common label to make the decision. For an early detection approach,
we tested a set of thresholds. For example, if the threshold is set to 5, and the first 5 user comments
are related to depression or anxiety, the system will conclude that the user is likely sufering from the
disorder.</p>
      <p>For this task, we submitted three runs, each with the same structure but difering in minor aspects of
the configuration of the early detection method.</p>
      <p>• Run 0. This run uses the MarIA model with sentiment features as the classification model and
uses a threshold of 5 for early detection. This means that during each round the labels of the
previous rounds are checked. If the user has 5 or more comments related to depression or anxiety,
the system considers it as depression or anxiety.
• Run 1. This run uses the same approach as for run 1, in this case, a threshold of 10.
• Run 2. This run takes the longest to decide because it consists of predicting all rounds and
making a final decision based on the most repeated label.</p>
      <p>Table 4 shows the results obtained in the oficial ranking based on decisions. We can see that the
conservative strategy of run 2 has obtained the best result with an M-F1 of 67.5, reaching the 15th place.
In this method, the system makes the decision after predicting all the rounds. On the other hand, the
threshold of 5 (run 0) for the early detection strategy has obtained an M-F1 of 64, reaching the 16th
position. On the other hand, the threshold of 10 has obtained the worst result with an M-F1 of 26.9.
Moreover, We can see that in this case, Run 2 is still the one that has obtained the best result in ERDE30
with a value of 0.166, followed by Run 0 and Run 1 with values of 0.194 and 0.501 respectively.
4.2. Task 2
For Task 2, which aims to identify the context of users sufering from mental illness, we followed the
same methodology as in Task 1. We analyzed diferent pre-trained Spanish language models, such as
BETO, MarIA and BERTIN, for multi-label context classification. We fine-tuned these models with and
without sentiment features and the logits of the mental illness classification model (MarIA model of
Task 1). Table 5 shows the results obtained.</p>
      <p>We can see that BETO stands out as the best model without the additional features, obtaining the
highest mean F1 score (28.2869), suggesting that it handles the multi-label context classification better
compared to MarIA and BERTIN in this configuration. The inclusion of sentiment and mental disorder
features does not always improve performance. In the case of MarIA, although recall improves slightly,
accuracy decreases significantly, suggesting a trade-of between these metrics. BERTIN shows an
improvement in the mean F1 score when additional features are added, with an M-F1 of 26.8134 vs.
25.0449, indicating that it can benefit from this additional information, although not as much as BETO
without these features.</p>
      <p>For this task, we present three runs based on those of Task 1. That is, when the system detects that a
user has a mental illness, it searches for the most repeated contexts from the previous rounds in order
to have a set of contexts related to the illness.</p>
      <p>Table 6 shows the BETO ranking report on the Task 2 validation partition. In this case, our model
shows uneven performance in diferent categories. The model performs well in the Social category,
but shows significant dificulty in Addiction and Emergency, with particularly low recall. The micro
and macro averages indicate limited overall performance, reflecting the need for improvement in
correctly classifying diferent classes. Overall, the model needs adjustments, such as more balanced data
collection, hyperparameter tuning, and improved preprocessing techniques, such as pre-identification
of mental illness, to improve its performance in identifying all categories.</p>
      <p>Table 7 shows the results obtained in the oficial ranking based on decisions. We can see that Run 2
has the highest accuracy with a value of 0.077 and the best macro precision (M-P) of 0.224, showing a
balance between accuracy and recall, although with low absolute values in all metrics. However, it is
followed by Run 0, with the best M-F1 of 22.4 and Run 1 with M-F1 of 4.4, respectively. On the other
hand, Run 2 is clearly the top performer in this latency-based evaluation. It has the lowest ERDE5
(0.203) and ERDE30 (0.166) values, indicating better performance in terms of errors relative to early
detection.</p>
      <p>In the decision-based evaluation, Run 2 shows a better balance between precision and recall, although
with low absolute values. In the latency-based evaluation, Run 2 excels in all key metrics, showing to
be the best model in terms of early detection eficiency and accuracy.</p>
      <p>From the results obtained, we can see that the simple number of comments may not be enough; the
context and severity of the comments are also important. In this case, a threshold of 5 is better than
10, but the prediction is still more robust in all rounds as in Task 1. We have also found that removing
certain negative comments from users marked as “none” runs the risk of the model not learning to
properly distinguish between negative comments that are normal and those that are indicative of a
mental disorder.</p>
      <p>In terms of carbon emissions, our approach has a mean duration of 3.156 with a deviation of 2.647
and a mean emission of 5.87e-5 with a deviation of 5.11e-5 across all runs in Task 1 and Task 2, because
submissions for both tasks are uploaded at the same time. In this case, our mean duration is at the
lowest of the rankings, but our mean emission is in the middle of the rankings.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This article summarizes UMUTeam’s participation in the MentalRiskES shared task at IberLEF 2024.
This task focuses on the early detection of mental illness from three perspectives: disease detection,
context detection, and suicidal ideation detection.</p>
      <p>In this shared task, we participated in Task 1 and Task 2, which focus on disease detection and
context detection, respectively. In both tasks, we have evaluated the normal tuning approach of diferent
pre-trained language models and also the tuning of these models with sentiment features obtained with
pysentimiento for Task 1. In addition, we evaluated the concatenation of pre-trained language models
with sentiment features and the classification model output of Task 1.</p>
      <p>In Task 1, we achieved the 15th position with an M-F1 of 0.675 in run 2 with an ERDE30 value of 0.162.
On the other hand, we obtained better results in Task 2, reaching the 10th position in the decision-based
ranking with an M-F1 of 0.203 in Run 2 and an ERDE30 value of 0.166.</p>
      <p>The results indicate that the simple number of comments may not be suficient; it is also important
to consider the context and severity of the comments. In this case, a threshold of 5 is preferable to 10,
although the prediction remains more robust in all rounds. We have also found that removing certain
negative comments from users marked as “None” may prevent the model from learning to properly
distinguish between normal negative comments and those that indicate a mental disorder.</p>
      <p>
        As a future line, we propose to incorporate context prior to current comment prediction and to
test diferent early detection methods. We also propose to evaluate diferent multilingual pre-trained
models based on Transformers. In addition, we think that it is relevant to consider the relationship
between signs of depression and hate speech [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], the use of humour [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and the demographic
and psychographic traits of the authors of the messages [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work is part of the research projects LaTe4PoliticES (PID2022-138099OB-I00) funded by
MICIU/AEI/10.13039/501100011033 and the European Regional Development Fund (ERDF)-a way of making
Europe and LT-SWM (TED2021-131167B-I00) funded by MICIU/AEI/10.13039/ 501100011033 and by the
European Union NextGenerationEU/PRTR, and ”Services based on language technologies for
political microtargeting“ (22252/PDC/23) funded by the Autonomous Community of the Region of Murcia
through the Regional Support Program for the Transfer and Valorization of Knowledge and Scientific
Entrepreneurship of the Seneca Foundation, Science and Technology Agency of the Region of Murcia.
Mr. Ronghao Pan is supported by the Programa Investigo grant, funded by the Region of Murcia, the
Spanish Ministry of Labour and Social Economy and the European Union - NextGenerationEU under
the “Plan de Recuperación, Transformación y Resiliencia (PRTR)”.
based on political ideology: An author analysis study on spanish politicians’ tweets posted in 2020,
Future Generation Computer Systems 130 (2022) 59–74.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sacco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Camilleri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Eberhardt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Umla-Runge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Newbury-Birch</surname>
          </string-name>
          ,
          <article-title>A systematic review and meta-analysis on the prevalence of mental disorders among children and adolescents in europe</article-title>
          ,
          <source>European Child &amp; Adolescent Psychiatry</source>
          (
          <year>2022</year>
          ).
          <source>doi:10.1007/s00787-022-02131-2.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Shannon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bush</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Villeneuve</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Hellemans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guimond</surname>
          </string-name>
          ,
          <article-title>Problematic social media use in adolescents and young adults: Systematic review and meta-analysis</article-title>
          ,
          <source>JMIR Ment Health</source>
          <volume>9</volume>
          (
          <year>2022</year>
          )
          <article-title>e33450</article-title>
          . URL: https://mental.jmir.org/
          <year>2022</year>
          /4/e33450. doi:
          <volume>10</volume>
          .2196/33450.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Martín-Rodilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , Overview of erisk 2023:
          <article-title>Early risk prediction on the internet</article-title>
          ,
          <source>in: International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer,
          <year>2023</year>
          , pp.
          <fpage>294</fpage>
          -
          <lpage>315</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Mármol-Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Moreno-Muñoz</surname>
          </string-name>
          ,
          <string-name>
            <surname>F. M.</surname>
          </string-name>
          <article-title>Plaza-del-</article-title>
          <string-name>
            <surname>Arco</surname>
            ,
            <given-names>M. D.</given-names>
          </string-name>
          <string-name>
            <surname>Molina-González</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          <string-name>
            <surname>Martín-Valdivia</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureña-López</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Montejo-Ráez</surname>
          </string-name>
          , Overview of mentalriskes at iberlef 2024:
          <article-title>Early detection of mental disorders risk in spanish</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>73</volume>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chiruzzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          , Overview of IberLEF 2024:
          <article-title>Natural Language Processing Challenges for Spanish and other Iberian Languages, in: Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2024), co-located with the 40th Conference of the Spanish Society for Natural Language Processing (SEPLN 2024), CEUR-WS</article-title>
          .org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Mármol Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Moreno</given-names>
            <surname>Muñoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Plaza-del Arco</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. D. Molina González</surname>
            ,
            <given-names>M. T. Martín</given-names>
          </string-name>
          <string-name>
            <surname>Valdivia</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureña-López</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Montejo</surname>
            <given-names>Ráez</given-names>
          </string-name>
          ,
          <article-title>MentalRiskES: A new corpus for early detection of mental disorders in Spanish</article-title>
          , in: N.
          <string-name>
            <surname>Calzolari</surname>
            , M.-
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lenci</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sakti</surname>
          </string-name>
          , N. Xue (Eds.),
          <source>Proceedings of the 2024 Joint International Conference on Computational Linguistics</source>
          ,
          <article-title>Language Resources and Evaluation (LREC-COLING 2024), ELRA</article-title>
          and
          <string-name>
            <given-names>ICCL</given-names>
            ,
            <surname>Torino</surname>
          </string-name>
          , Italia,
          <year>2024</year>
          , pp.
          <fpage>11204</fpage>
          -
          <lpage>11214</lpage>
          . URL: https://aclanthology.org/
          <year>2024</year>
          .lrec-main.
          <volume>978</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rajngewerc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Giudici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Furman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Luque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Alemany</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Martínez</surname>
          </string-name>
          ,
          <article-title>pysentimiento: a python toolkit for opinion mining and social nlp tasks</article-title>
          ,
          <source>arXiv preprint arXiv:2106.09462</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Fandiño</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Estapé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pàmies</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Palao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Ocampo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Carrino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. A.</given-names>
            <surname>Oller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Penagos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <article-title>Maria: Spanish language models</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          (
          <year>2022</year>
          ). URL: https://upcommons.upc.edu/handle/2117/367156#.YyMTB4X9A-0. mendeley. doi:
          <volume>10</volume>
          .26342/2022-68-3.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Cañete</surname>
          </string-name>
          , G. Chaperon,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-H.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pérez</surname>
          </string-name>
          ,
          <article-title>Spanish pre-trained bert model and evaluation data</article-title>
          ,
          <source>in: PML4DC at ICLR</source>
          <year>2020</year>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>J. D.</surname>
          </string-name>
          la Rosa y Eduardo G. Ponferrada y Manu Romero y Paulo Villegas y Pablo González de Prado Salas y María Grandury,
          <article-title>Bertin: Eficient pre-training of a spanish language model using perplexity sampling</article-title>
          ,
          <source>Procesamiento del Lenguaje Natural</source>
          <volume>68</volume>
          (
          <year>2022</year>
          )
          <fpage>13</fpage>
          -
          <lpage>23</lpage>
          . URL: http://journal.sepln.org/ sepln/ojs/ojs/index.php/pln/article/view/6403.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Jiménez-Zafra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>García-Cumbreras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Evaluating feature combination strategies for hate-speech detection in spanish using linguistic features and transformers</article-title>
          , Complex &amp; Intelligent
          <string-name>
            <surname>Systems</surname>
          </string-name>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Salmerón-Ríos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Fine grain emotion analysis in spanish using linguistic features and transformers</article-title>
          ,
          <source>PeerJ Computer Science</source>
          <volume>10</volume>
          (
          <year>2024</year>
          )
          <article-title>e1992</article-title>
          . doi:
          <volume>10</volume>
          .7717/peerj-cs.
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          ,
          <article-title>Compilation and evaluation of the spanish saticorpus 2021 for satire identification using linguistic features and transformers</article-title>
          ,
          <source>Complex &amp; Intelligent Systems</source>
          <volume>8</volume>
          (
          <year>2022</year>
          )
          <fpage>1723</fpage>
          -
          <lpage>1736</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>García-Díaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Colomo-Palacios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Valencia-García</surname>
          </string-name>
          , Psychographic traits identification
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>