<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>COTECMAR-UTB at eRisk 2025: Semantic-Centroid Symptom Ranking and Early Depression Detection using Adaptive Decision Rule⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luis Mendoza</string-name>
          <email>luimendoza@utb.edu.co</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joan Suarez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edwin Puertas</string-name>
          <email>epuertas@utb.edu.co</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Martinez</string-name>
          <email>jcmartinezs@utb.edu.co</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jairo Serrano</string-name>
          <email>jserrano@utb.edu.co</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Corporación de Ciencia y Tecnología para el Desarrollo de la Industria Naval</institution>
          ,
          <addr-line>Maritima y Fluvial (COTECMAR), Bolívar, Cartagena D.T. y C.</addr-line>
          ,
          <country country="CO">Colombia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad Tecnológica de Bolívar (UTB)</institution>
          ,
          <addr-line>Bolívar, Cartagena D.T. y C.</addr-line>
          ,
          <country country="CO">Colombia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Depression remains a major global health concern, with millions of people afected in various demographics. However, the timely detection of depression symptoms remains a challenge due to biases and limitations in traditional diagnostic methods. Social networks have become valuable sources for identifying early signs of depression, as they provide real-time user interactions that reflect emotional states. This paper explores the eRisk 2025 challenge, focusing on two primary tasks for early detection of depression in online conversations. Task 1 involves ranking sentences according to their relevance to depression symptoms, while task 2 addresses the analysis of emotional progression in real-time conversations. We apply Natural Language Processing (NLP) models, including Transformer architectures such as BERT, to capture semantic nuances in text. Furthermore, our methodology incorporates a Classifier with partial information (CPI) and a Decision Moment Classifier (DMC) to track emotional shifts over time, providing a framework for detecting depression risks early in conversational contexts. We present a comprehensive evaluation of our approach, discussing its challenges, successes, and potential for future improvements in early detection systems.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Depression detection</kwd>
        <kwd>early intervention</kwd>
        <kwd>social media analysis</kwd>
        <kwd>Natural Language Processing (NLP)</kwd>
        <kwd>Transformer models</kwd>
        <kwd>BERT</kwd>
        <kwd>emotional progression</kwd>
        <kwd>Decision Moment Classifier (DMC)</kwd>
        <kwd>Classifier with Partial Information (CPI)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Depression remains one of the leading causes of disability worldwide, afecting millions of individuals
across various age groups and socio-economic backgrounds. Despite its high prevalence, a significant
proportion of individuals sufering from depressive symptoms do not receive adequate treatment, often
due to barriers such as social stigma, limited access to mental health services, and underdiagnosis
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Traditional diagnostic methods based on clinical assessments and self-reporting are often subject
to biases and resource constraints, making it increasingly dificult to detect depression at an early
stage. Consequently, the need for innovative approaches to mental health diagnostics has become more
pressing.
      </p>
      <p>
        In recent years, the rise of social media platforms, such as Reddit, Quora, Facebook, and so on, has
opened new avenues for detecting early signs of depression. These platforms, where individuals share
their thoughts and emotional states in real time, provide rich sources of data that can be analyzed to
identify psychological distress [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The application of Natural Language Processing (NLP) and Machine
Learning (ML) techniques to these data has led to the development of automated systems for detecting
depression and other mental health issues, potentially ofering timely intervention [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ].
      </p>
      <p>
        This paper focuses on the eRisk 2025 challenge, which aims to advance early risk detection systems
for mental health through NLP and AI techniques. The challenge consists of two main tasks that
address diferent aspects of depression detection in online conversations: Task 1, which involves ranking
sentences based on their relevance to specific depression symptoms as outlined in the Beck Depression
Inventory-II (BDI-II) questionnaire, and Task 2, which focuses on tracking the emotional progression
of individuals in real-time, identifying subtle shifts in emotional states that may indicate worsening
depression [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ].
      </p>
      <p>The complexity of detecting depression from online texts arises from the nuanced and varied nature
of depressive symptoms, which can manifest diferently between individuals. To address this challenge,
this work integrates state-of-the-art NLP models, including Transformer-based architectures such as
BERT and its variants, to better capture the contextual and semantic meaning of text.</p>
      <p>In addition, the task of identifying depression symptoms in sequential conversational data, as in Task
2, introduces additional complexity. Unlike static data, conversational posts often involve incremental
changes in emotional state, requiring dynamic models capable of understanding emotional progression
over time.</p>
      <p>The contributions of this paper include the following.</p>
      <p>• A detailed exploration of the eRisk 2025 challenge, highlighting the significance of early detection
of depression through social media analysis.
• A robust evaluation of our methodology for both Task 1 and Task 2, demonstrating the challenges
and successes of applying NLP techniques to noisy data in the real world.
• A decision-making component that tracks emotional shifts over time, providing a framework for
early and continuous risk detection in conversational contexts.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>Task 1 and Task 2 of the eRisk 2025 challenge have driven the development of numerous methodologies
for the detection of depression on social networks, using cutting-edge natural language processing (NLP)
and machine learning (ML) techniques. These tasks have significantly influenced the state-of-the-art in
assessing depression symptoms and early detection in conversational interactions.</p>
      <p>
        Task 1 focuses on ranking sentences from social media posts according to their relevance to the 21
symptoms of depression outlined in the Beck Depression Inventory-II (BDI-II) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Transformer-based
models, such as BERT and its variants, such as RoBERTa and Distilbert, have become the dominant
approach for this task. These models generate sentence embeddings that capture contextualized
information, which can be used to rank sentences based on their relevance to specific depression
symptoms, such as sadness, pessimism, and agitation. The ability of Transformer models to encode
semantic relationships between words has significantly improved performance in depression symptom
detection, making them a powerful tool in the task of ranking relevant sentences [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ].
      </p>
      <p>
        Another sort of strategy used in Task 1 was the usage of synthetic data generated by Large
Language Models (LLMs) like GPT-O3-mini and ChatGPT 4o. These models augment existing datasets by
generating diverse and semantically rich examples of depression-related sentences for each symptom
in the BDI-II. This helps to enhance the training data, increasing its diversity and enabling models to
learn more nuanced patterns of depression expression in online text [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Furthermore, this was also
implemented with another sort of approach based on using early maladaptive schema (EMS), which
was adopted to achieve an enrichment of the ranking dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Task 2 addresses the early detection of depression in conversational interactions, with the goal of
predicting the likelihood of depression from a user’s posts in social media threads. This task typically
involves both binary classification and regression approaches. Models that combine various types of
data, such as lexical features, phonetic embeddings, and syntactic patterns, have shown great promise.
SVM (Support Vector Machines) models have been used to detect depression by incorporating these
multimodal inputs, which capture emotional expression, self-referential language, and interpersonal
dynamics in user interactions. Similarly, LSTM (Long Short-Term Memory) networks [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which are
a type of recurrent neural network, have been applied to this task as well. These models learn from
sequential data, capturing temporal dependencies that are crucial for identifying patterns indicative of
depression in conversations.
      </p>
      <p>
        In the context of early risk detection (ERD), the problem is addressed using the CPI (Classification
with Partial Information) and DMC (Decision Making Component) framework [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The CPI model
classifies users as at risk for depression based on the partial information available in their posts. The
DMC component, on the other hand, is responsible for deciding when to raise an alarm, determining
the optimal moment for classification by weighing the risk of false negatives against the need for timely
intervention. These components work together to identify and alert users at risk as early as possible,
balancing precision and recall in a real-world scenario.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Task 1: Search for Symptoms of Depression</title>
      <p>
        In Task 1 of the eRisk 2025 challenge, the goal is to classify and rank sentences from a dataset of social
media posts based on their relevance to the 21 symptoms of depression in the Beck Depression
InventoryII (BDI-II) [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. The aim is to develop models to identify and rank sentences that best represent
each depression symptom. The rankings are evaluated using metrics such as Mean Average Precision
(MAP), Recall Precision (R-Prec), Precision at 10 (P@10), and Normalized Discounted Cumulative Gain
(NDCG), which assess how accurately the models prioritize relevant content.
      </p>
      <sec id="sec-3-1">
        <title>3.1. Pipeline Explanation</title>
        <p>The initial challenge lay in designing an appropriate structure for integrating the available data. This
involved harmonising the .trec formatted files with their corresponding classification files, which
were annotated under two distinct evaluation schemes: unanimity and majority vote.</p>
        <p>Considering the data structure, a cleaning and filtering process was done on both datasets
corresponding to the 2023 and 2024 sets. This stage was imperative to ensure consistency in data quality,
format, and labelling prior to model development.</p>
        <p>After organising and cleaning the dataset, we proceeded to train various models using diferent
learning approaches. First, we employed classic machine learning techniques with the PyCaret framework;
then, we explored deep learning methods using PyTorch and Transformer-based architectures.</p>
        <p>In the next stage, attention was turned to the processing of the 2025 evaluation set, which resembles
the structural design of the 2024 data set. At this point, the focus was on filtering the data in preparation
for symptom-level classification.</p>
        <p>In this context, filtering refers to the application of the VADER algorithm - Valence Aware Dictionary
and sEntiment Reasoner - as a preprocessing step to identify texts that exhibit negative sentiment. This
selection criterion allowed us to isolate samples that were considered most likely to reflect psychological
distress, providing a more relevant subset for classification and scoring, an issue that was subsequently
addressed by two alternative modeling strategies.</p>
        <p>An overview of the workflow is presented in Figure 1, which illustrates the core stages of data
integration, such as cleaning, filtering, and running models in the machine learning and deep learning
paradigms.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Dataset overview</title>
        <p>The initial step involved reading and examining the accumulated text data stored in .TREC format. This
preliminary phase allowed us to understand the data structure we would be working with throughout
the challenge.</p>
        <sec id="sec-3-2-1">
          <title>3.2.1. Understanding the Document Structure</title>
          <p>The structure of the files provided for 2023 will be explained below using a sample.</p>
          <p>Data Overview</p>
          <p>Data Reading
Classification and Scoring</p>
          <p>Classification
Scoring algorithm
development and
execution</p>
          <p>Datasets
relationships
analysis</p>
          <p>Datasets</p>
          <p>Cleaning
2025 Filtering</p>
          <p>by VADER
Formatting the
requested file</p>
          <p>Machine
Learning</p>
          <p>Approach
Deep Learning</p>
          <p>Approach</p>
          <p>Model Training
2023 dataset
Treatment
2024 dataset
Treatment</p>
          <p>Training
dataset
+
&lt;DOC&gt;
&lt;DOCNO&gt;s_480_0_0&lt;/DOCNO&gt;
&lt;TEXT&gt;It came out in 1990 lol.&lt;/TEXT&gt;
&lt;/DOC&gt;
&lt;DOC&gt;
&lt;DOCNO&gt;s_480_1_0&lt;/DOCNO&gt;
&lt;TEXT&gt;I didn’t say he wasn’t charismatic.&lt;/TEXT&gt;
&lt;/DOC&gt;</p>
          <p>In the data set for 2024 was also included a thread-like structure comprising a central message and
reply sequences similar to those found on platforms such as Quora, Reddit, or Facebook. An example is
shown below:
&lt;DOC&gt;
&lt;DOCNO&gt;99001_0_0&lt;/DOCNO&gt;
&lt;PRE&gt;&lt;/PRE&gt;
&lt;TEXT&gt;My high school math teacher’s rule was that we had to
show all the steps...&lt;/TEXT&gt;
&lt;POST&gt;The exception to this rule was steps could be skipped if
she was able...&lt;/POST&gt;
&lt;/DOC&gt;
&lt;DOC&gt;
&lt;DOCNO&gt;99001_0_1&lt;/DOCNO&gt;
&lt;PRE&gt;&lt;/PRE&gt;
&lt;TEXT&gt;The exception to this rule was steps could be skipped
if she was able...&lt;/TEXT&gt;
&lt;POST&gt;I got a zero on a homework assignment because I just
wrote all the answers...&lt;/POST&gt;
&lt;/DOC&gt;</p>
          <p>Understanding the predecessor–response structure (PRE and POST) was key to extracting and
organising text entries during preprocessing. This same structure is present in the 2025 dataset, hence,
making an eficient, non-redundant parsing method, was regarded primary goal.</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>3.2.2. Structure of Classification Files</title>
          <p>Alongside the text corpora, classification files were provided for each dataset, formatted as follows:
query
1
This structure was interpreted as:
• query: Identifier referring to one of the 21 BDI-II symptom attributes.
• q0: A fixed value (0), not used further in processing.
• docid: Document identifier following the pattern s_X_Y_Z, where:
– s_X denotes a conversation or discussion thread.
– Y refers to a subgroup within that thread.</p>
          <p>– Z indicates the specific phrase within the subgroup.</p>
          <p>• rel: Relevance label; 1 for relevant, 0 for non-relevant.</p>
          <p>
            Two versions of the classification data were provided, based on diferent annotation philosophies: one
called majority vote, and the other unanimous agreement. Upon review [
            <xref ref-type="bibr" rid="ref5">5</xref>
            ], it was found that although
the documents retained the same symptom label under both schemes, their relevance annotations
sometimes difered.
          </p>
          <p>This discrepancy typically arose from two categories of text:
1. Texts describing events or states associated with someone other than the author. For example:
“I believe my aunt has been experiencing similar stress due to her previous job.”
2. Texts expressing opinions or reflections on events that are emotionally neutral or not personally
impactful. For instance:
“I found it sad that the film ended with the protagonists separated.”</p>
          <p>In light of these observations, the decision was to prioritise the use of texts labelled as relevant only
under the unanimous agreement criterion. This ensured a more contextually grounded training dataset,
favoring examples where the user’s emotional involvement was clearer and more likely to indicate
depressive symptoms.</p>
        </sec>
        <sec id="sec-3-2-3">
          <title>3.2.3. Relevance Discrepancy Statistics</title>
          <p>An audit of the datasets revealed notable numbers of disagreements between majority and unanimous
labels:
• For 2023, a total of 2,348 texts showed disagreement in relevance. Examples include:
DocID: s_975_61_2 Symptom: 1 Majority: 1 Consensus: 0</p>
          <p>DocID: s_993_582_1 Symptom: 14 Majority: 1 Consensus: 0
• For 2024, a total of 2,531 texts exhibited similar discrepancies. Examples include:
DocID: 49519_0_1 Symptom: 1 Majority: 1 Consensus: 0</p>
          <p>DocID: 414454_0_17 Symptom: 1 Majority: 0 Consensus: 1</p>
          <p>These discrepancies highlighted the importance of filtering and focusing on a single annotation
paradigm — in this case, the one based on unanimous consensus. Furthermore, as shown below,
a statistical analysis of each dataset was carried out to assess whether the usage of balancing and
compensation techniques was necessary or worth considering.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Class Distribution and Data Balancing Methodology</title>
        <p>Prior to the implementation of any regularisation strategies, a statistical analysis was conducted to
evaluate the distribution of symptom classes within the datasets for the years 2023 and 2024.</p>
        <p>This assessment exclusively considered texts deemed relevant under the unanimous annotation
criterion, as previously discussed. A summary of the number of symptom-annotated texts for each year
is illustrated in Figure 2.
(a) Frequency of relevant texts per symptom (2023)
(b) Frequency of relevant texts per symptom (2024)</p>
        <p>The analysis revealed a pronounced imbalance across the symptom classes. In particular, considering
2023 data set, symptoms 6, 11, and 13 were markedly underrepresented, whereas symptoms 9 and 10
appeared in disproportionately high quantities.</p>
        <p>Descriptive statistics indicated a mean of 117.9 texts per symptom, with a standard deviation of 60.43.
The interquartile range was defined by the first quartile ( Q1) at 76 and the third quartile (Q3) at 158.
Given this distribution, the 75th percentile (Q3) was chosen as the reference threshold for balancing
class frequencies.</p>
        <sec id="sec-3-3-1">
          <title>3.3.1. Balancing Strategy for the 2023 Dataset</title>
          <p>
            To mitigate the imbalance within the 2023 dataset, a two-stage augmentation strategy was adopted
for minority classes. Initially, the Easy Data Augmentation (EDA) technique was applied to generate
approximately 30% of the additional samples required. This included operations such as synonym
replacement and adverb insertion, following the methodology introduced in [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ].
          </p>
          <p>
            The remaining 70% of the augmented data was generated using prompt-based synthesis with the
o3-mini language model. Prompts were constructed following an EMS + BDI-II schema, enabling the
model to produce semantically coherent and symptom-aligned synthetic sentences. This methodology
draws inspiration from [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ], which employed Early Maladaptive Schemas for sentence generation.
          </p>
        </sec>
        <sec id="sec-3-3-2">
          <title>3.3.2. Balancing Strategy for the 2024 Dataset</title>
          <p>For the 2024 dataset, where the degree of class imbalance was less extreme, the Synthetic Minority
Oversampling Technique (SMOTE) was employed to interpolate new samples for underrepresented
classes.</p>
          <p>Simultaneously, classes exceeding the Q3 threshold were reduced through a targeted undersampling
procedure. Rather than applying random removal, a semantic preservation strategy was adopted to
retain the most representative samples. Specifically, TF-IDF vectorisation was used to compute centroid
vectors for each class. The cosine similarity of each text to its class centroid was then calculated and
only those samples with the highest similarity values, indicating maximum semantic representativeness,
were retained.</p>
          <p>This approach sought to ensure that core linguistic features for each symptom category were preserved
within the reduced subsets, avoiding the random elimination of important samples.</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>3.3.3. Final Training Corpus</title>
          <p>Following completion of the augmentation and reduction procedures, the balanced datasets for 2023
and 2024 were merged to form a unified training corpus. This final dataset contained approximately
450 samples per symptom, organised in a tabular format with two primary columns: one for the text
content and another specifying the corresponding BDI-II symptom label.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Machine Learning Modelling and Performance Evaluation</title>
        <p>Following the text preprocessing phase—including the removal of stopwords and non-informative
tokens—the modelling stage was initiated with the aim of developing a classifier capable of recognising
depressive symptoms and assigning texts to their corresponding BDI-II category.</p>
        <p>
          Two distinct modelling paradigms were explored. The first focused on traditional machine learning
techniques, implemented using the PyCaret and Scikit-learn libraries; the second centred
on deep learning approaches employing PyTorch and Transformer-based architectures. For the
classical pipeline (machine learning approach), a compact yet informative textual representation was
used, namely Term Frequency–Inverse Document Frequency (TF-IDF), limited to 5,000 features and
configured with bi-grams ( max_features=5000, ngram_range=(1,2)), as suggested in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
        <p>The PyCaret framework enabled rapid experimentation across a broad suite of supervised learning
algorithms using standardised pipelines and evaluation metrics. A total of eight classifiers were assessed:
Ridge Classifier, Support Vector Machine (linear kernel), Logistic Regression, Random Forest, Decision
Tree, k-Nearest Neighbours, Naïve Bayes, and AdaBoost. Each model was trained using five-fold
crossvalidation, and performance was primarily evaluated using the F1-score, as summarised in Table 1.</p>
        <p>Among the models tested, the Ridge Classifier yielded the highest performance, achieving an F1-score
of 0.9220, along with high precision and recall.</p>
        <p>Subsequently, the Ridge Classifier was further refined through hyperparameter optimisation using
GridSearchCV, a module within Scikit-learn that performs an exhaustive search across a predefined
parameter grid, employing cross-validation to ensure generalisability. This optimisation procedure led
to an enhanced final F1-score of 0.9283.</p>
        <p>The learning dynamics of the final model, including its convergence trajectory, were visualised
as shown in Figure 3, the learning curve demonstrated steady performance improvements and no
indications of overfitting.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Deep Learning Approach: Fine-Tuning Transformer-Based Models</title>
        <p>
          The second modelling strategy explored a deep learning approach grounded in pre-trained Transformer
architectures, with the aim of harnessing contextual embeddings to improve classification performance.
This methodology is based on contemporary advancements in transfer learning, whereby models trained
on large-scale language corpora are adapted to specific downstream tasks through fine-tuning [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>The selected backbone was distilbert-base-uncased, a distilled variant of BERT designed to
retain much of its representational power while significantly reducing model size and computational
demand [17]. This model was integrated via the Hugging Face transformers library. Input encoding
was handled using the DistilBertTokenizerFast, ofering eficient subword tokenisation aligned
with the base model’s vocabulary.</p>
        <p>To tailor the architecture for multi-class classification, the model’s output head was replaced with a
fully connected linear layer whose output dimension corresponded to the number of BDI-II symptom
classes.</p>
        <p>Training was conducted using the cross-entropy loss function, paired with the Adam optimiser. The
optimiser dynamically adjusted learning rates for each parameter based on estimates of the first and
second moments of gradients. During training, the randomly initialised output layer weights and
selected intermediate Transformer layers were updated via backpropagation.</p>
        <p>The dataset was partitioned as follows:
• Training set: 6,515 texts
• Validation set: 1,400 texts
• Test set: 1,407 texts</p>
        <p>The fine-tuning process spanned three epochs, with training and validation metrics monitored on a
per-batch basis. Initial results during the first epoch showed a training accuracy of 95.35% and validation
accuracy of 89.43%. By the final epoch, performance increased to 97.51% accuracy on the training set,
91.21% on the validation set, and 91.90% on the held-out test set.</p>
        <p>While these results were consistent with those obtained using classical machine learning methods,
it is noteworthy that the deep learning model required more computational resources and longer
training times. Interestingly, the optimised Ridge Classifier from the traditional pipeline achieved
comparable—and in some instances superior— F1-scores, while maintaining greater computational
eficiency and faster inference speed. This may be partially attributed to the relatively short length of
the texts in the dataset.</p>
        <p>Given the balance between performance gain and resource constraints, training over three epochs was
considered adequate for this iteration. Nonetheless, Transformer-based architectures remain promising
and merit further exploration in future studies.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. Semantic Scoring Methodologies for BDI-II Symptoms</title>
        <p>To assess the degree of semantic similarity between a given input phrase and a predefined set of
symptom-related expressions, two complementary methodologies were employed. These approaches
are grounded in contextual embeddings and make usage of established transformer-based models for
sentence representation.</p>
        <sec id="sec-3-6-1">
          <title>3.6.1. Contextual Similarity via Centroid Embeddings</title>
          <p>The first method relies on comparing an individual phrase against a class-level semantic centroid,
which acts as a condensed representation of prototypical expressions for a given symptom class. These
centroids were computed directly from the real training dataset, which consists of the merged datasets
from 2023 and 2024.</p>
          <p>
            The objective was to assign a continuous score to each input phrase , scaled in the range [
            <xref ref-type="bibr" rid="ref10">0, 10</xref>
            ],
which quantifies its semantic resemblance to a particular BDI-II symptom, as inferred from training
data.
          </p>
        </sec>
        <sec id="sec-3-6-2">
          <title>Formal Notation</title>
          <p>Let:
•  = {1, 2, . . . , 21} be the set of BDI-II symptom classes.
•  = {(, )}=1 the training dataset, where  ∈ R represents a sentence embedding ( = 384

as defined by ‘all-MiniLM-L6-v2‘), and  ∈  the corresponding symptom class.
• (· ) the sentence embedding function provided by ‘sentence-transformers‘.</p>
          <p>• ˆ the predicted symptom class for a new input , according to the trained classifier  ().
Semantic Centroid Definition For each class  ∈ , its semantic centroid was defined as the mean
embedding of all training samples belonging to that class:
  =
1</p>
          <p>∑︁
|| (,)∈
()
where  = { |  = } denotes the set of samples associated with class .</p>
          <p>Score Computation for New Inputs Given a new input phrase , its predicted class ˆ is obtained
via the trained classifier (both models defined formerly). The final score was then computed as the
cosine similarity between its embedding and the centroid of the predicted class:</p>
          <p>Score(, ˆ) = 10 · cos ((),  ^)
with cosine similarity defined as:
cos(, ) =  ·  (3)</p>
          <p>‖‖ · ‖ ‖</p>
          <p>
            The resulting score, scaled to [
            <xref ref-type="bibr" rid="ref10">0, 10</xref>
            ], served as an interpretable indicator of symptomatic relevance.
The pseudocode presented below illustrates the logic implemented to rank texts based on this criterion.
(1)
(2)
I n p u t :
− D a t a s e t D = { ( x _ i , y _ i ) }
− New input p h r a s e x
− T r a i n e d c l a s s i f i e r f ( x )
− Embedding model m ( . )
S t e p 1 : Group p h r a s e s by c l a s s s in S
S t e p 2 : F o r e a c h c l a s s s :
          </p>
          <p>
            Compute mu_s = mean o f m( x _ i ) where y _ i = s
S t e p 3 : P r e d i c t c l a s s : s _ h a t = f ( x )
S t e p 4 : Get embedding o f x : v_x = m( x )
S t e p 5 : S c o r e = 10 ∗ c o s i n e _ s i m i l a r i t y ( v_x , mu_s_hat )
O u t p u t : S c o r e in [
            <xref ref-type="bibr" rid="ref10">0 , 1 0</xref>
            ]
          </p>
        </sec>
        <sec id="sec-3-6-3">
          <title>3.6.2. Semantic Score via Symptom-Based Exemplars</title>
          <p>In contrast to the centroid-based method, the second approach employs a hybrid set of real and
synthetic exemplars to estimate semantic similarity. The synthetic exemplars were generated using the
OpenAI model O3-mini, this time leveraging a technique known as Generated Knowledge Prompting,
as introduced by Liu et al. in [18].</p>
          <p>
            The Semantic Score via Exemplars method assigns to each input phrase  a continuous score in
[
            <xref ref-type="bibr" rid="ref10">0, 10</xref>
            ], measuring its semantic closeness to a curated set of first-person exemplar sentences for each
BDI-II symptom.
          </p>
        </sec>
        <sec id="sec-3-6-4">
          <title>Formal Notation Let:</title>
          <p>•  = {1, 2, . . . , 21} denote the BDI-II symptom classes.
•  = {(1), (2), . . . , ()} the set of handcrafted or generated exemplars for class  ∈ .
• (· ) the embedding function for natural language input.</p>
          <p>• ˆ the class prediction for the input phrase  as determined by the classifier  ().
Score Computation Using Exemplars The score is computed as the average scaled cosine similarity
between the embedding of the input phrase and the embeddings of the  exemplars in the predicted
class ˆ:</p>
          <p>Score(, ˆ) =
|^| =1
10 ∑|︁^| ⎛ 1 + cos ︁( (), (^</p>
          <p>()))︁ ⎞
⎝ ⎠
2
(4)
where cos(· , · ) is defined as before. The following pseudocode illustrates the logic applied to rank
the texts according to the equation presented above.</p>
          <p>I n p u t :
− New input p h r a s e x
− T r a i n e d c l a s s i f i e r f ( x )
− Embedding model m ( . )
− D i c t i o n a r y o f e x e m p l a r s : e x e m p l a r s _ d i c t [ symptom_id ] = [ p h r a s e s ]
S t e p 1 : P r e d i c t c l a s s : s _ h a t = f ( x )
S t e p 2 : Get embedding o f x : v_x = m( x )
S t e p 3 : F o r e a c h e x e m p l a r x _ j in e x e m p l a r s _ d i c t [ s _ h a t ] :</p>
          <p>
            Compute v _ j = m( x _ j )
Compute s i m _ j = c o s i n e _ s i m i l a r i t y ( v_x , v _ j )
S c a l e s i m _ j t o [
            <xref ref-type="bibr" rid="ref10">0 , 1 0</xref>
            ] : s i m _ s c a l e d _ j = 1 0 ∗ ( 1 + s i m _ j )
          </p>
          <p>
            / 2
S t e p 4 : S c o r e = a v e r a g e ( s i m _ s c a l e d _ j )
O u t p u t : S c o r e i n [
            <xref ref-type="bibr" rid="ref10">0 , 1 0</xref>
            ]
          </p>
        </sec>
      </sec>
      <sec id="sec-3-7">
        <title>3.7. Final Dataset Processing and Evaluation Strategies</title>
        <p>
          After examining the 2025 test dataset, it was determined that it comprised approximately 17,384,603
texts. Given the computational power required to classify such a scale of data, a filtering mechanism
was introduced to reduce the volume of data processed. Specifically, we employed a pre-filtering step
designed to identify only those texts that exhibited a generally negative sentiment. To this end, the
VADER sentiment analysis tool (Valence Aware Dictionary and sEntiment Reasoner) was applied [19],
resulting in a reduced subset of 5,430,903 texts. This approach is consistent with prior methodologies
adopted in eRisk 2023 [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], where sentiment filtering efectively reduced the search space while preserving
meaningful signal.
        </p>
        <p>Having established the final evaluation corpus, the subsequent stage involved the application of
previously developed classification models and semantic scoring methodologies. Three distinct configurations
were implemented to generate the final results:
1. A Ridge-classifier , utilising the semantic centroid similarity strategy to assign scores to each
input phrase.
2. A second Ridge classifier , this time employing the exemplar-based semantic similarity method.</p>
        <p>Notably, this configuration demonstrated more optimistic scoring behaviour and thus was deemed
promising for further experimentation.
3. A deep learning classifier , specifically fine-tuned on relevant training data, combined with the
exemplar-based scoring methodology.</p>
      </sec>
      <sec id="sec-3-8">
        <title>3.8. Reflection on Evaluation Results</title>
        <p>The participation of COTECMAR-UTB in Task 1 of eRisk 2025 yielded an intermediate performance
outcome when benchmarked against the top-performing systems in the challenge. The submitted run,
titled ranked updated, was the best among the three submissions; it employed a Ridge classifier in
conjunction with an exemplar-based semantic similarity scoring mechanism. This run achieved 10th
place overall in the final ranking of participating systems.</p>
        <p>According to the oficial evaluation results, the COTECMAR-UTB system attained the following
scores under the unanimity-based assessment framework: AP = 0.042, R-Precision = 0.108, P@10 =
0.181, and NDCG = 0.243, as shown in Table 2. These figures positioned the team within the middle tier
of the ranked systems, showing potential yet leaving notable room for improvement. For comparison,
the same model evaluated under the majority vote scheme achieved AP = 0.077, R-Precision = 0.165,
P@10 = 0.414, and NDCG = 0.290.</p>
        <p>The rationale for emphasising the results under the unanimity scheme lies in its heightened
sensitivity to semantic and contextual precision. As discussed in preceding sections, the exemplar-based
scoring strategy is designed to capture fine-grained textual alignment with prototypical symptom
expressions—an alignment that is more rigorously evaluated under the unanimity setting, where all
expert assessors must agree on a document’s relevance.</p>
        <p>A comparative inspection reveals that leading teams—such as INESC-ID and
UET-PsycheWarriors—achieved substantially higher metrics (e.g., INESC-ID with AP = 0.269, P@10 = 0.509, and
NDCG = 0.561), demonstrating stronger alignment between sentence relevance scores and expert
annotations.</p>
        <p>NDCG</p>
        <p>While the exemplar-based approach adopted proved to be promising—particularly due to its
interpretability and synthetically enriched exemplars—several factors likely contributed to the observed
performance gap:
• Sparse representation coverage: The use of a limited set of exemplars per symptom, even
when generated synthetically, may have failed to suficiently encapsulate the broad diversity of
linguistic expressions associated with depressive symptoms.
• Context neglect: Although the system operated on sentence-level embeddings, the task allowed
for adjacent sentence context.
• Model depth and learning capacity: While Ridge classifiers ofer eficiency, they may be
outperformed by deeper models (e.g., fine-tuned Transformers) in capturing subtle nuances of
language. Incorporating deep contextual encoders with attention mechanisms could enhance the
precision of symptom detection.
• Score calibration: The cosine similarity score, although scaled, might not have been fully aligned
with the thresholds for relevance set by human assessors.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Task 2: Contextualized Early Detection of Depression</title>
      <p>
        The eRisk 2025 Task 2 introduces a new approach for early detection of depression symptoms by
analyzing user interactions on social media platforms [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This task difers from previous ones in that
it focuses on understanding the contextual nature of conversations. Rather than simply evaluating
isolated messages, the objective is to detect patterns of behavior across entire conversations, accounting
for both the temporal and emotional evolution of the user’s posts and comments. This makes the task
particularly challenging, as it requires capturing the subtle emotional shifts that may signal the onset
of depression, even when these shifts occur gradually over time.
      </p>
      <p>
        Building on previous editions of the eRisk Challenge, where early detection systems were applied to
detect risks such as pathological gambling [20], we chose to implement a Contextualized Early Detection
System using a combination of two models: CPI (Classifier with Partial Information) and DMC (Decision
Moment Classifier). This architecture has been applied in past challenges, where it proved efective in
detecting early signs of depression and other behavioral risks based on users’ interactions [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>The CPI model for our solution is built using a long-short-term memory (LSTM) network. LSTMs
are well suited for sequential data, as they can capture long-range dependencies in time-series data. In
our case, the model processes the messages exchanged by a user, considering the partial information
available at each moment. This allows the CPI model to classify whether a user’s message is indicative
of depression, using only the information available up to that point in the conversation.</p>
      <p>On the other hand, the DMC model adopts a fixed policy model approach, designed to assess the
emotional progression of the user over time. This model does not classify individual messages, but
instead tracks the trend of emotional content across multiple messages. By evaluating the overall
emotional trajectory of the conversation, the DMC determines if the user’s emotional state is improving
or worsening. This allows the system to make a decision about whether the user needs additional
intervention or support.</p>
      <p>The combination of the CPI and DMC models helps us create a system that not only classifies
individual messages, but also takes into account the contextual evolution of the conversation, leading to
a more accurate early detection of depression symptoms. This system can identify emotional changes
over time, even when individual messages do not explicitly indicate a risk of depression.</p>
      <p>The Figure 4 illustrates the system pipeline used in our solution. It details the sequential flow of data
and the various stages of processing, from data reading and text cleaning to the creation of batches,
re-labeling with VADER, vectorization with Word2Vec, and model training. Each stage plays a crucial
role in ensuring that the system can detect depression symptoms accurately and eficiently. The pipeline
will be explained in detail in the next section.</p>
      <sec id="sec-4-1">
        <title>4.1. Architecture explanation</title>
        <p>The architecture for the early detection of depression involves several key stages, from reading and
preprocessing the data to training the model and integrating them with the competition server.</p>
        <sec id="sec-4-1-1">
          <title>4.1.1. Data Reading</title>
          <p>The first step of the pipeline is to read the data, which is stored in XML files. These files contain textual
interactions of users, which can be posts or comments, ordered chronologically. Each file corresponds
to a single user and contains a sequence of writings, allowing the analysis of the user’s interaction
evolution over time. The dataset is labeled so that each user has a binary label: “1” (positive) if the user
is identified as at risk of depression, or “0” (negative) if the user belongs to the control group.</p>
          <p>The XML files are structured as follows:
&lt;INDIVIDUAL&gt;
&lt;ID&gt; ... &lt;/ID&gt;
&lt;WRITING&gt;
&lt;TITLE&gt; ... &lt;/TITLE&gt;
&lt;DATE&gt; ... &lt;/DATE&gt;
&lt;INFO&gt; ... &lt;/INFO&gt;
&lt;TEXT&gt; ... &lt;/TEXT&gt;
&lt;/WRITING&gt;
...
&lt;/INDIVIDUAL&gt;</p>
          <p>Where each XML document includes:
• ID: An anonymous user identifier. This ID is unique for each user and ensures the privacy of the
data.
• TITLE: The title of the post, if available. If the writing is a comment, the TITLE field will be
empty.
• INFO: Additional information about the origin of the post or comment (e.g., whether it comes
from a forum or a blog).</p>
          <p>• TEXT: The body of the post or comment written by the user.</p>
          <p>Regarding the label of the users:
• Depressed Users (labeled as “1”): These are users identified with signs of depression, either
because they explicitly mentioned their diagnosis or due to the analysis of their interactions over
time.
• Control Users (labeled as “0”): These users do not show signs of depression based on their
posts and comments.</p>
          <p>The dataset used in eRisk 2025 contains a considerable number of writings, this dataset is divided into
several editions, each with a specific number of control and positive users (those at risk of depression).
The total dataset size for training is as follows:
• T1 2017 Dataset: 752 control users and 135 positive users (with depression).
• T2 2018 Dataset: 741 control users and 79 positive users (with depression).</p>
          <p>• T2 2022 Dataset: 1,302 control users and 98 positive users (with depression).</p>
          <p>The users are represented by their writings, which are processed sequentially to identify patterns
indicating depression risk. Each dataset contains thousands of writings, enabling the models to train on
a large number of examples to detect signs of depression from the user’s texts.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Cleaning text</title>
          <p>After reading the dataset, the next step is to clean up the text in user posts, an essential process for
optimizing the data before it is processed by the model. This process includes:
• Conversion to lowercase: The text is normalized to lowercase to avoid inconsistencies due to
capitalization.
• Replacement of URLs and mentions: URLs and mentions of users or subreddits are replaced with
generic tokens, eliminating irrelevant information.
• Normalization of entities and special characters: Unicode entities and HTML characters are
converted to their readable form, eliminating unnecessary encoding.
• Noise removal: Irrelevant fragments such as URL parameters, external links, and empty words
that do not add value are removed based on Exploratory Data Analysis (EDA), where extensive
repetitions of words such as “Vik Vik ...” appeared.
• Replacement of numbers and removal of punctuation: Numbers are replaced with a generic token
and unnecessary punctuation is removed to avoid interference in the analysis.
• Filtering of short posts: Posts with fewer than five words are discarded, as they do not provide
suficient context.</p>
          <p>These steps ensure that the data is consistent and relevant, improving the accuracy of the model in
the early detection of depression.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.1.3. Batch Creation (Batches)</title>
          <p>The batch creation is a crucial experimental step in our pipeline, aimed at organizing user messages
in a way that maximizes computational eficiency and facilitates a more precise analysis of emotional
signals. This process involves several key aspects that optimize both the performance and accuracy of
the model.</p>
        </sec>
        <sec id="sec-4-1-4">
          <title>4.1.4. Grouping into Batches</title>
          <p>To develop the solution, we decided to group each user’s messages into batches of 10. This choice is
not arbitrary and is based on two main considerations. First, by dividing the messages into smaller
fragments, we optimize computational resources under Colab constraints. Processing large amounts of
data caused memory overload and increased processing time. Second, grouping messages allows us to
analyze the user’s overall emotional context, as a single message alone may not adequately reflect their
emotional state. By grouping messages, we can observe patterns of interaction that would be dificult
to detect in an isolated message.</p>
        </sec>
        <sec id="sec-4-1-5">
          <title>4.1.5. Progress Marker j/n</title>
          <p>To ensure proper evaluation of each batch, we use a j/n marker, where n represents the total number of
messages in a user’s conversation and j is the specific message number within the batch. This marker
provides a way to track the progress of processing and ensures that the temporal order of messages
within each batch is maintained. This approach also facilitates the integration of the full context of
each message within the batch, which is essential for precise emotional analysis.</p>
        </sec>
        <sec id="sec-4-1-6">
          <title>4.1.6. Reclassification within the Batch</title>
          <p>One of the most important goals of batch creation is the reclassification of messages . Not all messages
from a user clearly indicate signs of depression, so classification should be done considering the general
emotional tone of the batch as a whole. This approach allows the model to classify the messages
correctly by taking into account the flow and context of the interaction. Thus, a batch composed of
seemingly neutral messages could reveal significant emotional content when analyzed together. This
type of contextual analysis has proven essential for accurate user classification.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Relabeling with VADER</title>
        <p>Once the message batches have been created, the VADER (Valence Aware Dictionary and sEntiment
Reasoner) sentiment analysis tool is applied to relabel the emotional tone of the messages within each
batch. VADER is a particularly efective sentiment analysis model for informal texts, such as social
media messages, due to its ability to identify and classify emotions in complex and subjectively charged
contexts. It is an approach based on a lexical dictionary of emotions, together with syntactic rules that
allow the polarity (positive, negative, or neutral) and intensity of feelings expressed in the texts to be
analyzed [19].</p>
        <p>The analysis performed by VADER assigns a polarity score to each message, indicating whether
the overall tone is positive or negative. In addition, VADER calculates an intensity score for each
emotional category, allowing for a more nuanced analysis of the emotionality of messages. This feature
is especially useful for detecting subtle feelings or mixed emotions, which are common in user messages
and can be dificult to identify using simpler methods or methods based solely on keywords.</p>
        <p>In the context of detecting symptoms of depression, relabeling with VADER allows for refining
the classification of messages by adjusting labels based on emotional cues found in the texts. This
reassessment of the emotional charge of messages improves the accuracy of the model, especially in
situations where messages are not overtly negative, but contain emotional nuances that may indicate
depression.</p>
        <p>The use of VADER for this relabeling is a strategy proven in previous research, as the tool has been
shown to improve the ability of machine learning models to detect complex emotional states, such as
depression, by incorporating the analysis of polarity and emotional intensity. According to Hutto and
Gilbert [19], VADER is particularly efective at capturing the subtleties of informal texts on social media
platforms, making it ideal for tasks such as ours, where users often do not express emotions directly or
clearly.</p>
        <p>By reevaluating the emotional content of messages using VADER, we ensure that the model’s training
data reflects the emotional tone that is indicative of depression.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Embedding</title>
        <p>For this stage, Word2Vec was used with specific parameter settings that optimize both the quality of
semantic representations and computational eficiency. The model was trained with a vector_size
of 100, which provides an adequate balance between the ability to represent semantic relationships
and computational cost. The window size was set to 10, allowing for adequate capture of the context
of words within a suficiently wide range. The parameter min_count=2 was used to filter out
lowfrequency words, ensuring that only the most relevant ones influenced the representation. Training was
performed using 2 processing cores (workers=2), which reduced training time without compromising
model quality. In addition, the model was trained for 20 epochs, allowing adequate convergence without
the risk of overfitting. Finally, the random seed ( seed=42) was set to ensure the reproducibility of the
results. These parameters were selected after evaluating diferent configurations and seeking a balance
between accuracy and computational eficiency.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. CPI Training with LSTM</title>
        <p>The CPI model was trained using an LSTM (Long Short-Term Memory) to classify message batches.
The LSTM processes the vector representations of the messages, which have been obtained using
Word2Vec and labels adjusted by VADER. This type of neural network is particularly well suited for
time series tasks, such as ours, as it can capture interaction patterns over time. The ability of LSTMs to
remember relevant information throughout sequences allows them to identify emotional risks even
when signs of depression manifest gradually. The CPI model predicts the risk of depression based on
the partial information available in each message in the batch, learning to classify sequentially.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.5. DMC Modeling with Fixed Policy</title>
        <p>The DMC was modeled using a fixed decision policy, designed to evaluate the emotional trend of
messages over time. Unlike the CPI model, which classifies individual messages, the DMC does not
rely on a trained model but instead applies a set of rules to adjust decisions based on the evolution of
messages in the conversation.</p>
        <p>In this system, the CPI provides the grouped predictions of the user’s messages. These predictions,
which reflect the emotional tone of the messages, are passed to the DMC. The DMC then analyzes the
emotional trend by observing how the sentiment evolves across the messages. This allows the DMC to
make decisions regarding whether the user is at risk of depression (classification as 1) or not at risk
(classification as 0).</p>
        <p>The DMC measures the emotional trend by evaluating several key parameters:
• Median: The median of the predictions from previous batches is calculated to identify the general
trend.
• Trend: The diference between the first and last values in the sequence of predictions, which
indicates whether the user’s emotional state is improving or worsening.
• Standard deviation: Measures the dispersion of predictions, ensuring that decisions are not
made at times of high emotional variability.
• Maximum score: Verifies that the prediction score is above a minimum threshold to ensure that
emotional changes are significant.</p>
        <p>The _ function adjusts the intervention decision based on these parameters, using a
trend threshold and a maximum standard deviation to ensure that decisions are robust and that no
misclassifications are made due to minor emotional fluctuations. The fixed policy ensures that the
system acts consistently and in line with the user’s overall behavior, without the need for additional
training.</p>
        <p>The parameters selected for the DMC were determined after parameter search, evaluating diferent
combinations of threshold, slope weight, standard deviation limit, and minimum delay. The following
parameters were found to be optimal for this task: threshold = 0.6, slope weight = 0.4, std max = 0.2,
and min delay = 5. These values were selected based on the results of multiple tests, based on metrics
such as ERDE, precision, recall, and F1-score to evaluate model performance. The optimal parameter
combination was the one that yielded the best performance in terms of both detection accuracy and
decision latency.</p>
        <p>The DMC’s goal is to determine when a user shows clear signs of emotional deterioration or
improvement. If the model detects that the emotional trend indicates deterioration and that the dispersion
is low (i.e., that the emotional behavior is consistent), the system can make an intervention decision.</p>
        <p>This fixed policy-based approach is especially useful in scenarios such as ours, where detecting
emotional changes over time is the key to accurately classifying the risk of depression, without the
need for deep model training. The use of these adaptive rules improves the eficiency of the system,
while maintaining consistent and reliable classification based on the user’s emotional behavior.</p>
      </sec>
      <sec id="sec-4-6">
        <title>4.6. Results</title>
        <p>The results obtained during the evaluation phase of Task 2 are presented in Table 3. Key performance
metrics such as Precision (P@10), Recall (R), 1-score (F1), ERDE@5, ERDE@50, Latency, Speed, and
 are shown. Below, each of these metrics is briefly explained:
• Precision (P@10): Measures the proportion of relevant messages among the first 10 messages
recommended or classified by the model.
• Early Risk Detection Error: Measures the cost associated with early risk detection, evaluating
both the precision and the latency of the intervention depending on the number of messages
indicating by the @.
• Latency: Measures the number of writings processed before the true positive detection is made.</p>
        <p>It indicates how quickly the system can detect a risk case.
• Speed: Evaluates how fast the system processes messages and makes intervention decisions.</p>
        <p>It is calculated based on the median penalty of detected true positives, with a higher penalty
indicating slower detection.
• : A combined metric that takes into account both precision and latency, providing an
overall performance evaluation of the system in terms of its ability to act both precisely and
promptly.</p>
      </sec>
      <sec id="sec-4-7">
        <title>4.7. Conclusions on the Results</title>
        <p>The results obtained by our team in run 0 reflect a moderate level of performance. First, the model
achieved an F1 score of 0.40, which points to a reasonable balance between precision and recall; this
suggests that, although the system is capable of retrieving a reasonable number of relevant cases, its
overall efectiveness is limited by imbalances in these two key metrics. In particular, the recall value
of 0.65 indicates that the model managed to capture a significant proportion of the actual positive
cases, which is an encouraging sign for early detection. However, the low accuracy of 0.29 reveals that
the system stumbled upon a significant challenge in distinguishing the most relevant or informative
messages, resulting in a higher false positive rate. This imbalance reveals the need to further refine
the model’s decision-making processes, especially with regard to improving its capability to better
accurately filter and focus on high-risk content.</p>
        <p>In terms of latency, the value of 69.00 suggests that the system took a reasonable amount of time to
make decisions, but there is still room to optimize this metric to achieve faster interventions. ERDE@5
of 0.12 shows a relatively low cost for early risk detection, which means that the model was able to
make decisions without significant delays, although further optimization could reduce this cost even
more.</p>
        <p>In the ranking-based evaluation as shown in Table 4, the metrics P@10 and NDCG@10 were relatively
low, suggesting that the model had dificulty prioritizing the most relevant messages for the detection
of depression. This indicates that the system needs adjustments to improve the ranking of critical
messages.</p>
      </sec>
      <sec id="sec-4-8">
        <title>4.8. Future Work</title>
        <p>To improve the performance of the model in future editions of the challenge, the following strategies
are proposed:
• Optimization of the CPI model: Use more advanced architectures, such as BERT or DistilBERT,
which can better handle language complexities and provide better semantic representations of
messages.
• Adjusting the Decision Policy in the DMC: Explore approaches based on adaptive learning for
the fixed decision policy, so that the model can be dynamically adjusted according to the signals
in the messages.
• Improve relabeling with VADER: Evaluate additional sentiment analysis techniques or even
combine analysis tools to increase the accuracy of message relabeling.</p>
        <p>These strategies aim to improve system eficiency, increase prediction accuracy, and ensure that
intervention decisions are made at the right time.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>We would like to express our sincere gratitude to COTECMAR for providing the necessary space and
resources to carry out this research. We also extend our thanks to the Universidad Tecnológica de Bolívar
(UTB) for ofering the facilities and processing services for running the models, which significantly
contributed to the success of this work. Additionally, we acknowledge the support received through the
Convocatoria 950 of 2024 sponsored by Minciencias, which provided the resources for the scholarship
that made this research possible.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>Throughout the process of crafting this work, the authors used ChatGPT-4, ChatGPT-4o-mini, and
o3-mini AI models to assist with the enhancement and reorganisation of the text. These tools were
employed to improve pragmatic clarity, optimise coherence, and ensure consistency in scientific language.
Additionally, DeepL was utilised on certain occasions for translation purposes. Following the use of
these tools, the authors reviewed and edited the content as necessary and take full responsibility for
the publication’s content.
[17] V. Sanh, L. Debut, J. Chaumond, T. Wolf, Distilbert, a distilled version of bert: smaller, faster,
cheaper and lighter, 2019. URL: https://arxiv.org/abs/1910.01108. arXiv:1910.01108.
[18] J. Liu, A. Liu, X. Lu, S. Welleck, P. West, R. L. Bras, Y. Choi, H. Hajishirzi, Generated
knowledge prompting for commonsense reasoning, 2022. URL: https://arxiv.org/abs/2110.08387.
arXiv:2110.08387.
[19] C. Hutto, E. Gilbert, Vader: A parsimonious rule-based model for sentiment analysis of social
media text, Proceedings of the International AAAI Conference on Web and Social Media 8 (2014)
216–225.
[20] H. Fabregat, A. Duque, L. Araujo, J. Martinez-Romo, Nlp-uned-2 at erisk 2023: Detecting
pathological gambling in social media through dataset relabeling and neural networks., in: CLEF (Working
Notes), 2023, pp. 672–683. URL: https://ceur-ws.org/Vol-3497/paper-056.pdf.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Mansoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ansari</surname>
          </string-name>
          ,
          <article-title>Early detection of mental health crises through ai-powered social media analysis: A prospective observational study</article-title>
          ,
          <source>Journal of Personalized Medicine</source>
          <volume>14</volume>
          (
          <year>2024</year>
          )
          <fpage>153</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Detecting and measuring depression on social media using a machine learning approach: Systematic review</article-title>
          ,
          <source>JMIR Mental Health</source>
          <volume>9</volume>
          (
          <year>2022</year>
          )
          <article-title>e32786</article-title>
          . doi:
          <volume>10</volume>
          .2196/27244.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>M. M. Tadesse</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Detection of depression-related posts in reddit social media forum</article-title>
          ,
          <source>IEEE Access 7</source>
          (
          <year>2019</year>
          )
          <fpage>44883</fpage>
          -
          <lpage>44893</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2019</year>
          .
          <volume>2909180</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <article-title>Unravelling minds in the digital era: Mapping mental health disorders through machine learning using online social media, Social Network Analysis and Mining (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Martín-Rodilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , Overview of erisk at clef 2024:
          <article-title>Early risk prediction on the internet (extended overview)</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          (
          <year>2024</year>
          ). URL: https: //ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3740</volume>
          /paper-72.pdf,
          <source>presented at CLEF 2024, September 09-12</source>
          ,
          <year>2024</year>
          , Grenoble, France.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. H.</given-names>
            <surname>Ang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Gollapalli</surname>
          </string-name>
          , S.
          <article-title>-</article-title>
          K. Ng,
          <article-title>Nus-ids@erisk2024: Ranking sentences for depression symptoms using early maladaptive schemas and ensembles</article-title>
          , in: G. Faggioli,
          <string-name>
            <surname>N. F.</surname>
          </string-name>
          0001,
          <string-name>
            <given-names>P.</given-names>
            <surname>Galuscáková</surname>
          </string-name>
          , A. G. S. de Herrera (Eds.),
          <source>Working Notes of the Conference and Labs of the Evaluation Forum (CLEF</source>
          <year>2024</year>
          ), Grenoble, France,
          <fpage>9</fpage>
          -
          <issue>12</issue>
          <year>September</year>
          ,
          <year>2024</year>
          , volume
          <volume>3740</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>782</fpage>
          -
          <lpage>793</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3740</volume>
          /paper-73.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Beck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Steer</surname>
          </string-name>
          , G. Brown,
          <article-title>Beck Depression Inventory-II (BDI-II): Manual</article-title>
          , San Antonio, TX,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guecha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Potdar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miyaguchi</surname>
          </string-name>
          , Ds@gt erisk 2024:
          <article-title>Sentence transformers for social media risk assessment</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2407.08008. arXiv:
          <volume>2407</volume>
          .
          <fpage>08008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Thompson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cagnina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Errecalde</surname>
          </string-name>
          ,
          <article-title>Strategies to harness the transformers' potential: Unsl at erisk</article-title>
          <year>2023</year>
          ,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2310.19970. arXiv:
          <volume>2310</volume>
          .
          <fpage>19970</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>A.-M. Bucur</surname>
          </string-name>
          ,
          <article-title>Utilizing chatgpt generated data to retrieve depression symptoms from social media</article-title>
          ,
          <year>2023</year>
          . URL: https://arxiv.org/abs/2307.02313. arXiv:
          <volume>2307</volume>
          .
          <fpage>02313</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>J. M. Loyola</surname>
            ,
            <given-names>S. G.</given-names>
          </string-name>
          <string-name>
            <surname>Burdisso</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>L. C.</given-names>
          </string-name>
          <string-name>
            <surname>Cagnina</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          <string-name>
            <surname>Errecalde</surname>
          </string-name>
          , Unsl at erisk
          <year>2021</year>
          :
          <article-title>A comparison of three early alert policies for early risk detection</article-title>
          ,
          <source>in: Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2021</year>
          . URL: https://api.semanticscholar.org/CorpusID:237298519.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>J. M. Loyola</surname>
            ,
            <given-names>M. L.</given-names>
          </string-name>
          <string-name>
            <surname>Errecalde</surname>
            ,
            <given-names>H. J.</given-names>
          </string-name>
          <string-name>
            <surname>Escalante</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>Montes y Gomez, Learning when to classify for early text classification</article-title>
          ,
          <source>in: Proceedings of the Argentine Congress of Computer Science</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>24</fpage>
          -
          <lpage>34</lpage>
          . URL: https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -75214-
          <issue>3</issue>
          _3. doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>319</fpage>
          -75214-
          <issue>3</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , Overview of erisk 2025:
          <article-title>Early risk prediction on the internet, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and Interaction - 16th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2025</year>
          , Madrid, Spain, September 9-
          <issue>12</issue>
          ,
          <year>2025</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>II</given-names>
          </string-name>
          , volume To be
          <source>published of Lecture Notes in Computer Science</source>
          , Springer,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Perez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , Overview of erisk 2025:
          <article-title>Early risk prediction on the internet (extended overview)</article-title>
          ,
          <source>in: Working Notes of the Conference and Labs of the Evaluation Forum (CLEF</source>
          <year>2025</year>
          ), Madrid, Spain,
          <fpage>9</fpage>
          -
          <issue>12</issue>
          <year>September</year>
          ,
          <year>2025</year>
          , volume To be published of CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zou</surname>
          </string-name>
          , Eda:
          <article-title>Easy data augmentation techniques for boosting performance on text classification tasks</article-title>
          ,
          <year>2019</year>
          . URL: https://arxiv.org/abs/
          <year>1901</year>
          .11196. arXiv:
          <year>1901</year>
          .11196.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Raschka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Mirjalili</surname>
          </string-name>
          ,
          <article-title>Machine learning with pytorch and scikit-learn: Develop machine learning and deep learning models with python</article-title>
          ,
          <year>2022</year>
          . URL: https://github.com/rasbt/ machine-learning-book, iSBN:
          <fpage>9781801819312</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>