<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated Depression Detection in Text Data: Leveraging Lexical Features, Phonesthemes Embedding, and RoBERTa Transformer Model⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elizabeth Martinez</string-name>
          <email>jcmartinezs@utb.edu.co</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Cuadrado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daniel Peña</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Carlos Martinez-Santos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edwin Puertas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad Tecnologica de Bolivar, School of Engineering</institution>
          ,
          <addr-line>Cartagena de Indias 130010</addr-line>
          ,
          <country country="CO">Colombia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>Depression is a prevalent mental disorder characterized by persistent sadness, lack of interest, and diminished pleasure. Detecting depression is crucial for timely intervention and support. In this paper, we address the task of depression detection in text data, focusing on binary classification and regression. We present our approach, leveraging a dataset comprising labeled messages from Telegram groups related to mental disorders. We begin by exploring the existing literature on depression detection, highlighting the challenges faced and the methods employed. Our approach involves data pre-processing, lexical feature extraction, phonesthemes embedding, and using the RoBERTa transformer model. We achieved promising results in the training phase through rigorous experimentation and model refinement. However, we encountered challenges upon evaluating our approach in the MentalRiskEs evaluation. We identified areas for improvement, particularly in latency and speed of detection for real-time monitoring of depression-related risks. This research contributes to the ongoing eforts in automating depression detection and provides insights into the potential of text analysis techniques for mental health assessment. We remain committed to further enhancing our methodology and advancing the field to improve the well-being of individuals afected by depression.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Mental Risk</kwd>
        <kwd>Depression</kwd>
        <kwd>Lexical Features</kwd>
        <kwd>Phonesthemes Embedding</kwd>
        <kwd>Transformers</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Depression, a prevalent mental disorder afecting millions worldwide, poses significant
challenges regarding timely identification and support [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Traditional methods of depression
detection, which rely on clinical assessments and self-reporting, are limited by subjectivity
and resource constraints [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. As a result, there has been a growing interest in developing
automated approaches that utilize natural language processing and machine learning techniques
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Considerable research has been conducted in depression detection, exploring various linguistic
and contextual features, sentiment analysis, and machine learning algorithms [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ]. However,
accurately identifying depression from text remains a complex task due to the inherent subtleties
and contextual nuances associated with expressions related to mental health [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Furthermore,
the availability of annotated datasets for training and evaluation purposes is limited, which
further complicates the development of efective detection models [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        In this paper, we present our approach to depression detection using IberLEF 2023 Mental
Risk dataset , section 3, obtained from Telegram groups focused on mental disorders [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Our
methodology, section 4, consists of a pipeline that incorporates data pre-processing techniques,
such as removing stop words and punctuation, followed by extracting lexical features to capture
linguistic patterns [11] . To enhance the efectiveness of our approach, we also integrate
phonesthemes embedding, which encodes phonetic information and leverage the powerful
RoBERTa transformer named RoBERTuito model to capture contextual representations [12].
      </p>
      <p>
        During the training phase, section 5, our model achieved promising results, demonstrating the
accuracy of approximately 80% in binary classification Task 2a using SVM [ 13] and competitive
performance when incorporating lexical features, phonesthemes, and the RoBERTa transformer
[14]. However, in the evaluation conducted by IberLEF and the MentalRiskEs organization [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
section 6, our performance experienced a significant decline, indicating the need for further
exploration and improvement in our methodology. To address this, section 7, we shifted our
focus to analyzing individual messages separately, significantly improving binary classification
Task 2a. Our accuracy increased from 60.4% to 78.2%. It represents a notable gain of 17.8
percentage points. Furthermore, in the regression Task 2b, our Root Mean Square Error (RMSE)
decreased from 0.45 to 0.27 when incorporating lexical features, phonesthemes, and the RoBERTa
transformer. It is a substantial reduction of 0.18 in the RMSE metric. Nevertheless, additional
research and development are required to achieve optimal depression detection accuracy.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>In recent years, the automatic detection of depression using machine learning and natural
language processing methods has gained significant attention. Researchers have explored various
approaches to identify depression signals in text data, leveraging linguistic and contextual
features. This section provides an overview of previous studies in this domain, highlighting the
approaches, challenges, and results obtained.</p>
      <p>De Choudhury et al. [15] conducted one of the pioneering studies in this area. They utilized
data generated from interactions on Twitter to predict depression in users. The rise of social
media has provided much data for developing machine learning models focused on mental
health detection [16].</p>
      <p>Burdisso et al. [17] proposed a text classification framework for the early and efective
detection of depression in social media streams. Based on supervised learning techniques,
their approach outperformed standard models while providing computational eficiency and
explainability.</p>
      <p>Chiong et al. [18] investigated text pre-processing methods and feature extraction techniques
for depression detection using machine learning classifiers and social media texts. Their
approach efectively detected depression even without specific keywords related to depression in
the training datasets.</p>
      <p>Amanat et al. [19] proposed a model based on recurrent neural networks (RNN) and long
short-term memory (LSTM) for depression detection in textual data. Their approach achieved a
precision of 99.0 % in the early identification of depression, surpassing frequency-based models.</p>
      <p>Babu and Kanaga [20] conducted a review focused on sentiment analysis in social media for
depression detection using artificial intelligence. Their work emphasized multiclass
classification techniques and deep learning algorithms, highlighting the importance of obtaining more
accurate results in depression detection from social media text data.</p>
      <p>Mustafa et al. [21] employed word-based sentiment analysis and psychological attributes to
detect depression in social media. Their study used machine learning techniques to classify users
into three classes of depression: high, medium, and low. The results emphasized the significance
of feature selection and combination in enhancing classifier performance and precision.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Data</title>
      <p>A new dataset was made available for depression detection by the organizers of the Mental
Risk Challenge, IberLEF 2023. This dataset consists of labeled messages obtained from public
groups on the Telegram platform. The extraction and anonymization of conversations from
these groups were performed by IberLEF 2023, ensuring the privacy of the involved users.</p>
      <p>A team conducted a meticulous manual annotation process using the Prolific service by
IberLEF 2023 for the dataset labeling. Ten annotators carefully examined the history of each
user, determining the presence or absence of evidence indicating the targeted disorder. This
approach facilitated regression analysis, allowing the evaluation of prediction tools based on
alignment with the collective confidence of human judgments.</p>
      <p>We divided the dataset into three subsets, each associated with a diferent mental disorder.
Specifically, the depression corpus, comprising 335 users, was divided into trial, training, and
testing sets of 10, 175, and 150 users, respectively.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Architecture</title>
      <p>This section comprehensively describes the predictive model developed to address Task 2a
and Task 2b in the Mental Risk Challenge of the IberLEF 2023 competition. Task 2a involves
binary classification to detect depression in users based on their textual messages, while Task
2b focuses on estimating the probability of an individual sufering from depression.</p>
      <p>To tackle these tasks efectively, our model follows a systematic approach of several stages.
The first stage involves data reading and pre-processing. Multiple messages belonging to the
same user are concatenated into a single string, ensuring the continuity of information.</p>
      <sec id="sec-4-1">
        <title>4.1. Pre-Processing</title>
        <p>The text data from the IberLEF 2023 competition underwent a cleaning process to ensure its
quality and suitability for analysis. Specifically, we subjected the text data to a stop word
removal process.</p>
        <p>Stop words are commonly used words in a language that do not carry significant meaning.
We removed them to reduce noise and improve the accuracy of natural language processing
tasks. Examples of stop words include articles, prepositions, and conjunctions.</p>
        <p>By eliminating stop words from the text data, we aimed to enhance the clarity and focus of
the dataset, allowing the model to capture better the relevant information related to mental
health analysis.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Feature Extraction</title>
        <p>In sentiment analysis, feature extraction is crucial in capturing relevant linguistic aspects within
the text. This section focuses on incorporating various feature extraction techniques into our
model.</p>
        <sec id="sec-4-2-1">
          <title>4.2.1. Lexical Features</title>
          <p>Lexical features play a significant role in enhancing sentiment analysis by capturing important
linguistic characteristics within the text. Our model incorporates several lexical features,
including:
1. Personal Pronouns: The inclusion of first-person, second-person, and third-person
singular and plural pronouns provides valuable insights into self-referential and interpersonal
perspectives, which can be indicative of potential depressive tendencies.
2. Adverbs: Adverbs of time, negation, place, manner, and quantity are employed as lexical
features to provide contextual information, aiding in accurately interpreting temporal,
spatial, and emotional aspects within the text.
3. Adjectives: We considered negative and positive adjectives lexical features to capture the
emotional tone of the content and reflect pessimistic or optimistic sentiments.
4. Lexical Patterns: Specific lexical patterns, such as ’mention,’ ’URL,’ ’hashtag,’ ’emoji,’
and ’rt,’ reveal communication style and patterns associated with potential depressive
symptoms.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>4.2.2. Phonetic Embedding</title>
          <p>We utilized phonetic embedding as a feature extraction method to complement the sentiment
analysis of the textual data. This approach focuses on capturing the phonetic representations
of words, enabling the model to capture additional nuances of language. To achieve this,
pretrained Word2Vec models specifically trained on phonesthemes, the minor sound units in a
language, are leveraged. It enables the generation of vector representations that efectively
capture the phonetic structure of words.</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>4.2.3. Transformer RoBERTuito</title>
          <p>The Transformer RoBERTuito is employed to capture contextual relationships within the text.
This variant of the RoBERTa model utilizes deep learning techniques and extensive pre-training
on large text corpora. By leveraging its capabilities, RoBERTuito provides contextualized
embedding that encodes rich semantic and syntactic information. It enhances the model’s
understanding of sentiment-related concepts, idiomatic expressions, and subtle linguistic nuances,
significantly improving the accuracy of sentiment analysis predictions.</p>
        </sec>
        <sec id="sec-4-2-4">
          <title>4.2.4. Concatenated Features</title>
          <p>We concatenated the feature vectors obtained from the lexical features, phonetic embedding,
and Transformer RoBERTuito to consolidate the extracted features. This process results in a
single comprehensive feature vector that leverages a diverse range of linguistic information
for robust analysis and prediction. Combining the lexical features, phonetic embedding, and
contextual embedding from RoBERTuito, our model captures a comprehensive set of linguistic
cues and patterns, enabling a holistic understanding of the sentiment expressed in the text and
providing valuable insights into mental health conditions.</p>
          <p>Depending on the specific task, the consolidated feature vector serves as the input for
subsequent stages of the predictive model, such as classification algorithms or regression
models. These models utilize the extracted features to predict depression detection or estimate
the probability of an individual sufering from depression.</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Regularization</title>
        <p>We performed data augmentation to increase the diversity and representation of the dataset.
Specifically, we utilized the Synthetic Minority Over-sampling Technique (SMOTE). SMOTE
generated synthetic data points by interpolating between existing minority class samples. This
process efectively addressed class imbalances, improving the representation of
underrepresented instances.</p>
        <p>Furthermore, class balancing techniques were applied to address potential class imbalance
issues in the binary classification Task 2a as part of the regularization process. The aim was
to rectify the uneven distribution of instances across diferent classes, enhancing the model’s
performance and mitigating biases that could arise from an imbalanced dataset.</p>
        <p>It is important to note that the regression Task 2b, which focused on estimating the probability
of depression, did not involve class balancing techniques as it had a diferent objective. Therefore,
class balancing was not applicable in Task 2b.</p>
        <p>The model was trained and optimized by combining data augmentation and regularization
techniques to handle class imbalances and regression tasks efectively. This comprehensive
approach enhanced the model’s predictive capabilities for mental health analysis.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Classifiers</title>
        <p>This subsection provides an overview of the classifiers utilized for binary classification Task 2a
and regression Task 2b, focusing on their unique contributions and considerations. Scikit-learn
classifiers were chosen with default parameters and varying characteristics to ensure accurate
and reliable results in detecting depression severity. We will discuss the classifiers employed for
binary classification and regression, highlighting their unique contributions and considerations.</p>
        <sec id="sec-4-4-1">
          <title>4.4.1. Binary Classification Classifiers</title>
          <p>In order to assess their performance and predictive capability, we implemented several classifiers
for the binary classification Task 2a of detecting depression in text data. We utilized the following
classifiers:
• Random Forest (RF)
• Decision Tree (DT)
• Naive Bayes (NB)
• Logistic Regression (LR)
• Support Vector Machine (SVM)
• k-Nearest Neighbors (kNN)</p>
          <p>These classifiers were carefully selected based on their distinctive characteristics and
demonstrated efectiveness in accurately classifying instances in detecting depression in text data.</p>
        </sec>
        <sec id="sec-4-4-2">
          <title>4.4.2. Regression Classifiers</title>
          <p>For Task 2b, we employed a regressor to estimate the severity of depression in text data and
various regression classifiers. The classifiers utilized were as follows:
• Support Vector Machine (SVM)
• Ridge
• Linear Regression
• Lasso
• ElasticNet
• k-Nearest Neighbors (kNN)
• Decision Tree Regressor (DTR)</p>
          <p>These classifiers were selected based on their suitability for regression tasks and their potential
to provide precise estimations of depression severity in text data.</p>
        </sec>
        <sec id="sec-4-4-3">
          <title>4.4.3. Cross-Validation</title>
          <p>Cross-validation is a widely employed technique used to evaluate the performance of machine
learning models. Its primary objective is to estimate how well a model will generalize to unseen
data. This study employed cross-validation to reliably assess the classifiers and regressors used
in detecting depression in text data.</p>
          <p>Specifically, k-fold cross-validation was applied, dividing the dataset into k subsets (folds) of
similar sizes. Then, we trained the model and evaluated k times, employing each k subsets as a
test set once while we utilized the remaining k-1 subsets as the training set. This methodology
guarantees the utilization of all instances in the dataset for training and testing.</p>
        </sec>
      </sec>
      <sec id="sec-4-5">
        <title>4.5. Evaluation</title>
        <p>In this stage, we rigorously compared the results obtained from diferent models, and the model
with the best performance in each metric, or most metrics, is selected. Task 2a, which involves
binary classification, and Task 2b, which focuses on regression, are evaluated using specific
performance metrics.</p>
        <p>We utilized the following metrics for Binary Classification Task 2a: Accuracy, Precision,
Recall, and F1 score. These metrics are critical indicators for assessing the classifiers’ ability to
correctly classify instances and measure the balance between true positives, false positives, and
false negatives. By analyzing these metrics, we can determine which classifier performs better
in accurately detecting depression in text data.</p>
        <p>Regarding Regression Task 2b, the evaluation encompasses the following metrics: Root Mean
Square Error (RMSE), Pearson correlation coeficient, and Pearson’s recall. RMSE measures the
average diference between predicted and actual depression severity scores, while the Pearson
correlation coeficient quantifies the linear relationship between predicted and actual scores.
Additionally, Pearson recall evaluates the ability of regression classifiers to capture the recall of
depression severity levels accurately.</p>
        <p>By employing these comprehensive metrics, we ensure a thorough evaluation of the classifiers’
performance in both binary classification and regression tasks. The selection of the
bestperforming model in each task is based on an objective analysis of the model’s performance
across these metrics, enabling us to make informed decisions and derive reliable conclusions.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments Conducted and Training</title>
      <p>In the conducted experiments, we followed a step-by-step approach to classify depression
in text data. First, we evaluated the performance of classifiers using individual and parallel
feature extractions of phonesthemes and lexical features. Then, we examined the performance
of concatenated features, combining phonesthemes and lexical features. In the next step, we
incorporated transformer-based feature extraction, phonesthemes, and lexical features. We
repeated these steps for binary classification Task 2a and regression Task 2b. The experiments
involved evaluating various classifiers and measuring their performance using diferent metrics.
The results of these experiments provided insights into the efectiveness of diferent feature
combinations and classifiers for depression classification and regression tasks.</p>
      <sec id="sec-5-1">
        <title>5.1. Results of Training Process for Binary Classification Task 2a</title>
        <p>The binary classification Task 2a involved evaluating multiple models for diferent feature
combinations. The models assessed included Random Forest (RF) with DecisionTreeClassifier(200),
Decision Tree (DT) with DecisionTreeClassifier(4), Gaussian Naive Bayes (NB), Logistic
Regression (LR) with LogisticRegression(), Support Vector Machine (SVM) with SVC(), and K-Nearest
Neighbors (kNN) with KNeighborsClassifier(). These models were tested and compared in terms
of their performance for each feature combination.</p>
        <p>Table 1 displays the performance of the evaluated models for diferent feature combinations.
Notably, the SVM model consistently outperformed the other models. When considering the
phonesthemes approach, the SVM model achieved an accuracy of 0.80. Similarly, for the Lexical
approach, the SVM model demonstrated an accuracy of 0.78. Furthermore, the SVM model
maintained a high accuracy of 0.80 when incorporating phonesthemes, Transformers, and
Lexical features.</p>
        <p>The superior performance of the SVM model across multiple feature combinations highlights
its efectiveness for the task at hand. The high accuracy and precision make it the most
suitable choice for the phonesthemes and Lexical approaches and the combined phonesthemes,
Transformers, and Lexical approaches.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results of Training Process for Regression Task 2b</title>
        <p>The regression Task 2b involved evaluating various regression models for diferent feature
combinations. The models considered for evaluation were SVR (Support Vector Regressor), Ridge,
LinearRegression (LR), Lasso, ElasticNet, KNeighborsRegressor, and DecisionTreeRegressor.
These models were tested and compared using four distinct feature combinations: phonesthemes,
Lexical, phonesthemes, and Lexical, and phonesthemes, Transformers, and Lexical.</p>
        <p>As shown in Table 2, the results demonstrated that the Ridge regression model consistently
outperformed the other models across all feature combinations. Specifically, when combining
lexical, phonesthemes, and transformer features, the Ridge model achieved the best performance
with an RMSE (Root Mean Squared Error) of 0.29, a Pearson correlation coeficient of 0.81, and
a Pearson recall of 0.62.</p>
        <p>These results concluded that the Ridge regression model was the most suitable choice for the
regression task, particularly when considering the combination of lexical, phonesthemes, and
transformer features.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Results of UTB in MentalRiskEs Task Evaluation</title>
      <p>In this section, the results obtained by the UTB team in the evaluation of the MentalRiskEs
task are presented. The task was focused on the detection of mental disorders, with a specific
emphasis on the early identification of depression in Spanish comments from Telegram users.
Note that for each subtask, the exact same model was submitted by the UTB team for evaluation
three times.</p>
      <sec id="sec-6-1">
        <title>6.1. Task 2a: Depression Detection</title>
        <p>The UTB team approached Task 2a using various classification models to detect if users sufer
from depression. Table 3 showcases the classification-based evaluation results for Task 2a by
the UTB team. The table includes the team’s rank, run number, accuracy, macro-precision
(Macro-P), macro-recall (Macro-R), and macro-F1 (Macro-F1) scores. These metrics measure the
accuracy and overall performance of the team’s models in detecting depression.</p>
        <p>Furthermore, the team’s performance in terms of latency and speed of detection for Task
2a is presented in Table 4. The table includes the team’s rank, run number, ERDE5, ERDE30,
latencyTP, speed, and latency-weighted F1 scores. These metrics assess the team’s ability to
detect mental disorder risks promptly.</p>
        <p>The results indicate that the UTB team achieved competitive accuracy scores in depression
detection. However, we could explore further improvements in terms of latency and speed of
detection to enhance the real-time monitoring and identification of depression-related risks in
Telegram comments.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Task 2b: Regression-based Evaluation</title>
        <p>In Task 2b, the UTB team focused on regression-based evaluation to estimate the level of
afectation for users sufering from depression. Table 5 presents the team’s RMSE (root mean
square error) performance, which reflects the accuracy of their regression models. The table
includes the team’s rank, run number, and RMSE score.</p>
        <p>On the other hand, Table 6 presents the results of the ranking-based evaluation for the same
Task 2b in the NLPUTB project. The NLPUTB team also participated in three runs, identified
as 0, 1, and 2. However, in this evaluation metric, a p@30 score of 0.000 was obtained by the
NLPUTB team, indicating that correct rankings within the top 30 positions were not achieved.</p>
        <p>The UTB team’s results in the regression-based evaluation highlight their ability to accurately
estimate the level of afectation for users with depression. These findings demonstrate the team’s
competence in quantitatively utilizing regression models to assess mental health conditions.</p>
        <p>Overall, the UTB team showcased their proficiency in addressing classification and regression
tasks within the MentalRiskEs evaluation. Their results demonstrate their eforts in accurately
detecting mental disorders, estimating afectation levels, and considering the challenges posed
by the online nature of the problem.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Error Analysis and Discussion</title>
      <p>In this study, we initially developed an approach based on profiling for the binary classification
and regression tasks of mental health analysis. The profiling approach involved concatenating
messages and features to capture the overall behavioral patterns of individuals. While the initial
evaluation showed promising results, our approach performed diferently during the assessment
conducted by Mental Risk, where various research groups with diverse methodologies
participated. The evaluation incorporated additional metrics, such as p@k and latency, which we
had yet to consider in our approach initially. Consequently, the performance of our approach
significantly declined in all evaluation metrics.</p>
      <p>Upon further investigation and analysis, we hypothesized that the limitations of our initial
approach stemmed from its inability to capture the temporal dynamics of message content.
The concatenation of numerous messages without considering their chronological order and
the analysis of the entire dataset led to increased computational time, afecting the overall
performance. To address these challenges, we designed and implemented a new approach that
focused on considering the temporal aspect of message sequences.</p>
      <p>The new approach involved analyzing messages individually while preserving their
chronological order and incorporating temporal dependencies. By examining the messages in a time-based
sequence, we aimed to more efectively capture the evolving patterns and nuances of individuals’
mental states. Consequently, we observed notable improvements in several evaluation metrics.</p>
      <p>Table 7 summarizes the new approach’s performance compared to the regression task’s
profiling approach. The new approach, implemented using the Ridge regression model, achieved
an RMSE of 0.44, indicating a significant reduction in prediction error compared to the profiling
approach. However, it is essential to note that the Pearson correlation coeficient did not
demonstrate substantial improvement, suggesting that the revised approach may still benefit
from further refinement.</p>
      <p>For the binary classification Task 2a, Table 8 presents the classification-based evaluation
metrics. The new approach achieved an accuracy of 0.530, indicating an enhancement over the
profiling approach. However, the macro-Precision, macro-Recall, and macro-F1 scores were
slightly lower for the new approach, indicating that the revised methodology did not yield
significant improvements in the precision, recall, and F1 scores.</p>
      <p>Furthermore, Table 9 presents the latency-based evaluation metrics. While the new approach
demonstrated improvements in ERDE5 and ERDE30, the latency-weighted F1 score was slightly
lower than that of the profiling approach. It suggests that the revised approach may have
introduced specific latency-related challenges or trade-ofs.</p>
      <p>The findings from our study underscore the complexity of mental health analysis tasks
and the need for continuous exploration and refinement of approaches. Although the new
approach improved several evaluation metrics, we should have observed the areas where we
made the enhancements. Future research eforts will focus on addressing these limitations and
optimizing the performance of our approach across all evaluation metrics. One potential avenue
for improvement involves developing techniques to efectively capture the temporal dynamics
of message sequences while considering computational eficiency.</p>
    </sec>
    <sec id="sec-8">
      <title>8. Conclusion</title>
      <p>In conclusion, our study aimed to address the challenge of depression detection in text data
through binary classification Task 2a and regression Task 2b. Our approach involved a
comprehensive pipeline integrating data pre-processing, lexical feature extraction, phonesthemes
embedding, and the RoBERTa transformer model.</p>
      <p>During the training phase, our binary classification model achieved a good accuracy of
approximately 80%. However, when subjected to the evaluation of the MentalRiskEs task,
our performance experienced a decline, with lower rankings in Task 2a (23rd, 24th, and 25th
positions) and Task 2b (10th, 11th, and 12th positions). These results highlight the need for
further exploration and improvement in our methodology.</p>
      <p>To address the limitations identified during the evaluation, we explored alternative approaches.
By analyzing individual messages separately, we observed an improvement in our depression
detection performance. This shift increased accuracy, demonstrating the value of considering
individual messages as independent units of analysis.</p>
      <p>Furthermore, we examined the impact of our initial profiling approach and found no
significant diference in the evaluation metrics compared to our submitted results. This realization
emphasizes the ongoing need for discussion and exploration to enhance the binary detection
and regression tasks of depression.</p>
      <p>Moving forward, we will focus on refining our methodology to overcome the challenges
encountered during evaluation. It includes balancing the benefits of profiling and the need for
eficiency. Potential avenues for improvement include selective concatenation based on message
relevance or incorporating additional contextual information.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The authors would like to acknowledge the support provided by the master’s degree scholarship
program in engineering at the Universidad Tecnologica de Bolivar (UTB) in Cartagena, Colombia.
IberLEF 2023: Early Detection of Mental Disorders Risk in Spanish, Procesamiento del
Lenguaje Natural 71 (2023).
[11] E. A. Puertas Del Castillo, et al., Análisis de elementos fonéticos y elementos emocionales
para predecir la polaridad en fuentes de microblogging (2023).
[12] J. M. Pérez, D. A. Furman, L. A. Alemany, F. Luque, Robertuito: a pre-trained language
model for social media text in spanish, arXiv preprint arXiv:2111.09453 (2021).
[13] S. Huang, N. Cai, P. P. Pacheco, S. Narrandes, Y. Wang, W. Xu, Applications of support
vector machine (svm) learning in cancer genomics, Cancer genomics &amp; proteomics 15
(2018) 41–51.
[14] E. Puertas, J. C. Martinez-Santos, Phonetic detection for hate speech spreaders on twitter
(2021).
[15] M. De Choudhury, M. Gamon, S. Counts, E. Horvitz, Predicting depression via social
media, in: Proceedings of the Seventh International AAAI Conference on Weblogs and
Social Media (ICWSM), 2013.
[16] N. H. Di Cara, V. Maggio, O. S. P. Davis, C. M. A. Haworth, Methodologies for monitoring
mental health on twitter: Systematic review, J Med Internet Res 25 (2023) e42734. URL:
https://doi.org/10.2196/42734. doi:10.2196/42734.
[17] S. G. Burdisso, M. Errecalde, M. Montes-y Gómez, A text classification framework for
simple and efective early depression detection over social media streams, Expert Systems
with Applications 133 (2019) 182–197.
[18] R. Chiong, G. S. Budhi, S. Dhakal, F. Chiong, A textual-based featuring
approach for depression detection using machine learning classifiers and social
media texts, Computers in Biology and Medicine 135 (2021) 104499. URL: https://
www.sciencedirect.com/science/article/pii/S0010482521002936. doi:https://doi.org/
10.1016/j.compbiomed.2021.104499.
[19] A. Amanat, M. Rizwan, A. R. Javed, M. Abdelhaq, R. Alsaqour, S. Pandya, M. Uddin,
Deep learning for depression detection from textual data, Electronics 11 (2022). URL:
https://www.mdpi.com/2079-9292/11/5/676. doi:10.3390/electronics11050676.
[20] N. V. Babu, E. G. M. Kanaga, Sentiment analysis in social media data for depression
detection using artificial intelligence: a review, SN Computer Science 3 (2022) 1–20.
[21] R. U. Mustafa, N. Ashraf, F. S. Ahmed, J. Ferzund, B. Shahzad, A. Gelbukh, A multiclass
depression detection in social media based on sentiment analysis, in: S. Latifi (Ed.), 17th
International Conference on Information Technology–New Generations (ITNG 2020),
Springer International Publishing, Cham, 2020, pp. 659–662.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Wasserman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Iosue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wuestefeld</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Carli</surname>
          </string-name>
          ,
          <article-title>Adaptation of evidence-based suicide prevention strategies during and after the covid-19 pandemic</article-title>
          , World psychiatry
          <volume>19</volume>
          (
          <year>2020</year>
          )
          <fpage>294</fpage>
          -
          <lpage>306</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kim</surname>
          </string-name>
          , et al.,
          <article-title>Depression prediction by using ecological momentary assessment, actiwatch data, and machine learning: observational study on older adults living alone</article-title>
          ,
          <source>JMIR mHealth and uHealth 7</source>
          (
          <year>2019</year>
          )
          <article-title>e14149</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Bolton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gunnell</surname>
          </string-name>
          , G. Turecki,
          <article-title>Suicide risk assessment and intervention in people with mental illness</article-title>
          ,
          <source>Bmj</source>
          <volume>351</volume>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hirschberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>Advances in natural language processing</article-title>
          ,
          <source>Science</source>
          <volume>349</volume>
          (
          <year>2015</year>
          )
          <fpage>261</fpage>
          -
          <lpage>266</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Puertas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Moreno-Sandoval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Plaza-del Arco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Alvarado-Valencia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pomares-Quimbaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Alfonso</surname>
          </string-name>
          ,
          <article-title>Bots and gender profiling on twitter using sociolinguistic features</article-title>
          ,
          <source>CLEF (Working Notes)</source>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Moreno-Sandoval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Puertas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Plaza-Del-Arco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pomares-Quimbaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. AlvaradoValencia</given-names>
            , A.
            <surname>Ureña-López</surname>
          </string-name>
          ,
          <article-title>Celebrity profiling on twitter using sociolinguistic features notebook for pan at clef</article-title>
          <year>2019</year>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Puertas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Moreno-Sandoval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Redondo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Alvarado-Valencia</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>PomaresQuimbaya, Detection of sociolinguistic features in digital social networks for the detection of communities</article-title>
          ,
          <source>Cognitive Computation 13</source>
          (
          <year>2021</year>
          )
          <fpage>518</fpage>
          -
          <lpage>537</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kabir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ahmed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T. R.</given-names>
            <surname>Laskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. K.</given-names>
            <surname>Joarder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mahmud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <article-title>Deptweet: A typology for social media texts to detect depression severities</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>139</volume>
          (
          <year>2023</year>
          )
          <fpage>107503</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Shoaib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ei-Sappagh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Alenezi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gechev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <article-title>An advanced deep learning models-based plant disease detection: A review of recent research</article-title>
          ,
          <source>Frontiers in Plant Science</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>875</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>A. M. Mármol-Romero</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Moreno-Muñoz</surname>
            ,
            <given-names>F. M.</given-names>
          </string-name>
          <article-title>Plaza-del-</article-title>
          <string-name>
            <surname>Arco</surname>
            ,
            <given-names>M. D.</given-names>
          </string-name>
          <string-name>
            <surname>Molina-González</surname>
            ,
            <given-names>M. T.</given-names>
          </string-name>
          <string-name>
            <surname>Martín-Valdivia</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Ureña-López</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Montejo-Ráez</surname>
          </string-name>
          , Overview of MentalriskES at
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>