<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Attentive Multi-stage Learning for Early Risk Detection of Signs of Anorexia and Self-harm on Social Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Waleed Ragheb</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ome Aze</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Bringay</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maximilien Servajean</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AMIS, Paul Valery University - Montpellier 3</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IUT de Beziers, University of Montpellier</institution>
          ,
          <addr-line>Beziers</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIRMM UMR 5506, CNRS, University of Montpellier</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Three tasks are proposed at CLEF eRisk-2019 for predicting mental disorder using users posts on Reddit. Two tasks (T1 and T2) focus on early risk detection of signs of anorexia and self-harm respectively. The other one (T3) focus on estimation of the severity level of depression from a thread of user submissions. In this paper, we present the participation of LIRMM (Laboratoire d'Informatique, de Robotique et de Microelectronique de Montpellier) in both tasks on early detection (T1 and T2). The proposed model addresses this problem by modeling the temporal mood variation detected from user posts through multistage learning phases. The proposed architectures use only textual information without any hand-crafted features or dictionaries. The basic architecture uses two learning phases through exploration of state-of-theart deep language models. The proposed models perform comparably to other contributions.</p>
      </abstract>
      <kwd-group>
        <kwd>Classi cation</kwd>
        <kwd>LSTM</kwd>
        <kwd>Attention</kwd>
        <kwd>Temporal Variation</kwd>
        <kwd>Bayesian Variational Inference</kwd>
        <kwd>Anorexia</kwd>
        <kwd>Self-harm</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Anorexia is consider one of the most common eating disorder. It is characterized
by low weight, worry of gaining weight, and a powerful need to be skinny, leading
to food restriction. Many who su er from eating disorder see themselves as
overweight although they could be thin [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Individuals with eating disorders
have also been shown to have lower employment rates, in addition to an overall
loss of earnings. Eating disorder su erers who are experiencing an overall loss in
earnings associated with their illness are also magni ed by the excess of
healthcare costs. According to the National Eating Disorder Association (NEDA), up
to 70 million people worldwide su er from eating disorders [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Eating disorder
symptoms are beginning earlier in both males and females. As estimated, 1.1 to
4.2 percent of women su er from anorexia at some point in their lifetime [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Young people between the ages of 15 and 24 with anorexia have 10 times the
risk of dying compared to their same-aged peers.
      </p>
      <p>
        Self-harm is a very common problem, and many people are struggling to deal
with it [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Several illnesses are associated with self-harm, including borderline
personality disorder, depression, eating disorders, anxiety or emotional distress
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Self-harm occurs most often during the teenage and young adult begin around
age 14 and carry on into their 20s, though it can also happen later in life [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
There is also an increased risk of suicide in individuals who self-harm and it is
found in 40% to 60% of suicides [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Social media is becoming increasingly used not only by adults but also at
di erent age stages. Mental disordered patients also turn to online social media
and web forums for information on speci c conditions and emotional support.
Even though social media can be used as a very helpful tool in changing a
person's life, it may cause such con icts that can have a negative impact. This
puts responsibilities for content and community management for monitoring
and moderation. With the increasing number of users and their contents, these
operations turn out to be extremely di cult. Many social media try to deal with
this problem by reactive moderation. In reactive moderation, users report any
inappropriate, negative or risky user generated contents. However it may reduce
the workload or the cost of moderating, it is not enough especially for handling
mental disordered user's threads or posts.</p>
      <p>
        Previous researches on social media have established the relationship between
an individual's psychological state and hisnher linguistic and conversational
patterns [
        <xref ref-type="bibr" rid="ref18 ref19">19, 18</xref>
        ]. This motivate the task organizers to initiate the pilot task for
detecting depression from user posts on Reddit1 in eRisk-2017 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In
eRisk2018 the extension of the study was planned to include detection of anorexia. In
eRisk-2019, a continuation of anorexia tasks in addition to two other tasks are
proposed. One task is for early detection of signs of self-harm (T2). In this task
no training dataset is provided. Also, another new task for detection of severity
level of depression (T3) is presented. Tasks organizers proposed new evaluation
measures than what were used before.
      </p>
      <p>In this paper, we present the participation of LIRMM (Laboratoire
d'Informatique, de Robotique et de Microelectronique de Montpellier) in both
tasks for early detection of anorexia and self-harm in eRisk-2019. The originality
of our approach is to perform the detection through two main learning phases. In
the rst learning phase. we proposed Deep Mood Evaluation Module (DMEM)
that uses attention based deep learning models to construct a time series
representing temporal mood variation through users posts or writings. The second
phase is either to use machine learning or Bayesian inference model to obtain
1 Reddit is an open-source platform where community members (red-ditors) can
submit content (posts, comments, or direct links), vote submissions, and the content
entries are organized by areas of interests (subreddits).
the proper decision. The main idea is to give a decision once the models detect
clear signs of mental disorder from current and previous mood extracted from
the content.</p>
      <p>The rest of the paper is organized as follows. In Section 2, the related work
is introduced. Then in Section 3, a brief tasks (T1 and T2) description of early
risk detection and used datasets are presented. Section 4 presents the proposed
models. The experimental setup and all model variants used are introduced in
Section 5. In Section 6, the evaluation results and discussions are presented. We
conclude the study and experiments in Section 7.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Recent psychological studies showed the correlation between person's mental
status and mood variation over time [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. It is also evident that some mental
disordered may have chronic week-to-week mood instability. It is a common
presenting symptom for people with a wide variety of mental disorders, with
as many as 8 of 10 patients reporting some degree of mood instability during
assessment. These studies suggest that clinicians should screen for temporal
mood variation across most common mental health disorders.
      </p>
      <p>
        Concerning text representation, traditional Natural Language Processing
(NLP) modules start with feature extraction from text such as the count or
frequency of speci c words, prede ned patterns, Part-of-Speech tagging, etc.
These hand-crafted features should be selected carefully and sometimes with an
expert view. However these features are interesting [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], sometimes they loose the
sense of generalization. Another recent trend is the use of word and documents
vectorization methods. These strategies that convert either words, sentences or
even overall documents into vectors take into account all the text not just parts
of it. There are many ways to transform a text to high-dimensional space such
as term frequency and inverse document frequency (TF-IDF), Latent Semantic
Analysis (LSA), Latent Dirichlet Allocation (LDA), etc [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This direction was
revolutionized by Mikolov et al. [
        <xref ref-type="bibr" rid="ref16 ref17">16, 17</xref>
        ] who proposed the Continuous Bag Of
Words (CBOW) and skip-gram models known as Word2vec. It is a probabilistic
based model that makes use of a two layered neural network architecture to
compute the conditional probability of a word given its context. Based on this
work Le et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] propose Paragraph Vector model. The algorithm which is
also known as Doc2vec learns xed-length feature representations from
variablelength pieces of texts, such as sentences, paragraphs, and documents. Both word
vectors and documents vectors are trained using stochastic gradient descent and
back-propagation shallow neural network language models. The development
of Universal Language Model Fine Tuning (ULMFiT) is considered like
moving from shallow to deep contextual pre-training word representation [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This
idea has been proved to achieve Computer Vision (CV)-like transfer learning
for many NLP task. ULMFiT make use of the state-of-the art language model
AWD-LSTM (Average stochastic gradient descent - Weighted Dropout LSTM)
proposed by Merity et al. in 2017 [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The same 3-layer LSTM recurrent
architecture with the same hyperparameters and no additions other than tuned
dropout hyperparameters are used. The classi er layers above the base LM
encoder is simply a pooling layer (maximum and average pool) followed by three
fully-connected linear layers. The overall models signicantly outperforms the
state-of-the-art on six text classication tasks including three tasks for sentiment
analysis. In this paper, we will use these techniques for text representations.
      </p>
      <p>
        Attention mechanism is considered as one of the recent trends in NLP models
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It can be described as mapping a query and a set of key-value pairs to an
output, where the query, keys, values and output are all vectors. The output is
computed as a weighted sum of the values, where the weight assigned to each
value is computed by a compatibility function of the query with the
corresponding key. This can be seen as take a collection of vectors, whether it could be
a sequence of vectors representing a sequence of words, or an unordered
collections of vectors representing a collection of attributes and summarize them into
a single vector. This summarization is done by scoring each input sequence with
a probability-like scores obtained from the attention. This helps the model to
pay close attention to the sequence items with higher attention scores. In this
paper, we will evaluate the e ect of attention mechanisms on the model.
      </p>
      <p>In this paper, we will use deep attention based modi cation of ULMFiT
classi er to construct a time series representing temporal mood variation. We the
used classical machine learning and statistical models to get the nal decisions.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Tasks Description</title>
      <p>
        In CLEF eRisk 2019, three tasks are presented [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The rst task (T1) is for
early detection of signs of anorexia. It is a continuation of the same task in
eRisk2018. The second one (T2) is a new task in 2019 for early detection of signs of
self-harm. No training data is provided for this task. Another task was proposed
(T3) for measuring the severity of the signs of depression. In this section we will
describe the rst two tasks (T1 and T2) that we have participated on.
      </p>
      <p>
        Both tasks are considered as a binary classi cation problem. The datasets
are a dated textual data of user posts and comments -posts without titles- on
Reddit. The training and testing datasets are provided in stream of user writings
(posts and comments). The stream is ordered chronologically. A brief statistics
and summary for these datasets are provided in Table 1. Task organizers set
up a server that iteratively gives user writings to the participating teams. The
goal is not only to perform classi cation but also to do it as early as possible
using minimum amount of writings for each user. A decision must be sent after
processing each user writing to continue receiving more. This decision could be
positive risk case or postponed for future writings. A detailed description of
the tasks and used evaluation metrics can be found in the corresponding task
description paper [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
The temporal aspects of the eRisk tasks inspired us to model the temporal
mood variation through user's text content. The average number of days ranging
from the rst submission to the last submission is approximately 600 days. So,
determining the way in which user's posts and comments vary from positive
to negative and vice versa through time is worth inspecting. In the proposed
models, the main idea is to process user writings for each user and determine
the probability of how positive or negative it is. A detailed description of our
model can be found in the working notes paper of eRisk 2018 [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The proposed
architecture of our models comes in three main steps.
      </p>
      <sec id="sec-3-1">
        <title>Step 1 - Text Vectorization Module: It is considered as language mod</title>
        <p>eling step. The input of this step is the textual training datasets and the output
is text vectorization model.</p>
        <p>Step 2 - Mood Evaluation Module: This step is considered as the rst
supervised learning phase. Assign to each writing a probability like score
representing how positive (risky) the submission is. The output of this step is a time
series representing the mood variability over time. These time series will be the
training set of the second learning phase.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Step 3 - Temporal Modeling Module: Another learning phase is to build</title>
        <p>machine learning models to learn some patterns from these time series to come
up with the nal classi cation model.</p>
        <p>
          We tried to encapsulate text vectorization and mood evaluation modules
and proposed Deep Mood Evaluation Module (DMEM). This module is based on
ULMFiT architecture [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and the idea of transfer learning for language modeling
in addition to using attention layers for classi cations. In addition, we tried
Bayesian Variational Inference (BVI) [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] for the second learning phase.
We propose a modi cation of the basic architecture of the ULMFiT by adding
attention to the model. The proposed architecture will help the model to focus
on the important parts of the text that in uence the network decision. Figure
1 shows the proposed model and the separation between encoder layers (text
vectorization module) and classi er layers (mood evaluation module).
        </p>
        <p>i = fW i:Xig
Si = log[</p>
        <p>exp( i)
PN
j=1 exp( ji)</p>
        <p>]
Oi = Si</p>
        <p>Xi
Oi =</p>
        <p>X Si</p>
        <p>Xi
(1)
(2)
(3)</p>
        <p>Where W i is the weight of the attention layer of the ith sequence. The
attention scores Si is used to compute the scored sequence Oi = fo1; oi2; oi3; : : : ; oiN g
i
which has the same length as the input sequence.</p>
        <p>Since the input sequence to the attention layer (encoder output) resulted
from Bi-LSTM layers, the last element in the scored output SNi can be used
for representing the whole sequence. The whole sequence is represented by the
weighted sum of all output sequences Oi.</p>
        <p>The input sequence is passed to the embedding layer then the three Bi-LSTM
layers to form the output of the encoder. The encoder output has the form of
Xi = fx1; xi2; xi3; : : : ; xiN g where N is the sequence length. The attention layer
i
takes the encoded input sequence and computes the attention scores Si. The
attention layer can be viewed as a linear layer without bias.</p>
        <p>For classi cation layers, a simple concatenation between the maximum and
average pooling in addition to the scored output is inputted to a group of two
di erent sizes fully connected linear layers. The output of the last linear layer is
passed to the Softmax to form the network decision.</p>
        <p>
          Training the over whole models comes into three main steps proposed in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
1. The LM is initialized by training the encoder on a general-domain corpus
(Wikitext-103 dataset [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]). This helps to capture general features of the
language. Preserve low-level representations and adapt high-level ones
2. The pre-trained LM is ne-tuned using the training datasets for both tasks.
3. The classi er and the encoder is ne-tuned on the target task using di erent
strategies for each layer group.
        </p>
        <p>
          The training of the architecture is done using slanted triangular learning rates
(STLR), discriminative ne-tuning (Discr) and layers gradual unfreezing
proposed for ULMFiT with the same hyperparameter settings [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We train the
model on the forward language models for both the general-domain and task
speci c datasets. Training the attention layer uses the same learning rates and
cycles used in the classi cation layers group.
4.2
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Bayesian Variational Inference (BVI)</title>
        <p>
          We can represent the problem of classifying users from the already classi ed
(observed) writings as a variant of independent Bayesian classi er combination
[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Figure 2 shows the graphical model for the proposed BVI where the observed
random variable Wik represents if the ith writing for the kth user if it is classi ed
as positive or negative such that:
The variables , , and are the hyper-parameters re ecting our a priori
belief about the proportion of positive and negative users.
        </p>
        <p>
          We are interested in the posterior distribution of the random variable Uk,
that de nes if the user is positive or negative, which is unfortunately intractable.
We use a variational inference approach to compute an approximation such as
in [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. The approximation is obtained by solving the following equation for all
variables Zi conditionned on the observed data X:
        </p>
        <p>log qi(ZijX) = Ej6=i[log p(Z; X)] + const:</p>
        <p>So, we start from a number of positive and negative user writings (N d) where
d 2 f+; g for positives and negatives respectively. More speci cally:</p>
        <p>Then, the expected number of positive and negative writings for positive
users can be represented by N1+ and N1 respectively. The same for negative
users is N0+ and N0 . These values are computed as:</p>
        <p>Nrd = X E[1[uk = d]]:[1[wik = r]];</p>
        <p>d 2 f+; g; r 2 f0; 1g
k;i</p>
        <p>The hidden variable uk represents if the user will be classi ed as at-risk
(anorexia, self-harm) or not. So we can say:
(4)
(5)
(6)
(7)
(8)
(9)
N + =</p>
        <p>X 1[Wik = 1];
k;i</p>
        <p>N
= X 1[Wik = 0]
k;i</p>
        <p>We can estimate the expectation of the log of the probability to observe
positive writings independently of the user category as E[ln( )] and for negative
writings as E[ln(1 )] such that:</p>
        <p>Bernoulli( uk )
Beta( ; )
Bernoulli( )</p>
        <p>Beta( ; )
E[1</p>
        <p>E[ln( )] =
ln( )] =
( + N +)
( + N )
( +
( +
+ N + + N )
+ N + + N )</p>
        <p>Where is the digamma function de ned as the logarithmic derivative of the
gamma function. In addition, we can estimate the expectation of the log
probability for positive users to write positive writings as E[ln( 1)] and for negative
users as E[ln( 0)] where:</p>
        <p>E[1</p>
        <p>E[ln( i)] =
ln( i)] =
( + Ni+)
( + Ni )
( +
+ Ni+ + Ni )
+ Ni+ + Ni )
So, the expectation of a user to be positive or negative can be obtained as:
(11)</p>
        <p>Mk
ln( jk) = X Wik E[ln( j )] + (1</p>
        <p>
          Wik) E[ln(1
Where E[1[Uk = j]] is a normalized value for the two types of users (at-risk or
controlled). We can evaluate an optimal value for it iteratively by rst initializing
all factors, then updating each in turn using the expectations with respect to
the current values of the other factors [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
5
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <p>For each task, each team could participate with di erent ve runs. We create
di erent variants of our proposed architecture. In this section, we will present
all these variants, training procedures and model hyperparameters.
5.1</p>
      <sec id="sec-4-1">
        <title>Proposed Model Variants</title>
        <p>
          All the proposed model variants for both tasks are based on two supervised
learning phases (step 2 and step 3 in temporal mood variation model). For
selfharm detection task (T2), as there is no training data, we train our models
on the depression and anorexia datasets of eRisk-2018 [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. We assumed that
if a person with a clear signs of depression and/or anorexia could think about
harm himself. We used the DMEM module as the rst learning phase an all
the variants and tried di erent machine learning and statistical methods as the
second learning phase. Table 2 shows the used model for the second learning
phase in all the runs for both tasks. MLP stands for Multi-Layer Perceptrons
and RF is for Random Forest. All models that do not employ another learning
phase are marked by dashes. In these runs, we used simple counting thresholds
for successive positive classi ed writings.
5.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Model Training and hyperparameters</title>
        <p>We processed the training and testing streams of user writings by moving
window concatenation of size (N ). In other words, to give a decision about the
current writing at time (t), we process all user writing starting from (t N + 1).
This gives more information about the context of a writing and reduce the e ect
of noisy and irrelevant ones. Experiments show that (N = 5) to be a reasonable
choice for the window size.</p>
        <p>
          For DMEM, we use the same set of hyperparameter of AWD-LSTM proposed
by [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] replacing the LSTM with Bi-LSTM and keep the same embedding size of
400 and 1150 hidden activations. We used weighted dropout of 0:2 and 0:25 as
the input embedding dropout and the learning rate is 0:004. We ne-tuned the
LM by eithr anorexia or depression training datasets provided. We train the LM
for 14 epochs using batch size of 128 and limit the number of vocabulary to all
token that appear more than twice. For classi er, we used masked self-attention
layers and concatenation of maximum and average pooling. For the linear
block, we used hidden linear layer of size 100 and apply dropout of 0:4. We
used Adam optimizer [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] with 1 = 0:8 and 2 = 0:99. The base learning rate
is 0:01. We used the same batch size used in training LMs. For training the
classi er, we create each batch using weight random sampling to handle the
problem of imbalance in the datasets. We train the classi er on training set for
30 epochs and select the best model on validation set to get the nal model.
For T2 training, we combine the training datasets for depression and anorexia
of eRisk-2018.
        </p>
        <p>In the second learning phase, the used architecture of the MLP had two
hidden layers with ten neurons each. Concerning the RF classi er, ten estimators
were used. These models are used to classify time series of (N ) points. For
MLP, RF and BVI models in T1, positive users were reported for those with
classi cation probability higher than 0.8. This value increases to 0.9 in T2. We
set both thresholds to 0.6 in the last rounds. For some model variants (LIRMMC
and LIRMMD in T1 and LIRMMA and LIRMMB in T2), we apply counting
of successive positive writings and give a decision after either 5 or 10 following
writings respectively.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results &amp; Discussions</title>
      <p>
        In eRisk-2019 two di erent types are used for model evaluation. The rst one is
decision-based evaluations; where the classical classi cation measures - precision
(P), Recall (R) and (F1) - are computed for positive (at-risk) user. In addition
to these and due to the drawbacks of ERDE measure, a new latency weighted
F1 measure is introduced [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The other complementary evaluation is
rankingbased evaluation. Beside the red decision, scores are computed and used to
build a ranking of users in decreasing estimation of risk. We participated only
for decision-based evaluation. Tables 3 and 4 show the evaluation results of
all our proposed variants for both tasks. It is clear that using MLP for the
second learning phase is the best choice for both tasks. However, the usage of
high threshold in T2 make the models predict most of the positive user in late
writings. Also, applying BVI gets more comparable results than the runs with
simple counting of positive writings. But it needs more precise choice of threshold
for early detection in both tasks.
      </p>
      <p>Tables 5 and 6 show some statistics of other participants runs compared to
our proposed models. The ranks of the best run for each evaluation metric are
also included. The statistics of the anorexia task are for 54 runs of 13 teams.
The self-harm task statistics on results are for 33 runs of 8 teams. However the
proposed architecture does not include any hand-crafted features, it seems to be
comparable with other contributions for both tasks. Also, combining anorexia
and past eRisk depression training datasets for detecting signs of self-harm is
very competitive.
In this paper we present the participation of LIRMM in the CLEF eRisk-2019
T1 and T2 tasks. Both tasks are for early detection of signs of anorexia and
self-harm from users posts on Reddit respectively. We proposed ve runs for
each task and the results are interesting and comparable to other contributions.
The proposed framework architecture used the text without any handcrafted
features. It performs the classi cation through two phases of supervised learning
using state-of-the-art deep language modeling neural network. The rst learning
phase builds a time series representing the mood variation using attention-based
modi cation of the ULMFiT model. The second learning phase is another
classi cation model that learns patterns from these time series to detect early signs
of such mental disorders. In this phase, We tried set of machine learning (MLP
and RF) and statistical (BVI) models.</p>
      <p>Combining anorexia and previous eRisk depression datasets to detect early
signs of self-harm (T2) is interesting and shows the correlation of such mental
disorders. However, the proposed models need tuning of second learning phase
classi cation thresholds for earlier risk detection.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>We would like to acknowledge La Region Occitanie and l'Agglomeration Beziers
Mediterranee which nance the thesis of Waleed Ragheb as well as INSERM
and CNRS for their nancial support of CONTROV project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>1. The national eating disorders association (NEDA).: Envisioning a world without eating disorders</article-title>
          .
          <source>In: The newsletter of the National Eating Disorders Association. Issue</source>
          <volume>22</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.:</given-names>
          </string-name>
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
          <source>In: International Conference on Learning Representations (ICLR)</source>
          .
          <source>vol. abs/1409.0473 (Sep</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Doyle</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Treacy</surname>
            ,
            <given-names>M.M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheridan</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          :
          <article-title>Self-harm in young people: Prevalence, associated factors, and help-seeking in school-going adolescents</article-title>
          .
          <source>International journal of mental health nursing 24</source>
          <volume>6</volume>
          ,
          <issue>485</issue>
          {
          <fpage>94</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Dozat</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manning</surname>
          </string-name>
          , C.D.:
          <article-title>Deep bia ne attention for neural dependency parsing</article-title>
          .
          <source>vol. abs/1611</source>
          .01734 (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hawton</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zahl</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weatherall</surname>
          </string-name>
          , R.:
          <article-title>Suicide following deliberate self-harm: longterm follow-up of patients who presented to a general hospital</article-title>
          .
          <source>British Journal of Psychiatry</source>
          <volume>182</volume>
          (
          <issue>6</issue>
          ),
          <volume>537542</volume>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hoek</surname>
          </string-name>
          , H.:
          <article-title>Review of the worldwide epidemiology of eating disorders</article-title>
          .
          <source>In: Current Opinion in Psychiatry</source>
          . vol.
          <volume>29</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruder</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Universal language model ne-tuning for text classi cation</article-title>
          .
          <source>In: Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          . pp.
          <volume>328</volume>
          {
          <issue>339</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Joyce</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sulkowski</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.:</surname>
          </string-name>
          <article-title>The diagnostic and statistical manual of mental disorders: Fifth edition (dsm-5) model of impairment</article-title>
          .
          <source>In: Assessing Impairment: From Theory to Practice</source>
          . pp.
          <volume>167</volume>
          {
          <issue>189</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Klonsky</surname>
            ,
            <given-names>E.D.:</given-names>
          </string-name>
          <article-title>The functions of deliberate self-injury: A review of the evidence</article-title>
          .
          <source>Clinical psychology review 27</source>
          ,
          <volume>226</volume>
          {
          <volume>39</volume>
          (04
          <year>2007</year>
          ). https://doi.org/10.1016/j.cpr.
          <year>2006</year>
          .
          <volume>08</volume>
          .002
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of sentences and documents</article-title>
          .
          <source>In: ICML. JMLR Workshop and Conference Proceedings</source>
          , vol.
          <volume>32</volume>
          , pp.
          <volume>1188</volume>
          {
          <fpage>1196</fpage>
          . JMLR.org (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.: erisk 2017:
          <article-title>Clef lab on early risk prediction on the internet: Experimental foundations</article-title>
          .
          <source>In: 8th International Conference of the CLEF Association</source>
          . pp.
          <volume>346</volume>
          {
          <fpage>360</fpage>
          . Springer Verlag (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <article-title>Overview of eRisk { Early Risk Prediction on the Internet</article-title>
          .
          <article-title>In: Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Ninth International Conference of the CLEF Association (CLEF</source>
          <year>2018</year>
          ). Avignon, France (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <source>Overview of eRisk</source>
          <year>2019</year>
          :
          <article-title>Early Risk Prediction on the Internet</article-title>
          .
          <source>In: Experimental IR Meets Multilinguality, Multimodality, and Interaction. 10th International Conference of the CLEF Association</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2019</year>
          . Springer International Publishing, Lugano, Switzerland (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Maas</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daly</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>P.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potts</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Learning word vectors for sentiment analysis</article-title>
          .
          <source>In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1</source>
          . pp.
          <volume>142</volume>
          {
          <fpage>150</fpage>
          . HLT '
          <volume>11</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Merity</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keskar</surname>
            ,
            <given-names>N.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
          </string-name>
          , R.:
          <article-title>Regularizing and optimizing LSTM language models</article-title>
          .
          <source>In: International Conference on Learning Representations</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          . pp.
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
          . Curran Associates, Inc. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yih</surname>
          </string-name>
          , S.W.t.,
          <string-name>
            <surname>Zweig</surname>
          </string-name>
          , G.:
          <article-title>Linguistic regularities in continuous space word representations</article-title>
          .
          <source>In: Proceedings of the</source>
          <year>2013</year>
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HLT-</article-title>
          <year>2013</year>
          ).
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Moulahi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aze</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bringay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dare to care: A context-aware framework to track suicidal ideation on social media</article-title>
          .
          <source>In: Bouguettaya A</source>
          . et al. (eds) Web Information Systems Engineering - WISE
          <year>2017</year>
          .,Lecture Notes in Computer Science,. Springer, Cham. vol.
          <volume>10570</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Paul</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>You are what you tweet: Analyzing twitter for public health</article-title>
          . In: Adamic,
          <string-name>
            <given-names>L.A.</given-names>
            ,
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.A.</given-names>
            ,
            <surname>Counts</surname>
          </string-name>
          , S. (eds.) ICWSM. The AAAI Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Ragheb</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moulahi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aze</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bringay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Servajean</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Temporal mood variation: at the CLEF erisk-2018 tasks for early risk detection on the internet</article-title>
          .
          <source>In: Working Notes of CLEF 2018 - Conference and Labs of the Evaluation Forum</source>
          , Avignon, France,
          <source>September 10-14</source>
          ,
          <year>2018</year>
          . (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Simpson</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Psorakis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Dynamic bayesian combination of multiple imperfect classi ers</article-title>
          .
          <source>In: Decision Making and Imperfection</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>35</fpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Trotzek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koitka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Linguistic metadata augmented classi ers at the clef 2017 task for early detection of depression</article-title>
          .
          <source>In: Working Notes of CLEF 2017 - Conference and Labs of the Evaluation Forum</source>
          . vol.
          <source>CEUR-WS</source>
          <year>1866</year>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keskar</surname>
            ,
            <given-names>N.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiong</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
          </string-name>
          , R.:
          <article-title>Identifying generalization properties in neural networks</article-title>
          .
          <source>In: International Conference on Learning Representations</source>
          (
          <year>2019</year>
          ), https://openreview.net/forum?id=BJxOHs0cKm
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>