<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Early Mental Health Risk Assessment through Writing Styles, Topics and Neural Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Diego Maupome</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maxime D. Armstrong</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Raouf Belbahar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Josselin Alezot</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rhon Balassiano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marc Queudot</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>n Moss</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Quebec in Montreal UQAM</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of the RELAI team in the eRisk 2020 tasks. The 2020 edition of eRisk proposed two tasks: (T1) Early assessment of risk of self-harm and (T2) Signs of depression in social media users. The second task focused on automatically lling a depression questionnaire given user writing history. The RELAI team participated in both tasks, and addressed them using topic modeling algorithms (LDA and Anchor), neural models with three di erent architectures (Deep Averaging Networks (DANs), Contextualizers, and Recurrent Neural Networks (RNNs)), and an approach based on writing styles. For the second task related to early detection of depression, the system based on LDA performed well according to all the evaluation metrics, and achieved the best performance among participants according to the Average Di erence between Overall Depression Levels (ADODL) with a score of 83.15%. Overall, the submitted systems achieved promising results, and suggest that evidence extracted from social media could be useful for early mental health risk assessment.</p>
      </abstract>
      <kwd-group>
        <kwd>Early Risk Detection</kwd>
        <kwd>Topic Modeling</kwd>
        <kwd>Neural Networks</kwd>
        <kwd>Mental Health Risk Assessment</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The global goal of the eRisk challenges is the early detection of at-risk people
from their textual production on social media, using Natural Language
Processing (NLP) techniques. In 2020, two di erent tasks were put forth: early detection
of signs of self-harm (T1), and measuring the severity of the signs of
depression (T2) using textual data from related Reddit subreddits1. These tasks are
follow-ups of tasks 2 and 3 from 2019, respectively. T1 consists in sequentially
processing writings from a set of social media users, and detecting signs of
selfharm [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], classifying users as at-risk or not. The goal is not only to perform this
classi cation but also to do it as early as possible, i.e., based on as few writings
per user as possible. T2 consists in automatically lling the Beck's Depression
Inventory (BDI) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for a set of users, based on a history of their postings on
social media. In this work, we describe the participation of the RELAI team
from University of Quebec in Montreal (UQAM) at the Conference and Labs
of the Evaluation Forum (CLEF) 2020 eRisk tasks for early detection of signs
of self-harm and depression [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. The article is organized as follows. Sections 2
and 3 describe the proposed approaches and the research background they rely
on, present the applied methodologies, experimental setup and results obtained
on T1 and T2 respectively. Each Section concludes with a discussion about the
results, and suggests possible future improvements.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Early Signs of Self-Harm (T1)</title>
      <p>
        Self-harm is thought to a ect about 12% of adolescents [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In 2014-15,
hospitalizations due to self-in icted injuries in Canada were thrice as numerous as
suicides [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. While it is co-morbid with other mental health disorders [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], its
peculiar characteristics and cyclical nature have caused non-suicidal self-injury
to be included as an independent disorder in the DSM-V [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Further, only a
small fraction of young people will seek professional help either before or after
engaging in self-harm (9-12%) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This highlights the potential of the use of
automatic means of detection on social media [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Such is the aim of the current
task. We give hereafter a brief description of the corpus, the metrics, as well as
our participation.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Task and Data</title>
        <p>
          As previously mentioned, this task was rst introduced in the previous iteration
of eRisk (2019). In 2020, the dataset (training and test) consists of users
exhibiting signs of self-harm and control users. For information regarding the labeling
process, we direct the reader to [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Table 1 presents some statistics about the
2020 dataset. The test set is markedly di erent from the training set both in
class proportions and user verbosity. Indeed, the ratio of positive subjects in the
test is roughly double that of the training set. In addition, both positive and
negative test users have fewer and shorter documents compared to their
training counterparts. One of the chief concerns of the task is the early detection of
positive subjects. Therefore, during the test stage, a REST server2 was set up
by the task organizers to iteratively release user writings item-by-item during a
limited period of time. Thereby, the participants had to send a GET request to
retrieve the writing of each user. After each request, the processing/prediction
pipeline runs and gives back to the server, via a POST request, the predictions
about each individual. After each release of writings, a decision had to be
emitted. Classifying a user as su ering from self-harm (decision: 1) was considered as
nal, while predicting the user not at risk (decision: 0) was open to updates in
the following rounds. In order to evaluate the performance of the systems, and
to explore ranking-based measures, the task organizers also asked participants
to provide an estimated score of the level of self-harm with the decision.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Evaluation Metrics</title>
        <p>Several metrics allow to evaluate the systems in T1: standard classi cation
metrics like precision, recall and F1 score as well as speci c, time-aware classi cation
metrics such as ERDE, latencyT P , speed and Flatency (i.e. latency-weighted F1).
ERDE - Early Risk Detection Error - is a metric designed for eRisk tasks,
taking into account the correctness of predictions and the delay taken by the system
to make these predictions. The delay in decision for a given user is de ned by k,
the number of posts processed by the system before making a decision.
latencyT P takes into account the latency for true positive predictions only
because they represent users needing early intervention, as opposite to true negative
predictions. This measure is based on the median number of posts the system
has to analyze to detect true positives.</p>
        <p>
          The last two metrics are speed and Flatency [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Computing the speed requires a
penalty factor, which takes into account the number of a user's writings needed
by the system to make a decision. Flatency is a latency-weighted F1 score, which
combines the e ectiveness of the system with a delay, by multiplying the F1
score by the speed metric measuring the delay of the system.
        </p>
        <p>Since the 2019 edition, the organizers have added a ranking-based
evaluation process, which is not based on the binary predictions made but rather on
their associated score. Two metrics are proposed for evaluating this ranking:
the precision at k (P @k - percentage of true positive users among the k users
predicted by the system as presenting the highest risk); and the Normalized
Discounted Cumulative Gain (N DCG - evaluates a system based on the relevance
of the rankings built from its results). These metrics are computed after seeing
k writings; this year, they were reported with four di erent k values: 1, 100, 500,
1000.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Related Work</title>
        <p>
          In 2019, the best precision, F1 and ERDE [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] were obtained by a system based on
supervised learning for text classi cation called Sequential S3 - SS3 for
Smoothness, Signi cance, and Sanction - [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] submitted by the UNSL team. As for the
2 https://early.irlab.org/server.html
other evaluation metrics such as recall, latencyT P and speed, the system
submitted by LTL-INAOE achieved the best results with an approach based on the
similarity between a given piece of text and a set of phrases potentially related to
self-harm. Table 2 reports the best results obtained by the participating teams
of eRisk 2019 T2. In the next Sections, some details are given about how we
approached the problem of self-harm detection.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Topic Models</title>
        <p>
          Topics discussed by users of social media could provide insight into their
mental status [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. We chose to explore this hypothesis by using Latent Dirichlet
Allocation (LDA) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], the most widely known topic modeling algorithm, as well
as the Anchor variant [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. For both topic modeling approaches, the best results
were obtained when training the model using the entire textual production of
each user as a single document, concatenating the posts together.
Latent Dirichlet Allocation. Two di erent LDA models were tested, one
based on word stems and one based on word bigrams. Both models operate on
documents with stop-words and short words (3 characters or fewer) removed. The
rst model further stems the remaining words. The second LDA model is trained
instead on word bigrams. Once the LDA model is trained, users are mapped
to a vector of topics. Finally, a logistic classi er is trained on these vectors.
We use di erent training-validation splits to tune the hyper-parameters, namely
the number of topics extracted. The best results are obtained by the stemmed
unigram model, using a total of 14 topics.
        </p>
        <p>Anchor Variant. This method tries to nd a set of anchor words for each
topic discovered. The anchor words will be assigned high probability in only one
topic. The implementation of the Anchor-based system is similar to the LDA
models previously discussed, using stop-word removal and stemming. Further,
tokens used by over 60% of users are disregarded. As with the standard LDA
approach, we nd our best results in validation using 14 topics.
2.5</p>
      </sec>
      <sec id="sec-2-5">
        <title>Neural Encoders</title>
        <p>
          One of the principal challenges of the task is to combine the analyses of a user's
writings in order to arrive at a single prediction for said user. In this respect,
the exibility a orded by the back-propagation framework allowed us to explore
several manners in which to structure prediction models. Broadly speaking, we
distinguish two modes of aggregation encoding the documents making up a user
into a single user encoding. The rst mode, nested aggregation, uses two
encoders. The rst encoder encodes documents independently of each other. The
second encoder aggregates these encoded documents together. The second mode
of aggregation, at aggregation, uses a single encoder combining the words from
all documents simultaneously. We explored such aggregations with three di erent
architectures as encoders: Deep Averaging Networks (DANs), Contextualizers,
and Recurrent Neural Networks (RNNs). Contrary to the other two models,
DANs [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] cannot account for the position of items so we only used nested
aggregation with these models, having one DAN encode each of a user's document
independently, and a separate DAN aggregate these encoded documents. In the
case of Contextualizers, a at aggregation is more interesting, as even small
parts of documents can be put into the context of other passages so we opted
for a positional encoding consisting of a concatenation of three vectors of
sinusoids [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] for each word: one of them corresponding to the position of the word in
the document and the two others corresponding to the position of the document.
Rather than simply providing the position of a document in the user's history, we
provide a position with two components: one counting the units of time elapsed
since the writing of the post and a second one enumerates documents happening
within the same unit of time (a day in our case). As for RNNs, we borrow the
inter-document attention RNN approach described by [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] as the conditions are
very similar.
2.6
        </p>
      </sec>
      <sec id="sec-2-6">
        <title>Results and Discussion</title>
        <p>Results of the RELAI systems on the test set are presented in Table 3 for the
decision-based evaluation and Table 4 for the ranking-based evaluation. Overall,
the proposed models appear to have erred on the side of caution, achieving high
recall but relatively low precision.</p>
        <p>
          In terms of ranking-based evaluation, precision increases from the
measurement at 100 writings and remains high throughout for all models, suggesting
issues with the policies mapping scores to decisions. The number of positive
subjects being over 100, P @10 is less indicative of classi cation viability. While
all models achieve high N DCG@10, neural models especially, N DCG@100
remains modest throughout. Contrasting this with the high recall and low precision
achieved further illustrates the need for policy adjustments.
Given a user's history writing and based on evidence found in it, T2 participants
had to ll a standard depression questionnaire de ned from Beck's Depression
Inventory (BDI) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The questionnaire is composed of 21 questions with 4
possible answers (from 0 to 3), except for questions 16 and 18, where there are seven
possible answers (0, 1a, 1b, 2a, 2b, 3a, 3b). The answers to each question
represent an ordinal scale, each one associated to an integer value. The sum total
of a subject's answers is considered their score. Additionally, these scores are
associated with the following categories: minimal depression (depression levels
0-9), mild depression (10-18), moderate depression (19-29) and severe
depression (30-63). In T2, the proposed systems had to estimate each user's response
to each individual question. The predictions are therefore much more complex
than those expected in T1. We give a brief description of the corpus, the metrics
as well as our participation.
The second task of the eRisk 2020 lab was introduced in 2019, with the goal
of exploring much ner-grained prediction of the severity of depression
symptoms [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. For this purpose, each subject was asked to ll the BDI questionnaire.
The systems submitted by participants then had to estimate every user's answer
to each question given writing history of users. In order to assess the correctness
of the responses provided by the participants, a number of metrics are used.
Some are concerned with obtaining the exact answers whereas others are
concerned with proximity in individual answers or overall BDI score. These metrics
are described in the following Section. While participants in the previous
iteration did not have training data at their disposal, 20 users were made available for
eRisk 2020 participants, with 70 more used for evaluation. Their distributions
according to the standard categorization used on BDI scores are shown in Table
5. As in T1, the training and test set di er in this respect.
Four metrics evaluate the systems trying to address T2. The rst one is the
Average Hit Rate (AHR). For a given user, the hit rate is simply the number
of matches between the system automatic answers of the questionnaire and the
user answers, i.e., the rate of correct guesses over the total. The AHR is then
the mean hit rate across all users.
        </p>
        <p>The second metric is the Average Closeness Rate (ACR). The closeness rate is
a ner-grained measure of the disparity between the prediction and the ground
truth for each answer, as de ned by the ordinal scale on which the answers
are placed. To calculate this, for each question, one takes the system and user's
answers and computes the absolute di erences (ad) between them. The closeness
rate for a user is the mean closeness rate for each question, and the ACR is the
mean closeness rate across users.</p>
        <p>The third metric is the Average Di erence between Overall Depression Levels
(ADODL), which is the mean over all users of the Di erence between Overall
Depression Levels (DODL), i.e. the absolute di erence between the ground truth
and the system predictions.</p>
        <p>The last metric is the Depression Category Hit Rate (DCHR), which is the
fraction of the cases where the system score and the user's score fall in the same
category.
3.3</p>
      </sec>
      <sec id="sec-2-7">
        <title>Related Work</title>
        <p>
          In eRisk 2019, the highest AHR, was achieved by the SS3 system trained using
the dataset for the eRisk 2018 depression detection task [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Since the model
was designed as a "yes or no" classi er, the authors had to make some modi
cations to return a depression level between 0 and 63 to be able to ll a BDI
questionnaire. Additionally, a question-centered variant was built, achieving the
aforementioned AHR. The best ACR (distance-based variant) was also achieved
by a variant of the previous system using a probability distribution depending
on the value of the expected answer. In terms of ADODL and DCHR, the best
performances were reached with an unsupervised approach [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] using the distance
between the answers and all the sentences of a user's writing history. Table 6
reports the best results obtained by the participating teams of eRisk 2019
According to [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] these results show that it is possible to automatically
extract some depressions signals from social media activity. Although the
performance is still modest and far from a really e ective depression screening tool.
3.4
        </p>
      </sec>
      <sec id="sec-2-8">
        <title>Approaching the Task as One of Authorship Attribution</title>
        <p>The BDI was lled by only 20 users. Treating each of these users as one
observation to be mapped to the answers they gave to the questionnaire would lead to
a very limited number of examples. We hence approached the problem as one of
authorship attribution by two di erent methods, which rely on decision models
taking two documents and outputting the probability that both documents were
written by the same user. The proposed systems exploit the decision models in
di erent ways: one attempting to relate users to each other, and the other
attempting to relate a user to the text contained in the BDI questionnaire itself.
These methods are refered to as user-based and answer-based respectively.</p>
        <p>One key advantage of this authorship framework is that the training of
decision models does not require the annotation provided for the training set;
the models can be trained on unannotated data from the same domain. The
dataset from the eRisk 2018 depression risk detection task could hence be used
for the training of the authorship attribution models, using the 2020 training
data for validation. These models include LDA and a Contextualizer as well as a
stylometry-based approach. As for validation, in the user-based approach, some
users for whom the BDI is known were used as a knowledge base. The set of
users was therefore split in half, using one half as a knowledge base and the
other half for validation On the contrary, the answer-based approach allows to
validate on all 20 users.
3.5</p>
      </sec>
      <sec id="sec-2-9">
        <title>Topic Models</title>
        <p>One of the representation models used for this task was an LDA model. Using
topic modeling, the strategy is to create topic vectors for users and then measure
the distance between these in the user-based approach, or between these and
the topic representation of answers for the answer-based approach As previously
mentioned, the LDA model is trained on the eRisk 2018 depression risk detection
dataset. As in the Self-Harm task, each user's posts are grouped together into
larger documents. While the number of such groups will be the same for all
users, the choice of it is set by observing its e ect on validation results. The
pre-processing then involves of stop-words, short words (3 letters and fewer)
and stemming. Using the pre-processed documents, a dictionary and a
bag-ofwords are created to train the LDA model. A lter is applied when creating the
dictionary, removing words appearing in fewer than 20 documents or in over half
of the documents. We nd better results when requiring the model to nd 30
topics. The trained LDA model is then used to create vectors for the documents
from eRisk 2020 task 2 dataset. Each document is one of the reddit post included
in the dataset. Finally, the distance between every pair of document vectors is
measured using cosine similarity, which naturally falls in the unit interval, as the
topic vectors are strictly positive. For both approaches, we aimed to maximize
the ADODL metric. For the answer-based approach, the di erent experiments
show that the best ADODL is reached when combining each user's documents
into 19 groups, with an LDA model trained for 30 topics. The ADODL attained
by the user-based approach is approximately even when concatenating users'
posts into 10 to 19 groups. We opt to use 19 groups at test time as we posit this
will allow for ner-grained predictions.
3.6</p>
      </sec>
      <sec id="sec-2-10">
        <title>Contextualizer</title>
        <p>Contextualizer encoders were also used for this task. This time, the aggregation
considerations of the rst task were no longer relevant. We tested two di erent
approaches for the authorship decision task: encoding each document separately
(parallel) or together (simultaneous). Both encoders were trained for this
authorship task, ultimately using the depression questionnaire task as nal validation,
in both the user- and answer-based form. For the parallel encoder, the angular
similarity between the document vectors is used. The simultaneous encoder, on
the other hand, outputs the probability of the author being the same by design.</p>
        <p>In order to prevent over tting, we cease training of the authorship models by
monitoring their accuracy on unseen pairs of documents, including unseen pairs
of familiar documents, unseen documents by familiar users, as well as unseen
users. After extensive testing, we select the parallel encoder for the user-based
approach, and the simultaneous one for answer-based prediction.
3.7</p>
      </sec>
      <sec id="sec-2-11">
        <title>Stylometry</title>
        <p>
          This approach focuses on the writing style of a document in order to characterize
its author. To this end, several linguistic features served as document
representations, such as length of words and sentences, word and character frequencies
and word and sentence lengths. These features were largely inspired by
stylometric approaches to authorship attribution in instant messaging [
          <xref ref-type="bibr" rid="ref15 ref8">15, 8</xref>
          ] as well as
legal proceedings and lm reviews [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. They are presented in Table 7. As with
the LDA system, users' documents have been concatenated together, in order to
have the same number of documents per user while still accounting for all their
production. Features are normalized with respect to the length of these groups,
whether this length pertains to words, characters or sentences. These features
result in document representations of size 585. These vector representations are
then compared using cosine similarity. As previously mentioned, validation was
performed with the subjects for whom the BDI was available.
3.8
        </p>
      </sec>
      <sec id="sec-2-12">
        <title>Results and Discussion</title>
        <p>The results achieved on the test set are shown in Table 8. The more severe
metrics, the hit rates, were fairly low for all ve models. The user-based approach
produced superior results across metrics and authorship models. This is
unsurprising for LDA, where considerable parts of users' activity will likely di er in
subject matter from the BDI questionnaire. Given that the Contextualizer
encoder matches documents individually to answers, there might be gains in
performance to be obtained by considering the highest scoring document, rather than
the average, for each answer. Nevertheless, although the answer-based approach
was outperformed by the user-based one, it has the very appealing advantage of
not requiring annotated data, i.e. users with known BDIs.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>This paper has described the experiments performed by the RELAI team from
UQAM in the context of the eRisk 2020. Five models were submitted for each
of the two tasks.</p>
      <p>For the rst task related to early detection of self-harm, two topic modeling
systems were proposed, one using the standard LDA algorithm, and one relying
on its Anchor variant. The three remaining systems were based on neural
network, using three di erent architectures as encoders: Deep Averaging Networks
(DANs), Contextualizers, and Recurrent Neural Networks (RNNs). All models
are recall-oriented, which is arguably a safer decision policy. As evidenced by the
ranking-based evaluations, however, tweaking this policy could result in greater
precision. Globally, we achieved moderate results, the precision and recall
obtained leads to a F1-score between 0.439 and 0.550 which is decent comparing
to others systems. The Anchor model distinguished from our submitted models
by its rapidity to provide fast predictions with little content. This could be
explained by the presence of discriminative anchor words in provided user writings
which allow to predict rapidly if a user is at risk or not.</p>
      <p>For the second task related to early detection of depression, we approached
the problem as one of authorship attribution by two di erent methods:
userbased and answer-based. This approach a ords the freedom to build decision
models in a variety of ways. We relied again on LDA and the Contextualizer as
well as a stylometry-based approach, achieving the best result among
participants for ADODL (83.15%) with the LDA model with the user-based approach.
This metric is arguably the most relevant when it comes to overall assessment
of depression. Nonetheless, the ACR could be more interesting moving forward
as it pertains to informing a clinician on the exact symptoms a patient is
experiencing. Also, the LDA model shows a better balance between the di erent
metrics. Almost all of the other approaches submitted achieved higher results
than the average results for each metric. For example, the stylometric model
user-based approach performed the second best AHR. We could also note some
pertinent aspects. First, our systems are completely independent of the domain;
they make decisions only on extracted features from the provided texts without
requiring heavy processes of feature engineering or domain speci c hand-crafted
features. Also, the stark di erence in the proportions in terms of number of users
as well as the number of writings between the training set and the test set for
the two tasks could impact the performance of submitted models. Overall, the
test results show the promise of each approach. In future works, we will analyze
in more detail the results obtained for each task. We plan to incorporate more
carefully selected features to our decisions models which could grant a better
ability to identify users at risk. Finally, given the unique nature of T2, we will
explore di erent variations to improve predictions at a ner-grained level.
Reproducibility. The source code of the presented systems is available under
GNU GPL v3 licence to ensure reproducibility. It can be found in the following
repositories: https://gitlab.ikb.info.uqam.ca/ikb-lab/nlp/eRisk2020</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abed-Esfahani</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maslej</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goegan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>French</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Transfer Learning for Depression: Early Detection and Severity Prediction from Social Media Postings</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Arora</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halpern</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mimno</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moitra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sontag</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A practical algorithm for topic modeling with provable guarantees</article-title>
          .
          <source>In: International Conference on Machine Learning</source>
          . pp.
          <volume>280</volume>
          {
          <issue>288</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>A.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendelson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erbaugh</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>An Inventory for Measuring Depression</article-title>
          .
          <source>Archives of General Psychiatry</source>
          <volume>4</volume>
          (
          <issue>6</issue>
          ),
          <volume>561</volume>
          {
          <volume>571</volume>
          (06
          <year>1961</year>
          ). https://doi.org/10.1001/archpsyc.
          <year>1961</year>
          .
          <volume>01710120031004</volume>
          , https://doi.org/10.1001/archpsyc.
          <year>1961</year>
          .01710120031004
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research 3(Jan)</source>
          ,
          <volume>993</volume>
          {
          <fpage>1022</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Brown</surname>
          </string-name>
          , R.C.,
          <string-name>
            <surname>Plener</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          :
          <article-title>Non-suicidal self-injury in adolescence</article-title>
          .
          <source>Current psychiatry reports 19(3)</source>
          ,
          <volume>20</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Burdisso</surname>
            ,
            <given-names>S.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Errecalde</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , y Gomez,
          <string-name>
            <surname>M.M.:</surname>
          </string-name>
          <article-title>A text classi cation framework for simple and e ective early depression detection over social media streams</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>133</volume>
          , 182 {
          <fpage>197</fpage>
          (
          <year>2019</year>
          ). https://doi.org/https://doi.org/10.1016/j.eswa.
          <year>2019</year>
          .
          <volume>05</volume>
          .023, http://www.sciencedirect.com/science/article/pii/S0957417419303525
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Burdisso</surname>
            ,
            <given-names>S.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Errecalde</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montes</surname>
            y Gomez,
            <given-names>M.</given-names>
          </string-name>
          : UNSL at eRisk
          <year>2019</year>
          :
          <article-title>a unied approach for anorexia, self-harm and depression detection in social media</article-title>
          .
          <source>In: Working Notes of the Conference and Labs of the Evaluation Forum-CEUR Workshop Proceedings</source>
          . vol.
          <volume>2380</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Cristani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Ro o, G.,
          <string-name>
            <surname>Segalin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bazzani</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinciarelli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murino</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Conversationally-inspired stylometric features for authorship attribution in instant messaging</article-title>
          .
          <source>In: Proceedings of the 20th ACM international conference on Multimedia</source>
          . pp.
          <volume>1121</volume>
          {
          <issue>1124</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Doyle</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Treacy</surname>
            ,
            <given-names>M.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sheridan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Self-harm in young people: Prevalence, associated factors, and help-seeking in school-going adolescents</article-title>
          .
          <source>International journal of mental health nursing 24(6)</source>
          ,
          <volume>485</volume>
          {
          <fpage>494</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Iyyer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manjunatha</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boyd-Graber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daume</surname>
            <given-names>III</given-names>
          </string-name>
          , H.:
          <article-title>Deep unordered composition rivals syntactic methods for text classi cation</article-title>
          .
          <source>In: Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)</source>
          . pp.
          <volume>1681</volume>
          {
          <issue>1691</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <article-title>Overview of eRisk 2019 Early Risk Prediction on the Internet</article-title>
          . In:
          <article-title>International Conference of the Cross-Language Evaluation Forum for European Languages</article-title>
          . pp.
          <volume>340</volume>
          {
          <fpage>357</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <source>Overview of eRisk</source>
          <year>2020</year>
          :
          <article-title>Early Risk Prediction on the Internet</article-title>
          . In: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickho</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (eds) (ed.)
          <string-name>
            <surname>Experimental IR Meets Multilinguality</surname>
          </string-name>
          , Multimodality, and
          <source>Interaction Proceedings of the Eleventh International Conference of the CLEF Association (CLEF</source>
          <year>2020</year>
          ). Springer International Publishing (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Maupome</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meurs</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>Using Topic Extraction on Social Media Content for the Early Detection of Depression</article-title>
          . CLEF (Working Notes)
          <volume>2125</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Maupome</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Queudot</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meurs</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          :
          <article-title>Inter and intra document attention for depression risk assessment</article-title>
          .
          <source>In: Canadian Conference on Arti cial Intelligence</source>
          . pp.
          <volume>333</volume>
          {
          <fpage>341</fpage>
          . Springer (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mudit</surname>
            <given-names>Bhargava</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.M.</given-names>
            ,
            <surname>Asawa</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Stylometric Analysis for Authorship Attribution on Twitter</article-title>
          . In: Big Data Analyctics: Second International Conference. pp.
          <volume>37</volume>
          {
          <issue>47</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Ortega-Mendoza</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Far as</surname>
            ,
            <given-names>D.I.H.</given-names>
          </string-name>
          , Montes-y
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>LTL-INAOE's Participation at eRisk 2019: Detecting Anorexia in Social Media through Shared Personal Information</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sadeque</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bethard</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Measuring the latency of depression detection in social media</article-title>
          .
          <source>In: Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining</source>
          . pp.
          <volume>495</volume>
          {
          <issue>503</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Skinner</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McFaull</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Draca</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frechette</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaur</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pearson</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thompson</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Suicide and self-in icted injury hospitalizations in Canada (1979 to</article-title>
          <year>2014</year>
          /15).
          <article-title>Health promotion and chronic disease prevention in Canada: research, policy</article-title>
          and practice
          <volume>36</volume>
          (
          <issue>11</issue>
          ),
          <volume>243</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shazeer</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaiser</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polosukhin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Attention is all you need</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>5998</volume>
          {
          <issue>6008</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Xian</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vickers</surname>
            ,
            <given-names>S.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giordano</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>I.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramaswamy</surname>
          </string-name>
          , L.:
          <article-title># selfharm on Instagram: Quantitative Analysis and Classi cation of Non-Suicidal Self-Injury</article-title>
          .
          <source>In: 2019 IEEE First International Conference on Cognitive Machine Intelligence (CogMI)</source>
          . pp.
          <volume>61</volume>
          {
          <fpage>70</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Yunita</surname>
            <given-names>Sari</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Stevenson</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.V.</surname>
          </string-name>
          :
          <article-title>Topic or Style? Exploring the Most Useful Features for Authorship Attribution</article-title>
          .
          <source>In: Proceedings of the 27th International Conference on Computational Linguistics</source>
          . pp.
          <volume>343</volume>
          {
          <issue>353</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Zetterqvist</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The DSM-5 diagnosis of nonsuicidal self-injury disorder: a review of the empirical literature</article-title>
          .
          <source>Child and adolescent psychiatry and mental health 9(1)</source>
          ,
          <volume>1</volume>
          {
          <fpage>13</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>