<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Data-Driven Approach for Measuring the Severity of the Signs of Depression using Reddit Posts</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Paul van Rijen</string-name>
          <email>paul.vanrijen@hesge.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Douglas Teodoro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nona Naderi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luc Mottin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julien Knafou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matt Jeffryes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Ruch</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BiTeM group, HES-SO / HEG Geneva, Information Sciences</institution>
          ,
          <addr-line>Geneva</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>SIB Text Mining, Swiss Institute of Bioinformatics</institution>
          ,
          <addr-line>Geneva</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Toronto</institution>
          ,
          <addr-line>Toronto</addr-line>
          ,
          <country>Canada contact:</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In response to the CLEF eRisk 2019 shared task on measuring the severity of the signs of depression from threads of user submissions on social media, our team has developed a data-driven, ensemble model approach. Our system leverages word polarities, token extraction via mutual information, keyword expansion and semantic similarities for classifying Reddit posts according to the Beck's Depression Inventory (BDI). Individual models were combined at the post level by majority voting. The approach achieved a baseline performance for the assessed metrics, including Average Hit Rate and Depression Category Hit Rate, being equivalent to the median system in the limit of one standard deviation.</p>
      </abstract>
      <kwd-group>
        <kwd>Depression severity assessment</kwd>
        <kwd>Social networks</kwd>
        <kwd>Natural language processing</kwd>
        <kwd>Machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Depression is increasingly recognized as a major burden in public healthcare worldwide
[
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. In 2015 the World Health Organization (WHO) estimated that the total number
of people living with depression was 322 million [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Depression is ranked as the single
largest contributing factor to non-fatal health loss worldwide [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Major depressive
disorder is associated with increased morbidity, disability and costs, increased mortality
due to other co-occurring medical conditions including cardiovascular and pulmonary
diseases and is a leading cause of suicide [
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2–4</xref>
        ]. In addition to the high burden of
disease, the majority of patients (50% globally) do not receive appropriate care [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
Barriers to proper diagnoses and treatment include social stigmas and a low detection rate in
primary care [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Accurate and early detection of depression can help to lower these
barriers and thus mitigate the associated health risks.
      </p>
      <p>
        Social media networks, such as Facebook, Twitter and Reddit, enable people to share
their opinions and sentiments about a wide range of topics online [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. In recent years
various studies have explored the potential of data from social media networks for
detecting signs of depression [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. In addition, the scientific community has put forward
various shared tasks such as CLPsych [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and CLEF eRisk [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. In CLEF eRisk
2018, the objective was to predict whether a user was depressed or not given a set of
posts in a chronological order. Trotzek et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] achieved the top F1-score using a bag
of words ensemble method.
      </p>
      <p>
        For 2019, the CLEF eRisk includes a task aimed towards measuring the severity of
the signs of depression from threads of user submissions on social media [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. The
eRisk task 3 involves filling in a Beck's Depression Inventory (BDI) questionnaire [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ],
which assesses the presence of feelings like sadness, pessimism, loss of energy, etc., in
an individual, using a set of social media posts. Hence, the task changed from a standard
classification task, as in tasks Early Detection of Signs of Anorexia (task 1) and
Selfharm (task 2) of CLEF eRisk 2019 [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], to a combination of an information retrieval
and an interactive dialogue task, where the system should simulate how a user would
answer/fill in the questionnaire [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. In response to this challenge our team developed
a data-driven, multi-model approach based on word-polarities, mutual information and
semantic similarities.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <sec id="sec-2-1">
        <title>Beck's Depression Inventory</title>
        <p>
          The BDI questionnaire has 21 questions in the following categories: sadness,
pessimism, past failure, loss of pleasure, guilty feelings, punishment feelings, self-dislike,
self-criticalness, suicidal thoughts or wishes, crying, agitation, loss of Interest,
indecisiveness, worthlessness, loss of energy, changes in sleeping pattern, irritability, changes
in appetite, concentration difficulty, tiredness or fatigue, and loss of interest in sex. The
answers vary in a [
          <xref ref-type="bibr" rid="ref1 ref2 ref3">0-3</xref>
          ] scale, where 0 means the absence of the feeling and 1 to 3 the
presence from a milder (1) to a stronger (3) form.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Task data</title>
        <p>As shown in Table 1, the dataset for eRisk 2019 consists of Reddit posts from 20 users
and contains the user identifier, the post timestamp, title and post content. The data was
annotated at the user level with the depression severity according to Beck’s Depression
Inventory. The dataset was shared with the participants without the labels for the system
development and was also used for the model evaluation during the test phase.</p>
      </sec>
      <sec id="sec-2-3">
        <title>BDI questionnaire answering models</title>
        <p>In this section, we describe the models used to automatically fill in the BDI
questionnaire using the user’s Reddit posts.</p>
        <p>Model 1 - Word polarity. In this model, we aim to leverage word polarities for first
classifying Reddit posts as depressive and next associate posts to relevant BDI
dimensions.</p>
        <p>
          Resources. For this model, we made use of the Multi Perspective Question Answering
(MPQA) subjectivity lexicon [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ][
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] of over 8000 cues that can be used to express
private states including emotions, evaluations and stances. In this lexicon, cues are
annotated with positive, negative, or neutral polarity. In addition, this lexicon provides
information regarding the part-of-speech of the cues and whether they are stemmed or
not. In our model, we only considered single-word cues to determine whether a post is
depressive or not.
        </p>
        <p>
          For analyzing the posts with the BDI dimensions, we created a lexicon that provides
cues for each dimension by first randomly selecting three subjects' writings
(subject2341, subject5897 and subject9694). The MPQA single-word cues that appeared in
these writings were used as cues for the BDI dimensions lexicon. Next we expanded
the list of cues with the following resources: WordNet [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] to find synonyms, a sexual
desires vocabulary [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and the F.E.A.S.T.'s Eating Disorders Glossary [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. The
annotation process of assigning BDI dimensions to each of the cues was done by three
team members. In the final version of lexicon, we included only the annotations that
were agreed upon by at least two annotators. Some cues can be associated with multiple
BDI dimensions. For instance, the cue ‘hate’ was associated to both the ‘agitation’ and
‘self-dislike’ dimensions, as illustrated by the following examples: Post 1: ‘Hate when
people do that.’, Post 2: ‘My life is already disintegrating, and I hate my grades.’. The
final lexicon contained 668 words in total for all BDI dimensions. On average, the
lexicon included 30 terms for each dimension. The majority of cues (583) were annotated
with only a single dimension.
        </p>
        <p>
          Classifier. First, we tagged the words in each post according to their polarity using the
MPQA subjectivity lexicon. Since no training data was available during the official
phase, we empirically set a threshold of 0.1 for the ratio of negative to positive words
for classifying the posts as depressive. Only considering the depressive posts, we then
tagged words with BDI dimensions using the developed lexicon. Finally, we calculated
questionnaire responses by normalizing the tag counts for each BDI dimension into a
[
          <xref ref-type="bibr" rid="ref1 ref2 ref3">0-3</xref>
          ] score.
        </p>
        <p>
          Model 2 - Mutual information. In this model, we attempt to create a training dataset
from Reddit to classify posts as depressive or not. We used the mutual information
measure to extract relevant tokens from depressive posts [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. Kraskov et al. proposes
a model to estimate the mutual information M(X,Y) from samples of random points
distributed according to some joint probability density µ(x,y) based on entropy
estimates from k-nearest neighbor distances.
        </p>
        <p>Data. Two subreddit collections, containing 107,129 posts, were extracted as
candidates for providing positive and negative depression tokens. The positive collection
included 12 mental health related subreddits, such as Anxiety, depression,
eating_disorders, self-harm, social anxiety, and SuicideWatch. The negative collection included
32 general subreddits, such as all, AskReddit, explainlikeimfive, funny, movies, and
worldnews.</p>
        <p>Training collection. Each post of the positive and negative collections was tokenized,
stopword-removed, and stemmed, and unigram tokens were extracted and associated to
the respective subreddit. Using the mutual information criteria, the 200 most
informative tokens from each collection were used to tag each post. If a post from the positive
collection contained more positive tokens, it was deemed as positive. Similarly, if a
post from the negative collection contained more negative tokens, it was deemed as
negative. The training set was then created with the positive and negative posts tagged
with the 200 most informative tokens extracted from both collections. The final training
collection contains 3,318 positive and 58,328 negative posts.</p>
        <p>Classifier. A logistic regression classifier was trained to categorize posts into
depressive or not using the positive and negative posts. Then, keywords from the BDI
categories were expanded using WordNet and used to tag the positively classified posts.
This model did not take into account the nuances of the positive answers for a BDI
category, i.e., it considered the task as binary, assigning answers as 0 (negative) or 2
(positive) for a post.</p>
        <p>
          Model 3 - Semantic similarity. Word embeddings have shown to capture semantic
similarities and in recent years, various models have been proposed to generate these
embeddings, such as word2vec [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], GloVe [26], and BERT [27]. Here, we propose to
find the most semantically similar user posts to the questionnaire responses in order to
estimate how a user may respond to the questionnaire. Given word embeddings, we
generate the representation of each user post by averaging over the embeddings of
words in the post. We use a similar approach to represent the questionnaire response
vectors, i.e., we average the embeddings of words in each questionnaire response. We
will then compute the similarity of a user post and a questionnaire response using cosine
similarity. We use pre-trained GloVe word embeddings1 [26] (trained on 2 billion
tweets and has 200 dimensions) to represent the words. In order to filter out the
irrelevant posts, in the first step, we remove the posts that are not similar to the questionnaire
responses using only the noun and verb vectors and a threshold that was chosen
empirically (0.8). For the remaining posts, we compute the vector-based distance of each post
and questionnaire responses and choose the most similar response for that post.
Treating each post as the average of word embeddings does not consider word orders and
        </p>
        <p>https://nlp.stanford.edu/projects/glove/
not likely produce a good representation for the longer posts, but it has shown to provide
a relatively good baseline. This post-level representation can be improved by
leveraging state-of-the-art sentence embedding models [28].</p>
        <p>Model 4 – Ensemble. Two ensemble approaches were tested - micro and macro-voting
using models 1 to 3. In micro-voting, models were combined at the post level. If more
than two models classified a post as positive for depression, and if two (majority) or
three (strict) models classified the post as positive for a category, the category was
deemed as positive for that user. In macro voting, results were combined using the
average category prediction from the three models. The official run (BiTeM run 0) was
generated using the strict micro-voting ensemble.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Tuning individual models</title>
        <p>For each of the individual models we applied a threshold k, which determines the
minimal number of positive posts (i.e., categorized as depressed) the system would need to
consider a user as depressed. Only then, responses would be given to questions in the
21 BDI categories. Hence, a positive post could be considered as a proxy for a
depressive episode. As there was no training data for this task, we used an empirical k=5, that
is, five depressive episodes would be needed to regard a user as depressed. Then, for
each category, if multiple answers (0 to 4) were retrieved for a deemed positive user,
then the system assigned the response with the highest value for the category.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        System effectiveness metrics considered for this task are Average Hit Rate (AHR),
Average Closeness Rate (ACR), Average Difference between Overall Depression Levels
(ADODL) and Depression Category Hit Rate (DCHR). The AHR measures the ratio of
cases where the computed questionnaire produces exactly the same answer as the real
questionnaire. The ACR measures the averaged absolute distance on an ordinal scale
between the automated answer and the real answer. ADODL assesses the system’s
performance by first calculating the overall depression score (sum of all answers) and,
next, the absolute difference (ad_overall) between the automated score and the real
overall depression score. Depression levels are normalized as follows; DODL = (63
ad_overall) / 63. DCHR measures the fraction of cases in which the automated
questionnaire resulted in same depression severity categorization as the real questionnaire.
Table 2 shows the four depth-of-depression categories and the associated depression
levels used in this task[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
Our team submitted one official run to the Task 3 of the eRisk challenge. This run
combined the results of the three models described above using voting. The voting was
performed in a micro-average fashion, i.e., the results for each model were combined
at post level. Table 3 shows the results of our model. Overall, it has a baseline
performance for all the metrics, being equivalent to the median model of the participants
within the limit of one standard deviation.
After the official phase our team conducted some further experiments. In addition to
the micro and macro ensemble models, we evaluated the performance on the three
models separately. For each of these models we applied a k=5 threshold, i.e., if the model
classified at least 5 posts as positive for a category then the category was deemed as
positive for that user. Table 4 describes the results of the various models. Overall, both
ensemble micro models outperform the individual models and the ensemble macro
model. The ensemble micro majority model outperforms the model used in the official
phase on the ACR and ADODL metrics but at significant penalty in DCHR.
We repeated the unofficial experiments under the assumption that the questionnaire is
binary, i.e., the user expressed any feelings of depression or not. Table 5 shows the
model results under this assumption. Similar to the non-binary results, the ensemble
micro majority model outperforms the model used in the official phase but again at a
significant penalty in the DCHR metric.
      </p>
      <sec id="sec-3-1">
        <title>Impact of k on the individual model’s performance</title>
        <p>Fig 1. shows how the ADODL metric varies in function of k, i.e., the number of positive
posts necessary to confirm a category as positive. Models 1 and 2 presents their highest
ADODL around k=5 whereas model 3 in k=3. As there was no training set during the
official submission phase, we set these values empirically to k=5 for all the individual
models based on a manual analysis of some results. We suspected that k=1 would create
too many false positives. A similar pattern is also seen for the other metrics (not shown
here for brevity).
Social media can provide valuable resources for assessing individual’s mental health
that could be useful for early detection and consequent healthcare provision. We
developed a simple model for measuring the severity of the signs of depression from Reddit
posts based on word polarities, mutual information and semantic similarities. The
ensemble model used in the official phase achieved modest results. This could be
explained by the significant negative effect of weak individual models during the
construction of the ensemble model.</p>
        <p>Nevertheless, both micro ensemble models significantly improved upon the
individual models’ results for all the metrics apart from DCHR. Indeed, for the DCHR metric,
model 1 presented the best performance in the standard questionnaire answer, being
able to predict the correct depression severity category for 30% of the users. The
ensemble macro model did not improve upon the ensemble micro models. One possible
cause can be the relatively small set of candidate models from which the results were
averaged to calculate the ensemble category predictions.</p>
        <p>Answering the BDI questionnaire without training data proved to be a challenging
task. Indeed, even when considering the questionnaire as binary, the participant models
were outperformed by a naïve all-positive answer baseline on some of the metrics.
Model 2 performed the best in the DCHR and correctly predicted depression in 45% of
the cases when considering the questionnaire as binary. This remarkable improvement
over the 20% performance in DCHR in the standard questionnaire could be explained
by the fact that this model, in contrast to model 1 and 3, considered the task as binary
already in its conception, not taking the nuances in positive answers into account.</p>
        <p>Finally, as expected, training the models for some parameters would significantly
improve their performance. Indeed, most of the individual and ensemble model
parameters, such as cut-off, k, and voting weight, were set empirically during the official
phase and the results reported here do not try to tune them based on the gold standard
answers. As shown in Fig. 1, tuning only k, for example, would result on an average
improvement of up to 13% for the ADODL metric if we consider k=1 as the baseline.
This effect is also seen for the AHR, ACR and DCHR metrics, which can have an
average relative performance increase of up to 32%, 14% and 33%, respectively, with
tuning.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>The task T3 of CLEF eRisk 2019 aimed to measure the severity of the signs of
depression using user threads available in social media. The organizers provided a dataset
containing Reddit posts from 20 users and the goal was to automatically fill the 21
questions of Beck's Depression Inventory for each of the users. Our team developed a
data-driven, ensemble model combining sentiment lexicons, mutual information and
embedding similarities in order to overcome the lack of training samples. The model
achieved a baseline performance, being equivalent to the median system from the
overall challenge. Nevertheless, answering the BDI questionnaire without training data
showed to be a challenging task, with an average hit rate of less than 42% for the top 1
system (32% in our case). Indeed, for some metrics, our system was outperformed by
a naïve all-positive answer baseline in a binary classification. As next steps, we aim to
leverage the post evidences created during this task to improve the performance of our
classification model.
26. Pennington, J., Socher, R., Manning, C.: Glove: Global vectors for word representation. In:
Proceedings of the 2014 conference on empirical methods in natural language processing
(EMNLP). pp. 1532–1543 (2014).
27. Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: Bert: Pre-training of deep bidirectional
transformers for language understanding. arXiv preprint arXiv:1810.04805. (2018).
28. Ethayarajh, K.: Unsupervised random walk sentence embeddings: A strong but simple
baseline. In: Proceedings of The Third Workshop on Representation Learning for NLP. pp. 91–
100 (2018).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Marcus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yasamy</surname>
          </string-name>
          , M.T.,
          <string-name>
            <surname>Van Ommeren</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chisholm</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saxena</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Depression:
          <article-title>A global public health concern</article-title>
          .
          <source>World Health Organization Paper on Depression. 6-8</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. World Health Organization:
          <article-title>Depression and other common mental disorders: global health estimates</article-title>
          . World Health Organization (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Forte</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baldessarini</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tondo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vázquez</surname>
            ,
            <given-names>G.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pompili</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Girardi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Long-term morbidity in bipolar-I, bipolar-II, and unipolar major depressive disorders</article-title>
          .
          <source>Journal of Affective Disorders</source>
          .
          <volume>178</volume>
          ,
          <fpage>71</fpage>
          -
          <lpage>78</lpage>
          (
          <year>2015</year>
          ). https://doi.org/10.1016/j.jad.
          <year>2015</year>
          .
          <volume>02</volume>
          .011.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kessler</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berglund</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demler</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koretz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merikangas</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rush</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walters</surname>
            ,
            <given-names>E.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.S.:</given-names>
          </string-name>
          <article-title>The Epidemiology of Major Depressive Disorder: Results From the National Comorbidity Survey Replication (NCS-R)</article-title>
          .
          <source>JAMA</source>
          .
          <volume>289</volume>
          ,
          <fpage>3095</fpage>
          -
          <lpage>3105</lpage>
          (
          <year>2003</year>
          ). https://doi.org/10.1001/jama.289.23.3095.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bokhour</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dell</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osei-Bonsu</surname>
            ,
            <given-names>P.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glickman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisen</surname>
            ,
            <given-names>S.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elwy</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          :
          <article-title>Impact of Stigma on Veteran Treatment Seeking for Depression</article-title>
          .
          <source>American Journal of Psychiatric Rehabilitation</source>
          .
          <volume>17</volume>
          ,
          <fpage>128</fpage>
          -
          <lpage>146</lpage>
          (
          <year>2014</year>
          ). https://doi.org/10.1080/15487768.
          <year>2014</year>
          .
          <volume>903875</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Vermani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marcus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katzman</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Rates of Detection of Mood and Anxiety Disorders in Primary Care: A Descriptive, Cross-Sectional Study</article-title>
          .
          <source>Prim Care Companion CNS Disord</source>
          .
          <volume>13</volume>
          , (
          <year>2011</year>
          ). https://doi.org/10.4088/PCC.10m01013.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gramlich</surname>
          </string-name>
          , J.:
          <article-title>5 Facts about Americans and Facebook</article-title>
          . Pew Research Center.
          <volume>10</volume>
          , (
          <volume>5</volume>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Choudhury</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gamon</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Counts</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Horvitz</surname>
          </string-name>
          , E.:
          <article-title>Predicting Depression via Social Media</article-title>
          . In: Seventh
          <source>International AAAI Conference on Weblogs and Social Media</source>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Guntuku</surname>
            ,
            <given-names>S.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yaden</surname>
            ,
            <given-names>D.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kern</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ungar</surname>
            ,
            <given-names>L.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eichstaedt</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Detecting depression and mental illness on social media: an integrative review</article-title>
          .
          <source>Current Opinion in Behavioral Sciences. 18</source>
          ,
          <fpage>43</fpage>
          -
          <lpage>49</lpage>
          (
          <year>2017</year>
          ). https://doi.org/10.1016/j.cobeha.
          <year>2017</year>
          .
          <volume>07</volume>
          .005.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Coppersmith</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hollingshead</surname>
            , K., Mitchell,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>CLPsych 2015 shared task: Depression and PTSD on Twitter</article-title>
          .
          <source>In: Proceedings of the 2nd Workshop on Computational Linguistics and Clinical Psychology: From Linguistic Signal to Clinical Reality</source>
          . pp.
          <fpage>31</fpage>
          -
          <lpage>39</lpage>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
            ,
            <given-names>J.: eRISK</given-names>
          </string-name>
          <year>2017</year>
          :
          <article-title>CLEF lab on early risk prediction on the internet: experimental foundations</article-title>
          .
          <source>In: International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          . pp.
          <fpage>346</fpage>
          -
          <lpage>360</lpage>
          . Springer (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <article-title>Overview of eRisk: Early Risk Prediction on the Internet</article-title>
          . In:
          <article-title>International Conference of the Cross-Language Evaluation Forum for European Languages</article-title>
          . pp.
          <fpage>343</fpage>
          -
          <lpage>361</lpage>
          . Springer (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Trotzek</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koitka</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          :
          <article-title>Word Embeddings and Linguistic Metadata at the CLEF 2018 Tasks for Early Detection of Depression and Anorexia</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Early Detection of Risks on the Internet: An Exploratory Campaign</article-title>
          .
          <source>In: European Conference on Information Retrieval</source>
          . pp.
          <fpage>259</fpage>
          -
          <lpage>266</lpage>
          . Springer (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. CLEF eRisk:
          <article-title>Early risk prediction on the Internet | CLEF 2019 workshop</article-title>
          , https://early.irlab.org/.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>A.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steer</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carbin</surname>
            ,
            <given-names>M.G.</given-names>
          </string-name>
          :
          <article-title>Psychometric properties of the Beck Depression Inventory: Twenty-five years of evaluation</article-title>
          .
          <source>Clinical psychology review. 8</source>
          ,
          <fpage>77</fpage>
          -
          <lpage>100</lpage>
          (
          <year>1988</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Nona</surname>
            <given-names>Naderi</given-names>
          </string-name>
          , Julien Gobeill, Douglas Teodoro, Emilie Pasche, Patrick Ruch:
          <article-title>A Baseline Approach for Early Detection of Signs of Anorexia and Self-harm in Reddit Posts</article-title>
          .
          <source>In: Proceedings of the CLEF 2019 Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Sutcliffe</surname>
            ,
            <given-names>R.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peñas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovy</surname>
            ,
            <given-names>E.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forner</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigo</surname>
          </string-name>
          , Á.,
          <string-name>
            <surname>Forascu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benajiba</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osenova</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Overview of QA4MRE Main Task at CLEF 2013</article-title>
          . In: CLEF (Working Notes) (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Wilson,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Wiebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Hoffmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Recognizing contextual polarity in phrase-level sentiment analysis</article-title>
          .
          <source>In: Proceedings of Human Language Technology Conference and Conference on Empirical Methods in Natural Language Processing</source>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>20. MPQA Resources, http://mpqa.cs.pitt.edu/#subj_lexicon.</mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Fellbaum</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>WordNet: An electronic lexical database Cambridge</article-title>
          . MA: MIT Press.
          <article-title>(</article-title>
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <article-title>feeling sexual excitement or desire - synonyms and related words | Macmillan Dictionary</article-title>
          , https://www.macmillandictionary.com/thesaurus-category/british/feeling
          <article-title>-sexual-excitement-or-desire.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>23. Eating Disorders Glossary, http://glossary.feast-ed.org/.</mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Kraskov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stögbauer</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grassberger</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Estimating mutual information</article-title>
          .
          <source>Phys. Rev. E</source>
          .
          <volume>69</volume>
          ,
          <issue>066138</issue>
          (
          <year>2004</year>
          ). https://doi.org/10.1103/PhysRevE.69.066138.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Efficient Estimation of Word Representations in Vector Space</article-title>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>