<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detection of early sign of self-harm on Reddit using multi-level machine</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hojjat Bagherzadeh</string-name>
          <email>bagherzadehhosseinabad.hojjat@mail.um.ac.ir</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ehsan Fazl-Ersi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abedin Vahedian</string-name>
          <email>vahedian@um.ac.ir</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Computer Engineering, Ferdowsi University of Mashhad</institution>
          ,
          <country country="IR">Iran</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the participation of the EFE research team in task1 of CLEF eRisk 2020 competitions. This challenge basically focuses on the early detection of symptoms of self-harm from users' posts on social media. Identifying mental illnesses especially in the early stages can help people and avoid risky behaviors. Personal notes on social media are often indicative of one's psychological state, therefore using natural language processing techniques on users' posts one can develop an early risk detection system. The proposed method is basically consisting of Word2Vec representation, an ensemble of SVM and deep neural network and also attention layers. The obtained results are very competitive and show the strength of the system provided in the early diagnosis of selfharm.</p>
      </abstract>
      <kwd-group>
        <kwd>Early Risk Detection</kwd>
        <kwd>Self-Harm</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>SVM</kwd>
        <kwd>Attention</kwd>
        <kwd>Word2Vec</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Self-harm, also known as self-injury, is de ned as intentional bodily harm and
can a ect all people, regardless of age, gender and, race. It can be considered as a
common mental health issue, which can lead to several mental illnesses including
depression, anxiety and, emotional distress [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] . Previous ndings suggest that
people's narratives or writing patterns can re ect their mental state [
        <xref ref-type="bibr" rid="ref29 ref30">29, 30</xref>
        ]. So
with help of sentiment analysis, researchers try to identify mentally ill individuals
based on people's writings on the internet and this is the main objective behind
the CLEF eRisk challenges [
        <xref ref-type="bibr" rid="ref17 ref18 ref20">17, 18, 20</xref>
        ]. This year's challenge has two tasks. Task 1
deals with early detection of self-harm and task 2 tries to measure the severity of
the signs of depression. The EFE team participated in Task 1 of the competition
and that is given a sequence of writings for each user, the system attempts
to detect signs of self-harm in users as early as possible. User's writings are
processed in the same order in which they were sent and it lets chronologically
monitor user's activity.
      </p>
      <p>
        This paper presents the participation of the EFE team in self-harm early
detection challenges of CLEF 2020 [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. The method is an ensemble of deep
neural networks and Support Vector Machine (SVM) as the classi er of the
approach which takes its features from vector representation of the text of people
posts. These vectors are Word2Vec representation of cleaned text tweaked by
attention layers at di erent steps. Evaluation results of proposed method runs
and all other competition runs of the task 1 is discussed.
      </p>
      <p>The rest of the paper is organized as follows. Section 2 covers related work,
while section 3 gives a brief description of Task 1 of early risk detection and
the used datasets. And part 4 introduces the proposed method. Analysis of the
results of experiments are presented in section 5 and nally, the conclusion and
suggested directions for future works are presented in section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Early risk detection based on sentiment analysis is a trending research eld
having many growing applications. In recent years, given the popularity of social
media networks as a source for news and information such health-related datasets
have been made available which attracted substantial attention and led to the
introduction of online competitions such as CLEF [
        <xref ref-type="bibr" rid="ref1 ref17">1, 17</xref>
        ] , CLPsych Shared Task
[
        <xref ref-type="bibr" rid="ref23 ref8">8, 23, 36</xref>
        ]. Research suggests that individuals with mental issues can be identi ed
by what they publicly share on online social media platforms because of the
language patterns they use in their written texts. Thus, with advances in Natural
Language Processing (NLP) researchers can now provide tools that have the
capability of detecting mental illness in early stages.
      </p>
      <p>
        NLP modules basically have two steps. The rst step is a vector
representation of text such as term frequency-inverse document frequency (TF-IDF) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
pre-de ned patterns, or Part-Of-Speech tagging which needs expert views over
the context. Also, there are more generic text representations namely Word2Vec
[
        <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
        ], Doc2Vec [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and LDA (Latent Dirichlet Allocation) [
        <xref ref-type="bibr" rid="ref24 ref3">3, 24</xref>
        ] that are
based on counting all words of context. All of these try to represent the text
as a high-dimensional vector that is appropriate for machine-learning engines.
The second step is the learning process. SVM, neural networks and inference
models are some of many learning models which are being wildly used in the
NLP process.
      </p>
      <p>
        Dealing with the detection of mental illness, especially self-harm is a
challenging task as it is commonly relied on self-report. Most people who have self-harm
also su er from other mental illnesses, such as depression and anxiety, and thus
make it di cult to be distinguished [
        <xref ref-type="bibr" rid="ref10 ref13">13, 10</xref>
        ]. Wang et al. [34] were detecting
selfharm content on Flicker using word embedding and deep neural networks. UNSL
team [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], one of the participants in eRisk 2019, designed a special
dictionarybased text classi er. Bouarara and his team [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] analyzed users tweets to detect
suicidal or self-harm behaviors to prevent any risk attempt using a sentiment
classi cation model. Research ndings state that users' writing patterns can
express their mental state [
        <xref ref-type="bibr" rid="ref6 ref7">7, 6, 32</xref>
        ]. Furthermore, based on researches in this eld,
EFE team participated in in CLEF eRisk 2020 competition using a combination
of 2 deep neural networks and SVM models trained by Word2vec text
representation which promising results were obtained.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Dataset and Competition</title>
      <p>
        The training dataset of early self-harm detection as Task 1 of CLEF 2020 was
provided by the competition organizer and was the Task 2 of CLEF 2019
competition [
        <xref ref-type="bibr" rid="ref19 ref20">20, 19</xref>
        ]. The dataset consists of dated textual data of users' online posts
labeled as self-harm and non-self-harm. The labeling only determines the status
of each user and doesn't suggest any label for each writing. Table 3 shows a brief
statistic and summary of Task 1 train data [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        The training dataset includes whole writings of each user and a label
indicating the self-harm status of each one. The test stage though has an iterative
strategy and a new round of writing for each user is being released only after the
current run results are sent by the competitor. The evaluation measures being
used in this challenge other than precision, recall, and F1, is ERDE. A detailed
description of the tasks and evaluation metrics can be found in the corresponding
task description paper [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Proposed Method</title>
      <p>
        The proposed method for eRisk 2020 Task 1 is a multi-level approach which is
a combination of deep neural networks and SVM machine along with attention
mechanism, as shown in Fig. 1. At each round, rst, at word level, the text
cleaning steps including lowercasing, tokenizing, removing additional phrases
and stemming is done to obtain a cleaned version of submitted texts and then
vector representation of post using Word2Vec model [
        <xref ref-type="bibr" rid="ref15 ref25">15, 25</xref>
        ] is computed. Then,
at user level these representation of posts is fed to the rst level machines and
scores of indicating the level of self-harness for each post is achieved. Next, at
user level, these scores are aggregated to create user level features. Using an
attention mechanism and Chi-Squared feature selection technique [
        <xref ref-type="bibr" rid="ref25">25, 33</xref>
        ] the
most appropriate features are being selected as input to the user level learning
machines. Finally, values of user level SVM machines are making the nal
decision based on a scoring fusion function. Therefore, at each round based on the
user's writings from the beginning until this round a decision about considering
the user as a self-harm case is being made. Further details for each level are as
follows.
Word level or input layer is where create a numerical representation of words
and posts used by learning machines at the next level. First, the text of people's
post is tokenized and cleaned using NLP tools which mostly involve converting
plurals, removing web address and hashtags, lowercasing, stemming and
lemmatizing [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Then these clean words are fed to the word embedder which in
this experiment is a customized Word2Vec model with 100-dimension vector
space that is trained on Twitter and Reddit posts. Word2Vec is two-layer
neural network that is trained to reconstruct or predict surrounding contexts of
words in a sentence [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and the inner layer weights of trained network are used
as numeric representation of words. Therefore, each post P is represented by
[ W 2V1; W 2V2; :::; W 2Vm]p and each W 2Vip is a 100-dimension vector which is
Word2Vec representation of i-th word of p-th post.
In post level, the Convolutional Neural Network (CNN) [
        <xref ref-type="bibr" rid="ref14 ref16 ref27">27, 14, 16</xref>
        ] , Long
ShortTerm Memory (LSTM) [
        <xref ref-type="bibr" rid="ref12">35, 12</xref>
        ] and SVM [
        <xref ref-type="bibr" rid="ref28 ref31">31, 28</xref>
        ] are the main learning machines.
The one dimensional CNN and LSTM networks process words of each post based
on their chronological order and gives a score for each post which determines the
probability of belonging to the self-harm class. Each W 2Vip Word2Vec
representation of each words of post is fed directly to the CNN and LSTM networks.
Additionally, SVM is not being able to sequentially process words' posts and
thus the aggregated version of words representation which is a weighted average
of each words' post as shown in equation 1 is fed to the SVM machine. Then the
SVM like the other two neural machines gives a score to each post.
      </p>
      <p>Rp =
np
X (Wword)p;i
i=1</p>
      <p>W 2Vip
(1)</p>
      <p>Where Rp represents the numerical representation of p-th post and (Wword)p;i
is weight of i-th word of the p-th post which is computed in the training process
based on the importance of every word in the degree of positivity of each post.
And W 2Vip is the Word2Vec representation of i-th word of p-th post. Therefore,
the output of this level for each post of a user is a three-value vector in a way
that each learning machine produce one value.
4.3</p>
      <p>
        User level
At user level, posts' score of each user form a 3 n matrix which n is the number
of post sent by user up to the current round. In the training stage, by using
common statistical measures such as average, standard deviation and variance
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] 37 features are generated for each user and among these features, with the
help of Chi-squared feature selection method [
        <xref ref-type="bibr" rid="ref25">25, 33</xref>
        ] , the 7 most descriptive
ones are being selected and used as inputs in the SVM machine. In the test
stage, at each round, these 7 descriptive features in forms of 3 7 matrix are
created based on user writings from the beginning.
      </p>
      <p>Before applying statistical measures on scores of post level, scores are changed
by the attention mechanisms. Considering the fact that not all the posts sent
by a user is related to one's mental state, posts scores are weighted by their
correlation to self-harm category. Therefore, the score's posts are being weighted
and then used to generate the selected statistical features which are fed into the
nal SVM.</p>
      <p>At each round, the three-digit value output of the nal SVM is used by the
scoring fusion function which calculates a value indicating the level of self-harm
of the user based on one's writing from the beginning. These values are sent to
the nal decision system to alerts a `1' for users with self-harm mental status.</p>
      <p>The nal decision system works in a way that it stores scores of scoring
fusion function at each round and then decide to alert `1' for users whenever the
conditions depicted below are met.</p>
      <p>The average scores of user's up to this round.</p>
      <p>The number of ascending cases of user's scores.</p>
      <p>The number of values above the maximum threshold level.</p>
      <p>The average higher scores of user's up to this round.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Experimental setup and results</title>
      <p>The main parts of proposed system are shown in Fig. 1. EFE team have
participated in task 1 with three runs, each of which has small di erences.</p>
      <p>The rst run model is exactly as depicted in Fig. 1 and used eRisk 2018 task
1 &amp; 2 as the training dataset for optimizing the post level and used eRisk
2019 task 1 &amp; 2 as the training dataset for optimizing the user level and
con guring the attention mechanism.</p>
      <p>The second run con guration is as same as the rst run, and the only
difference is for the training datasets. eRisk 2018 depression task was used to
train the rst post level of the machine and the eRisk 2019 self-harm was
used to train the user level of the model.</p>
      <p>The third run model only has the SVM machines and the neural network
models of the system are omitted. The datasets for post and user levels are
the same as run 2 training sets.</p>
      <p>
        Evaluation metrics for this challenge were two groups that are fully explained
in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The rst group measures are precision, recall, and F1 metrics which
consider accuracy of the model on unbalanced datasets. And the second group which
is early risk detection error (ERDE), latencyT P and latency weighted F 1
consider accuracy of the model in presence of time.
      </p>
      <p>Table 5 shows the o cial results of the 3 runs of the proposed system. As
can be seen in the table, the second run has the best performance among the
others and that is because of choosing the self-harm dataset for training the user
level. However, the rst run has shown comparable results, which indicates the
connection between depression and self-harm and other mental illnesses.</p>
      <p>There are 56 runs of 12 teams participating in Task 1 of 2020 eRisk CLEF
challenge. Table 3 shows the statistic of participant results compare to the
proposed method. As shown in Table 5 the method has gained comparable results
in F1 and latency-weighted F1 measure and achieved 4th rank in F1 and 3rd
rand in latency-weighted F1 measure in almost the shortest time needed for
processing and completing the challenge. Because of the imbalance in the dataset,
the key to great performance is to maintain the balance between P and R, which
is achieved in run 2 of the proposed model. Given the fact that this was the rst
attempt participating in such competition, we paid a lot of attention to giving
early answers and as a result, the best outcomes of the model were not obtained.
This also explains why the latency-weighted F1 rank is the third, while the F1
rank is the forth
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and future work</title>
      <p>
        In this article using the presented model, EFE team participated in task 1 of
eRisk2020 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The task was to detect the sign of self-harm in users based on
their writings as early as possible. By engaging in this challenge, the capability of
social media's content as a potential source for applications related to health and
safety issues has been demonstrated. The proposed system is an ensemble
multilevel method based on SVM, CNN, and LSTM network, which are ne-tuned by
the attention layers.
      </p>
      <p>The test results show the positive e ects of using attention mechanisms in
the post layer and the user layer on the system, especially since not all posts sent
by a person re ect his or her mental state. Another main di culty is that there
is always a trade-o between early decision making and more precise decision
making. In this way, on the one hand, there is the need to detect the sign of
mental illness in the user as early as possible and on the other hand, the more
writings the system processes about the user, the more accurate the answer will
be.</p>
      <p>Finally, considering the fact that this is the rst attempt of EFE team at such
challenges, it's been found that a lot of work can be done to improve the system
for real situations. Future research direction in improving the model is by working
on better encoding text into numerical representation and also creating better
attention mechanisms at di erent levels of the system. Also, another research
interest is to nd an optimum, under which both accuracy and giving the fastest
answer are maintained.
32. Schwartz, H., Eichstaedt, J., Kern, M., Park, G., Sap, M., Stillwell, D.,
Kosinski, M., Ungar, L.: Towards Assessing Changes in Degree of Depression through
Facebook (jan 2014)
33. Sun, J., Zhang, X., Liao, D., Chang, V.: E cient method for feature selection in
text classi cation. In: Proceedings of 2017 International Conference on Engineering
and Technology, ICET 2017. vol. 2018-Janua, pp. 1{6. Institute of Electrical and
Electronics Engineers Inc. (mar 2018)
34. Wang, Y., Tang, J., Li, J., Li, B., Wan, Y., Mellina, C., O'Hare, N., Chang, Y.:
Understanding and Discovering Deliberate Self-Harm Content in Social Media. In:
Proceedings of the 26th International Conference on World Wide Web. pp. 93{
102. WWW '17, International World Wide Web Conferences Steering Committee,
Republic and Canton of Geneva, CHE (2017)
35. Wang, Y., Zhang, C., Zhao, B., Xi, X., Geng, L., Cui, C.: Sentiment Analysis of
Twitter Data Based on CNN. Shuju Caiji Yu Chuli/Journal of Data Acquisition
and Processing 33(5), 921{927 (sep 2018)
36. Zirikly, A., Resnik, P., Uzuner, .O., Hollingshead, K.: CLPsych 2019 Shared Task:
Predicting the Degree of Suicide Risk in Reddit Posts. Tech. rep. (2019)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. CLEF eRisk:
          <article-title>Early risk prediction on the Internet</article-title>
          , https://erisk.irlab.org/
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Beel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Langer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gipp</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>TF-IDuF: A Novel Term-Weighting Scheme for User Modeling based on Users' Personal Document Collections (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Latent Dirichlet Allocation</article-title>
          . In: Dietterich,
          <string-name>
            <given-names>T.G.</given-names>
            ,
            <surname>Becker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z</surname>
          </string-name>
          . (eds.)
          <source>Advances in Neural Information Processing Systems</source>
          <volume>14</volume>
          , pp.
          <volume>601</volume>
          {
          <fpage>608</fpage>
          . MIT Press (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bouarara</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          :
          <article-title>Detection and Prevention of Twitter Users with Suicidal SelfHarm Behavior</article-title>
          .
          <source>International Journal of Knowledge-Based Organizations</source>
          <volume>10</volume>
          (
          <issue>1</issue>
          ),
          <volume>49</volume>
          {61 (nov
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Burdisso</surname>
            ,
            <given-names>S.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Errecalde</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montes-Y-Gomez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : UNSL at eRisk
          <year>2019</year>
          :
          <article-title>a Unied Approach for Anorexia, Self-harm and Depression Detection in Social Media</article-title>
          .
          <source>Tech. rep. (</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Choudhury</surname>
          </string-name>
          , M.D.,
          <string-name>
            <surname>De</surname>
          </string-name>
          , S.:
          <article-title>Mental Health Discourse on reddit: Self-Disclosure, Social Support, and</article-title>
          <string-name>
            <surname>Anonymity. unde ned</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Coppersmith</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hollingshead</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>From ADHD to SAD: Analyzing the Language of Mental Health on Twitter through Self-Reported Diagnoses pp</article-title>
          .
          <volume>1</volume>
          {
          <issue>10</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Coppersmith</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hollingshead</surname>
            , K., Mitchell,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>CLPsych 2015 Shared Task: Depression and</article-title>
          PTSD on Twitter pp.
          <volume>31</volume>
          {
          <issue>39</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Edmondson</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brennan</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>House</surname>
            <given-names>"</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A.O.</surname>
          </string-name>
          :
          <article-title>"non-suicidal reasons for selfharm: A systematic review of self-reported accounts"</article-title>
          .
          <source>"Journal of A ective Disorders" "191"</source>
          , "
          <volume>109</volume>
          {
          <fpage>117</fpage>
          " (
          <article-title>"2016")</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Gratz</surname>
            ,
            <given-names>K.L.</given-names>
          </string-name>
          :
          <article-title>Risk factors for deliberate self-harm among female college students: The role and interaction of childhood maltreatment, emotional inexpressivity, and a ect intensity/reactivity</article-title>
          .
          <source>American Journal of Orthopsychiatry</source>
          <volume>76</volume>
          (
          <issue>2</issue>
          ),
          <volume>238</volume>
          {250 (apr
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Holosko</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thyer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Commonly Used Statistical Terms</article-title>
          .
          <source>In: Pocket Glossary for Commonly Used Research Terms</source>
          , pp.
          <volume>145</volume>
          {
          <fpage>156</fpage>
          . SAGE Publications, Inc. (jan
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Jianqiang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiaolin</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xuejun</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Deep Convolution Neural Networks for Twitter Sentiment Analysis</article-title>
          .
          <source>IEEE Access 6</source>
          ,
          <issue>23253</issue>
          {23260 (jan
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kairam</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaye</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guerra-Gomez</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shamma</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          :
          <article-title>Snap decisions? How users, content, and aesthetics interact to shape photo sharing behaviors</article-title>
          .
          <source>In: Conference on Human Factors in Computing Systems - Proceedings</source>
          . pp.
          <volume>113</volume>
          {
          <fpage>124</fpage>
          .
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery (may
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kshirsagar</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morris</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Detecting and
          <string-name>
            <given-names>Explaining</given-names>
            <surname>Crisis</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Distributed Representations of Sentences and Documents</article-title>
          .
          <source>31st International Conference on Machine Learning, ICML 2014 4</source>
          ,
          <issue>2931</issue>
          {2939 (may
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sato</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , Cheng,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          :
          <article-title>CNN for situations understanding based on sentiment analysis of twitter data</article-title>
          .
          <source>In: Procedia Computer Science</source>
          . vol.
          <volume>111</volume>
          , pp.
          <volume>376</volume>
          {
          <fpage>381</fpage>
          .
          <string-name>
            <surname>Elsevier</surname>
            <given-names>B.V.</given-names>
          </string-name>
          (jan
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
            ,
            <given-names>J.: eRISK</given-names>
          </string-name>
          <year>2017</year>
          :
          <article-title>CLEF lab on early risk prediction on the internet: Experimental foundations</article-title>
          .
          <source>In: Lecture Notes in Computer Science (including subseries Lecture Notes in Arti cial Intelligence and Lecture Notes in Bioinformatics)</source>
          . vol.
          <volume>10456</volume>
          LNCS, pp.
          <volume>346</volume>
          {
          <fpage>360</fpage>
          . Springer Verlag (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
            ,
            <given-names>J.: eRISK</given-names>
          </string-name>
          <year>2017</year>
          :
          <article-title>CLEF Lab on Early Risk Prediction on the Internet: Experimental Foundations</article-title>
          . In: Jones,
          <string-name>
            <given-names>G.J.F.</given-names>
            ,
            <surname>Lawless</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Cappellato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          , N. (eds.)
          <string-name>
            <surname>Experimental IR Meets Multilinguality</surname>
          </string-name>
          , Multimodality, and Interaction. pp.
          <volume>346</volume>
          {
          <fpage>360</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <article-title>Overview of eRisk 2019 Early Risk Prediction on the Internet</article-title>
          .
          <source>In: Lecture Notes in Computer Science (including subseries Lecture Notes in Arti cial Intelligence and Lecture Notes in Bioinformatics)</source>
          . vol.
          <volume>11696</volume>
          LNCS, pp.
          <volume>340</volume>
          {
          <fpage>357</fpage>
          . Springer Verlag (sep
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <source>Overview of eRisk at CLEF</source>
          <year>2019</year>
          :
          <article-title>Early Risk Prediction on the Internet</article-title>
          . In: Linda Cappellato Nicola Ferro,
          <string-name>
            <surname>D.E.L.H.M.</surname>
          </string-name>
          (ed.)
          <article-title>Conference and Labs of the Evaluation Forum. CEUR-WS.org (</article-title>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.: eRisk 2020:
          <article-title>Self-harm and depression challenges</article-title>
          .
          <source>In: Lecture Notes in Computer Science (including subseries Lecture Notes in Arti cial Intelligence and Lecture Notes in Bioinformatics)</source>
          . vol.
          <volume>12036</volume>
          LNCS, pp.
          <volume>557</volume>
          {
          <fpage>563</fpage>
          . Springer (apr
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Losada</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parapar</surname>
          </string-name>
          , J.:
          <source>Overview of eRisk</source>
          <year>2020</year>
          :
          <article-title>Early Risk Prediction on the Internet</article-title>
          . In: A.
          <string-name>
            <surname>Arampatzis E. Kanoulas</surname>
            ,
            <given-names>T.T.S.V.H.J.C.L.C.E.A.N.L.C.N.F.</given-names>
          </string-name>
          <year>e</year>
          . (ed.) Experimental Meets Multilinguality, Multimodality, and
          <source>Interaction Proceedings of the Eleventh International Conference of the CLEF Association</source>
          . Springer International Publishing (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Lynn</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goodman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niederho</surname>
            <given-names>er</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Loveys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Resnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Schwartz</surname>
          </string-name>
          , H.:
          <article-title>CLPsych 2018 Shared Task: Predicting Current and Future Psychological Health from Childhood Essays</article-title>
          . pp.
          <volume>37</volume>
          {
          <issue>46</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Maas</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Daly</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>P.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Potts</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Learning Word Vectors for Sentiment Analysis</article-title>
          .
          <source>Tech. rep. (</source>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          .
          <source>Advances in Neural Information Processing Systems</source>
          (oct
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yih</surname>
          </string-name>
          , W.T.,
          <string-name>
            <surname>Zweig</surname>
          </string-name>
          , G.:
          <article-title>Linguistic Regularities in Continuous Space Word Representations</article-title>
          .
          <source>Tech. rep. (</source>
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Mohan</surname>
          </string-name>
          , V.:
          <article-title>Preprocessing Techniques for Text Mining - An Overview (feb</article-title>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Monika</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deivalakshmi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Janet</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Sentiment Analysis of US Airlines Tweets Using LSTM/RNN</article-title>
          . In
          <source>: Proceedings of the 2019 IEEE 9th International Conference on Advanced Computing</source>
          ,
          <string-name>
            <surname>IACC</surname>
          </string-name>
          <year>2019</year>
          . pp.
          <volume>92</volume>
          {
          <fpage>95</fpage>
          .
          <article-title>Institute of Electrical and Electronics Engineers Inc</article-title>
          . (dec
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Moulahi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aze</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bringay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Dare to care: A context-aware framework to track suicidal ideation on social media</article-title>
          . pp.
          <volume>346</volume>
          {
          <issue>353</issue>
          (10
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Paul</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dredze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>You Are What Your Tweet: Analyzing Twitter for Public Health</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>38</volume>
          ,
          <fpage>265</fpage>
          {
          <volume>272</volume>
          (01
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Ragheb</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aze</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bringay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Servajean</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Attentive Multi-stage Learning for Early Risk Detection of Signs of Anorexia and Self-harm on Social Media</article-title>
          .
          <source>Tech. rep. (</source>
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>