<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Bucharest, Romania
" rlopes@ipb.pt (R. P. Lopes)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>CeDRI at eRisk 2021: A Naive Approach to Early Detection of Psychological Disorders in Social Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rui Pedro Lopes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Research Center for Digitalization and Intelligent Robotics (CeDRI), Instituto Politécnico de Bragança</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>This paper describes the participation of the CeDRI team in eRisk 2021 tasks, particularly, the Task 1: Early Detection of Signs of Pathological Gambling and Task 2: Early Detection of Signs of Self-Harm. The main diference between these two is that the first is a “test only” challenge, where no training data is supplied. The second task has labeled data available, which can be used for training. Both tasks were addressed using the same algorithms, using a custom training set for Task 1 and the provided data in the second. The algorithms were TfIdf vectorizer with a Logistic Regression layer, Word2Vec vectorizer with LSTM and Word2Vec vectorizer with CNN. All vectorizers and Neural Networks were trained solely with the training data. As expected, the algorithms did not state-of-the-art, but the experience allowed to reflect in several aspects related to the importance of proper dataset preparation and processing.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Early Risk Detection</kwd>
        <kwd>Tf-Idf</kwd>
        <kwd>Word2Vec</kwd>
        <kwd>Recursive Neural Networks</kwd>
        <kwd>Dataset Heuristics</kwd>
        <kwd>DL4J</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The term social network refers to a person’s connections to other people. In fact, creating and
maintaining social networks provide opportunities to connect with others who have similar
interests. Although initially applied in the context of “real-world” or physical, the concept
expanded to also include platforms that support online communication, such as Instagram,
Twitter or Reddit. Digital platforms further enhance these opportunities, allowing forming
relationships with people never met in person. Geographical barriers are attenuated or eliminated,
allowing to actively engage with people around the world. They can explore their curiosity,
pick up hobbies, or just spend time online. The possibility to write, participate or communicate
without restrictions also provides a means to unburden or receive emotional support. Some
people resort to social networks to talk about their state of mind, their feelings, distresses and
other problems.</p>
      <p>
        In opposition to verbal and direct communication, the content available in the social networks
is persistent, allowing asynchronous access data and providing a good means for psychological
and health related studies and analysis [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. According to several findings, people’s mental
state can be inferred from their social networks narratives [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. Based in this, the CLEF eRisk
challenges harness this opportunity to explore issues of evaluation methodologies, performance
metrics and other aspects related to building test collections and defining challenges for early
risk detection [
        <xref ref-type="bibr" rid="ref6 ref7 ref8 ref9">6, 7, 8, 9</xref>
        ].
      </p>
      <p>This year’s challenge has three tasks. Task 1, on early risk detection of pathological gambling,
and Task 2, on early risk detection of self-harm, consist of sequentially processing pieces of
evidence and detect early traces of pathological gambling and self-harm , respectively, as soon
as possible. Task 3, measuring the severity of the signs of depression, consists of estimating the
level of depression from a thread of user submissions. The CeDRI team participated in Task
1 and Task 2, where users’ posts are processed in the same order in which they are sent, to
chronologically monitor the users’ activity.</p>
      <p>This paper presents the participation of the CeDRI team in the pathological gambling and
in the self-harm early detection challenges of CLEF 2021. In task 1, two runs where executed,
using a Long-short Term Memory (LSTM) and Convolutional Neural Network (CNN) deep
neural networks, both with Word2Vec embeddings. Task 2 used three runs, with LSTM, CNN
with Word2Vec embeddings, like the previous task, and a logistic regression layer with Tf-Idf
vectorizer. Although the results were very close within the runs, the best results in Task 1 was
latency-weighted F1=0.141 (with the LSTM) and in Task 2 latency-weighted F1=0.206 (with the
CNN).</p>
      <p>The rest of the paper is organized as follows. Section 2 covers the considerations regarding the
datasets, while section 3 introduces the proposed method. Analysis of the results of experiments
are presented in section 4 and finally, the conclusion and suggested directions for future works
are presented in section 5.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Dataset</title>
      <p>The machine learning area is characterized by three main approaches of learning [10]:
• supervised - maps an input to an output based on example input-output pairs;
• unsupervised - patterns are learned without any explicit feedback;
• reinforcement - learns from a series of reinforcements, such as rewards and punishments.</p>
      <p>These are applied in several areas and with several purposes, such as classification, prediction,
estimation, afinity grouping, clustering and profiling. The eRisk challenge Task 1 and 2 is
mainly a classification problem, widely approached with supervised learning methods. In these
problems, a learning agent is shown what to do through an annotated set of training examples,
and it is expect an automated learning algorithm to generalize from these examples.</p>
      <p>For this, it is fundamental to understand and make sure that the training data is adequate
and it is well labeled.</p>
      <sec id="sec-2-1">
        <title>2.1. Text pre-processing</title>
        <p>Social networks’ posts often include tokens that do not represent words, such as URLs, HTML
entities, users’ handles, or others. Some of these do not bring relevant information to infer
the psychological condition of the user and may afect the performance of classification. The
pre-processing applied in both tasks included the following operations:
• unescape html entities (ex: &amp;lt; or &amp;#60;)
• remove handles (@abcd @pqrs)
• remove URLs (https://erisk.irlab.org)
• normalize lengthening (111111 -&gt; 11; kkkkkkkkkkk -&gt; kk)
• remove numbers
• convert to lowercase (Tomorrow -&gt; tomorrow)
• strip punctuation
• tokenize
• perform stemming</p>
        <p>The vocabulary is substantially reduced, as well as the word variations (Table 1). The same
pre-processing approach was applied in both tasks (sections 2.2 and 2.3).
We will be having our next meeting this evening at
5:00pm EST (9:00pm GMT). Meetings are 1 hour.
Participants must use Skype audio and video. If you’d like to
join, [DM me](http://www.reddit.com/message/compose/?to=
Jef W55&amp;amp;subject=ProblemGamblingSupportGroup ) with
your Skype name so you can be added to the call. Thanks. Jef
→</p>
        <p>Pre-processed text
[next, meet, even, pm, pm,
gmt, meet, hour, particip,
must, skype, audio, video,
you’d, like, join, me, gambl,
support, group, skype, name,
ad, call, thank, jef]</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Task 1: pathological gambling dataset</title>
        <p>The challenge consists of sequentially processing pieces of evidence and detect early traces
of pathological gambling signs in texts written in Social Media. This was an “only test” task,
so no training data was provided. The test collection format is a collection of writings (posts
or comments) from a set of Social Media users, labeling two categories of users, pathological
gamblers and non-pathological gamblers, and, for each user, the collection contains a sequence
of writings (in chronological order) [11].</p>
        <p>Since the challenge did not provide labeled data, a custom dataset, based on Reddit, was
built. For that, the Python Pushshift.io API Wrapper (PSAW - https://github.com/dmarx/
psaw) was used to retrieve posts from the Pushshift initiative (https://pushshift.io), in Comma
Separated Values (CSV) format. This allowed to remove the limit of 1000 posts that could be
downloaded from Reddit directly. The dataset was built based on the r/GamblingAddiction
and r/problemgambling communities. In addition, a random set of posts was also downloaded
to complement the dataset with non-gambling related content (Table 2).</p>
        <p>There is a considerable number of posts available after downloading, in a total of 73064
referring gambling issues and 47103 posts of random subjects. However, extracting data from
the CSV files failed in many posts, having only 7079 posts and 2306, respectively. This was due
to incompatibility issues between the post text and the CSV encoding, related to the appearance
of commas (‘,’) in the text and unterminated ‘"’, which made the issue of extracting the columns
very dificult and error sensitive. Because of balancing issues, the dataset was build with 2306
posts labeled with False and 2306 posts with True.</p>
        <p>Each post was stored in a single file, prefixed with pos or neg followed by a number (e.g.
pos_1762.txt, neg_2032.txt). It was decided not to associate or track the users, so each
post is individual and not related to any other.</p>
        <p>After building the dataset, the most frequent tokens in the gambling related posts (1a) and in
the non-gambling related posts (1b) were calculated (Figure 1). As expected, tokens like gambi,
monei, or stop appear in the vocabulary for gambling posts. For random, like, know and
would are very frequent.</p>
        <p>(a) gambling related posts.
(b) Non-gambling related posts.</p>
        <p>Next, the same operation was performed for bi-grams, to better understand the context of
the words (Figure 2).</p>
        <p>In these, feel like is transversal to both types of posts, although credit card and
gambli addict, for example, are clearly indicating the type of posts.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Task 2: self-harm dataset</title>
        <p>The training data provided XML files for 340 subjects, 41 of which belonging to the self-harm
group, 299 to the control group (Table 3). The total number of writings in the self-harm group
is 7,192 posts in contrast to 163,506 in the control group. The diference between the groups
is also very significant in the average number of writings per subject: 175.4 in the self-harm
group and 546.8 in the control group. The average length of the users’ (subjects’) writings is
(a) Gambling related posts.
179.1 and 129.2 respectively, and the number of tokens is 15.12 and 10.6. Although the control
subjects write more posts, they are, in average, shorter. The dataset is also provided with the
test writing, in the same format. They are also present in table 3, for completeness.</p>
        <p>In addition, not all posts are of the same language. Using OpenNLP’s language detection
model, a total of 81 diferent languages were counted. Table 4 show the 15 more frequent
languages within the writings.</p>
        <p>The dataset uses binary labels on the subjects, as having (positive) and not-having (negative)
self-harm (ground truth). As seen in table 3, each subject has an arbitrary number of posts,
and it is not expected that all of them will be strictly related to whether an user self-harms or
not. The main approach in this work, is to use a machine learning approach that uses text to
predict whether a message belongs to a positive or negative user, so the classifier should not be
trained with just the ground truth. Some selection on the posts have to be made, so that only
the self-harm related writings are kept as positive samples in the training set.</p>
        <p>Based on Non-Suicidal Self-Injury (NSSI) words [12], a selection was made on the posts to
extract individual writings to be used as positive examples. The examples were written in
two directories (pos/ for positive and neg/ for negative) with the following name schema:
subject280_2.txt, where the first number is the subject number and the second is this
subject’s post number. After selecting writings based on NSSI words, and excluding all languages
except English, a total of 391 positive labeled writings remained. For balance, the same number
of negative labeled writings were selected.</p>
        <p>The tokens frequency were also extracted from both the positive and the negative writings.
In this case, the bi-grams (Figure 3) and tri-grams (Figure 4) are presented.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed methods</title>
      <p>This section presents the models and experiments conducted for the eRisk 2021 task. First, the
classification methods require that text be converted to vectors.</p>
      <sec id="sec-3-1">
        <title>3.1. Vectorizers</title>
        <p>• TfIdf:
All the methods rely on the vectorization of the subjects’ writings. Two vectorizers were trained,
based on TfIdf and Word2Vec, both with the same text pre-processing techniques (section 2.1).
– minimum word frequency = 2;
(a) Self-harm related posts.
(b) Non-self-harm related posts.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Classifiers</title>
        <p>
          Three classification models were build for the tasks. The first is a simple Logistic Regression
layer using the TfIdf vectorizer, used in both Tasks 1 and 2:
• output dimension = 2;
• weight initialization algorithm = XAVIER;
2:
2:
• activation = SOFTMAX;
• optimization algorithm = STOCHASTIC_GRADIENT_DESCENT;
• updater = Nesterovs(0.1, 0.9)
• batch size = 32;
Another classifier was built using a CNN with Word2Vec vectors as input, used only in Task
• weight initialization algorithm = RELU;
• activation = LEAKYRELU;
• updater = Adam(0.01);
• convolution mode = SAME;
• l2 = 0.0001;
• convolution layer 1 = [128, 100], kernel size = [
          <xref ref-type="bibr" rid="ref3">3, 128</xref>
          ]
• convolution layer 2 = [128, 100], kernel size = [
          <xref ref-type="bibr" rid="ref4">4, 128</xref>
          ]
• convolution layer 3 = [128, 100], kernel size = [
          <xref ref-type="bibr" rid="ref5">5, 128</xref>
          ]
• merge(cl1, cl2 and cl3)
• global pooling with dropout = 0.5
• loss function = MCXENT
• dense layer = [
          <xref ref-type="bibr" rid="ref2">100, 2</xref>
          ], activation = SOFTMAX
Finally, a classifier based on LSTM with Word2Vec vectors as input, used in both Tasks 1 and
• updater = Adam(5e-3)
• l2 = 1e-5;
• weight initialization algorithm = XAVIER;
• lstm layer = [128, 256], activation = TANH);
• lstm output layer = [
          <xref ref-type="bibr" rid="ref2">256, 2</xref>
          ], activation = SOFTMAX; loss function = MCXENT
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Analysis of the results</title>
      <p>In task 1, according to the eRisk 2021 evaluation report, the maximum number of all users
writings was 2000. Of these, only 271 were processed, in 1 day 5 hours, 44 minutes and 10
seconds, until the servers were shutdown. The unavailability, at the time, of an additional GPU
made the processing time much slower and, as such, 1 day was not enough to process the whole
set. Two runs were executed, based on LSTM and TfIdf (Table 5).</p>
      <p>The final results are far from the best in all metrics. Nevertheless, the LSTM performed better,
although marginally, than TfIdf, a much simpler classifier.</p>
      <p>In task 2, the maximum number of all users writings were 1999. Of these, only 369 were
processed, taking 1 day 9 hours, 51 minutes and 27 seconds. As before, and although a GPU
was available in this task, the system was not able to process the totality of test users until the
server was shutdown. Three runs were executed, based on LSTM, CNN and TfIdf (Table 6).</p>
      <p>Run</p>
      <p>Method</p>
      <p>It seemed that the CNN performed better in some metrics, although marginally, compared
with LSTM, with TfIdf getting very low scores. Moreover, the algorithms seems to be highly
inclined to emit positive decisions, with perfect recall but extremely low precision. Although it
is not clear, this may be due to the fact that the posts are processed individually, without any
consideration of the previous writings. Some window or accumulator approach could be used
to understand if this is the issue.</p>
      <p>Overall, the three methods can be improved. They were rather close, which gives the
indication that the main issue is with the selection of the training dataset. A deeper understanding is
necessary regarding the dataset and, after that, new methods can be devised and tested.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>This paper describes the CeDRI submission to the CLEF eRisk 2021 task 1 and 2 on detecting
early signs of pathological gambling and self-harm in social media posts. Three methods were
presented that seek to classify each writing independently of the others using only information
about the text. The first task is a “test only”, so it was necessary to build a training set based
on posts collected from Reddit. Task 2 required the processing and filtering of the writings in
order to isolate the posts that refer to self-harm from the others, and use these for training the
classifiers.</p>
      <p>Due to the simple classifiers used, state-of-the-art results were not expected. The main
purpose was to try to understand the efectiveness of building training sets based on simple
heuristics filters. For future work, the inclusion of more features, such as Part of Speech (PoS)
frequency, post date and time, and others should be studied.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work has been supported by FCT – Fundação para a Ciência e Tecnologia within the Project
Scope: UIDB/05757/2020.
2020, pp. 272–287. URL: https://link.springer.com/10.1007/978-3-030-58219-7_20. doi:10.
1007/978-3-030-58219-7_20, series Title: Lecture Notes in Computer Science.
[10] S. J. Russell, P. Norvig, Artificial intelligence: a modern approach, Pearson series in artificial
intelligence, fourth edition ed., Pearson, Hoboken, 2021.
[11] D. Losada, F. Crestani, A Test Collection for Research on Depression and Language Use,
in: Proc. of Experimental IR Meets Multilinguality, Multimodality, and Interaction, 7th
International Conference of the CLEF Association, CLEF 2016, Evora, Portugal, 2016, pp.
28–39.
[12] M. M. Greaves, C. Dykeman, A Corpus Linguistic Analysis of Public Reddit Blog Posts on
Non-Suicidal Self-Injury, arXiv:1902.06689 [cs] (2019). URL: http://arxiv.org/abs/1902.06689,
arXiv: 1902.06689.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Marengo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Montag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sindermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Elhai</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Settanni, Examining the links between active Facebook use, received likes, self-esteem and happiness: A study using objective social media data</article-title>
          ,
          <source>Telematics and Informatics</source>
          <volume>58</volume>
          (
          <year>2021</year>
          )
          <article-title>101523</article-title>
          . URL: https: //linkinghub.elsevier.com/retrieve/pii/S0736585320301829. doi:
          <volume>10</volume>
          .1016/j.tele.
          <year>2020</year>
          .
          <volume>101523</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Faelens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hoorelbeke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Soenens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Van Gaeveren</surname>
          </string-name>
          ,
          <string-name>
            <surname>L</surname>
          </string-name>
          . De Marez, R. De Raedt,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Koster</surname>
          </string-name>
          ,
          <article-title>Social media use and well-being: A prospective experience-sampling study</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>114</volume>
          (
          <year>2021</year>
          )
          <article-title>106510</article-title>
          . URL: https://linkinghub.elsevier.com/retrieve/ pii/S0747563220302624. doi:
          <volume>10</volume>
          .1016/j.chb.
          <year>2020</year>
          .
          <volume>106510</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <article-title>A review on assessment, early warning and auxiliary diagnosis of depression based on diferent modal data</article-title>
          , in: Z.
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          Hei (Eds.),
          <source>Twelfth International Conference on Graphics and Image Processing (ICGIP</source>
          <year>2020</year>
          ), SPIE,
          <source>Xi'an, China</source>
          ,
          <year>2021</year>
          , p.
          <fpage>75</fpage>
          . URL: https://www.spiedigitallibrary.org/conference-proceedings-of-spie/11720/ 2589413/A-review
          <article-title>-on-assessment-early-warning-and-auxiliary-diagnosis-</article-title>
          <source>of/10.1117/12</source>
          . 2589413.full.
          <source>doi:10.1117/12</source>
          .2589413.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Moulahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Azé</surname>
          </string-name>
          , S. Bringay,
          <article-title>DARE to Care: A Context-Aware Framework to Track Suicidal Ideation on Social Media</article-title>
          , in: A.
          <string-name>
            <surname>Bouguettaya</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Klimenko</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Dzerzhinskiy</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>S. V.</given-names>
          </string-name>
          <string-name>
            <surname>Klimenko</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          (Eds.),
          <source>Web Information Systems Engineering - WISE</source>
          <year>2017</year>
          , volume
          <volume>10570</volume>
          , Springer International Publishing, Cham,
          <year>2017</year>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>353</lpage>
          . URL: http://link.springer.com/10.1007/978-3-
          <fpage>319</fpage>
          -68786-5_
          <fpage>28</fpage>
          . doi:
          <volume>10</volume>
          .1007/ 978-3-
          <fpage>319</fpage>
          -68786-5_28, series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , G. Bors, “
          <article-title>Less is more”: Mining useful features from Twitter user profiles for Twitter user classification in the public health domain</article-title>
          ,
          <source>Online Information Review</source>
          <volume>44</volume>
          (
          <year>2019</year>
          )
          <fpage>213</fpage>
          -
          <lpage>237</lpage>
          . URL: https://www.emerald.com/insight/content/doi/10.1108/OIR-05-2019-0143/ full/html. doi:
          <volume>10</volume>
          .1108/OIR-05-2019-0143.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , J. Parapar, eRISK
          <year>2017</year>
          :
          <article-title>CLEF Lab on Early Risk Prediction on the Internet: Experimental Foundations</article-title>
          , in: G. J.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Lawless</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , volume
          <volume>10456</volume>
          , Springer International Publishing, Cham,
          <year>2017</year>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>360</lpage>
          . URL: http://link.springer.com/10.1007/978-3-
          <fpage>319</fpage>
          -65813-1_
          <fpage>30</fpage>
          . doi:
          <volume>10</volume>
          . 1007/978-3-
          <fpage>319</fpage>
          -65813-1_30, series Title: Lecture Notes in Computer Science.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          , Overview of eRisk 2018:
          <article-title>Early Risk Prediction on the Internet (extended lab overview</article-title>
          ) (
          <year>2018</year>
          )
          <fpage>20</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <article-title>Overview of eRisk 2019 Early Risk Prediction on the Internet</article-title>
          , in: F. Crestani,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Savoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rauber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. Heinatz</given-names>
            <surname>Bürki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction, Lecture Notes in Computer Science</source>
          , Springer International Publishing, Cham,
          <year>2019</year>
          , pp.
          <fpage>340</fpage>
          -
          <lpage>357</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -28577-7_
          <fpage>27</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          , Overview of eRisk 2020:
          <article-title>Early Risk Prediction on the Internet</article-title>
          , in: A.
          <string-name>
            <surname>Arampatzis</surname>
            , E. Kanoulas,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tsikrika</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Vrochidis</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lioma</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eickhof</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Névéol</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cappellato</surname>
          </string-name>
          , N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , volume
          <volume>12260</volume>
          , Springer International Publishing, Cham,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>