<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Natural Language Processing Based Framework for Early Detection of Anorexia via Sequential Text Processing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Prateek Sarangi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sumit Kumar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shraddha Agarwal</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tanmay Basu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Data Science and Engineering, Indian Institute of Science Education and Research</institution>
          ,
          <addr-line>Bhopal</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>0</volume>
      <fpage>9</fpage>
      <lpage>12</lpage>
      <abstract>
        <p>Task 2 of eRisk shared tasks in CLEF 2024 aims to develop text mining solutions for early prediction of anorexia using sequentially posted texts over social media. Anorexia is an eating disorder, a kind of mental illness, where people have distorted perceptions of their body weights and accordingly manipulate their food habits, which often results in deficient body weight. The aim here is to identify anorexia by processing the interactions over social media of an individual through text mining models. The organisers provided training data with ground truths for training the models and test data for evaluating their performance. The BioNLP research group at the Indian Institute of Science Education and Research Bhopal (IISERB) participated in Task 2 and submitted five runs for five text mining frameworks. Five diferent classifiers and individual feature engineering techniques were used to develop frameworks. The bag-of-words model and transformer-based embedding were used as features for individual classifiers. The performance of multiple classifiers was evaluated using the training corpus. Then Random Forest, Adaptive Boosting, Logistic Regression, Support Vector Machine (SVM), and Longformer classifiers were chosen to run on the test set. Experimental results show that SVM and AdaBoost classifiers using the TF-IDF-based weighting schemes achieved the highest precision score among all submitted runs in Task 2 of eRisk 2024. However, the performance of our models in terms of the other metrics, like recall, f-score, ERDE, etc., is not reasonably good compared to the other runs. Hence, we plan to develop transformer-based embeddings from scratch using data collected from multiple social media platforms.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;BioNLP</kwd>
        <kwd>Information Extraction</kwd>
        <kwd>Text Classification</kwd>
        <kwd>Mental Health</kwd>
        <kwd>Anorexia Detection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The rise of social media has revolutionised communication and opened up new opportunities for
sociological and psychological research, particularly in mental health [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Anorexia nervosa, a severe
eating disorder characterised by an intense fear of weight gain and a distorted body image leading to
self-starvation and significant health issues, is a crucial area of focus. Social media’s widespread use
provides a unique, real-time view into behaviours and expressions that may be useful to identify such
mental disorders[
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ]. The vast amount of multiple generated content on platforms like Facebook,
Twitter, and Reddit ofers an extensive data pool for detecting early signs of mental health conditions
like anorexia [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Such contents, including written posts, comments, shared photos, and likes, can be
analysed for patterns indicating the onset of anorexia, which may be helpful for early interventions [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Recognising social media’s potential as a public health tool, the Conference and Labs of the Evaluation
Forum (CLEF) launched the eRisk initiative in 2018 to explore this potential. Initially focused on
depression, the initiative later included anorexia, showcasing a broader commitment to using big data
for mental health monitoring and intervention.
      </p>
      <p>
        In 2024, the CLEF eRisk second shared task highlights the importance of sequential text processing
to identify anorexia markers as they appear in social media posts [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]; in the way the mental health
professionals analyse patient behaviours over a while for treating a mental illness [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This method
aids not only in early detection but also in understanding the progression of anorexia through digital
footprints. The Bio-NLP group at the Indian Institute of Science Education and Research Bhopal (IISERB)
has developed comprehensive text mining frameworks for this task. These frameworks explore text
feature extraction methods like transformer-based embeddings and classical bag-of-words models for
semantic interpretation of social media text to identify the indicators of anorexia [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Subsequently,
various text classifiers viz. Random Forest, Adaptive Boosting, Logistic Regression, SVM, and Longformer
were trained on multiple features derived from the text data to categorise the posts to either anorexia
or control group. The proposed frameworks are presented in section 2 of the paper. The organisers
provided a training corpus with ground truths for developing the models and a separate test corpus for
evaluating the performance of the submitted runs. The empirical analysis of the proposed frameworks
on both the training and test corpora are provided in section 3.3 in terms of the evaluation metrics
provided by eRisk organisers, like ERDE, Latency, F-score, etc. [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ].
      </p>
      <p>The experimental evaluation shows that two of our runs containing Adaptive Boosting and SVM
classifiers using the TF-TDF model achieve the first and second ranks in precision among all the runs
submitted for this task. Moreover, our Longformer model achieves the best latency and speed, and
another run having a Logistic Regression and Entropy-based bag of words model achieves the best
speed among all 44 runs submitted for task 2 of eRisk 2024. However, our models could not achieve
reasonable recall, F1 and ERDE scores compared to many other runs submitted for this task. Our models
show poor performance for ranking-based evaluations. Hence, we need to investigate this direction
to improve the performance. We plan to train a transformer-based model from scratch using the data
collected from various social media sites like Reddit and Twitter to identify the subtle nuances in the
sequential texts processed over time. Furthermore, in future, we need to consider the timestamp of the
posts as a significant indicator in the model for semantic interpretation of the sequential texts.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Proposed Frameworks</title>
      <p>In the efort to identify early signs of anorexia through sequential text processing of social media
conversations, particularly from Reddit, we have developed various text classification frameworks. Our
methodologies are designed to harness the vast amount of unstructured text in these digital interactions,
formatted in XML and provided by the organisers of the eRisk 2024 challenge.</p>
      <sec id="sec-2-1">
        <title>2.1. Feature Engineering</title>
        <p>The feature engineering approaches we have explored are crafted to capture the intricate linguistic and
semantic patterns from sequential social media texts to identify the signs of anorexia. Both classical
Bag of Words (BoW)—-based features and recent transformer-based embeddings are used to train the
classification models. The following three feature selection techniques generated features from the
given training corpus.</p>
        <sec id="sec-2-1-1">
          <title>2.1.1. TF-IDF Weighting Scheme of BoW</title>
          <p>
            Each document, representing a user’s aggregated posts, is converted into a vector with each unique
term as a feature. This model provides a fundamental basis for more advanced analyses. The BoW
approach[
            <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
            ] considers each unique term as a feature, known as unigram, to constitute the set
of features of a text collection known as vocabulary. Sometimes, it considers two, three, or multiple
terms as one feature based on the significance of the text sequence, known as bigram, trigram, or
n-grams, respectively. In the experimental analysis, we explored the classifiers’ performance using
unigrams, bigrams and trigrams. After generating the dictionary of a corpus, the Term Frequency-Inverse
Document Frequency (TF-IDF) weighting is used to develop the vector of a given text[
            <xref ref-type="bibr" rid="ref14 ref15 ref16">14, 15, 16, 17</xref>
            ].
TF counts the number of terms in a given text, whereas IDF of a term, say, t, is defined as
IDF() = log
︂(
          </p>
          <p>)︂
 ()</p>
          <p>Where  is the total number of texts in the text collections, and  () is the number of texts in a
corpus containing the term t.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>2.1.2. Entropy Based Weighting Scheme of BoW</title>
          <p>
            This scheme assigns weights to the terms in a document based on their entropy, which measures the
amount of information or uncertainty associated with each term. Many researchers use the
entropybased term weighting technique to form a term-document matrix from a text collection [
            <xref ref-type="bibr" rid="ref13 ref15 ref16">13, 15, 16, 17, 18</xref>
            ].
This method is developed in the spirit that the more important term is the more frequent one that
occurs in fewer documents, taking the distribution of the term over the corpus into account [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. The
weight of a term in a document is determined by the entropy of the term frequency of the term in that
document [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. The weight ( ) of the ℎ term in the ℎ document is defined by the Entropy 1 [
            <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
            ]
model as follows:
 = log (︀  + 1)︀ ×
︃(
1 +

=1
∑︀  log  )︃
log( + 1)
,
where,  =
          </p>
          <p>
            (1)


∑︀ 
=1
Here, N is the total number of documents in the corpus, and  is the frequency of ℎ term in the ℎ
document of the corpus. Generally, the BoW model generates many terms, making the term-document
matrix sparse and high dimensional, which can badly afect the performance of the text classifiers
[17]. Hence,  2-statistic-based term selection technique was used for both TF-IDF and Entropy-based
term weighting schemes in the experiments to identify essential terms from the term-document matrix,
which is a widely used technique for term selection [
            <xref ref-type="bibr" rid="ref13 ref16">13, 16, 19</xref>
            ].
          </p>
        </sec>
        <sec id="sec-2-1-3">
          <title>2.1.3. Transformer-Based Embeddings</title>
          <p>Utilising cutting-edge transformer architectures like Longformer2, we create dense vector
representations of text that capture deep contextual nuances far beyond traditional models. Bidirectional Encoder
Representations from Transformers (BERT) is a contextualised word representation model based on a
masked language model and pre-trained using bidirectional transformers on general domain corpora,
i.e., English Wikipedia and books [20]. The Longformer model[21] was chosen because it performs
better than BERT at understanding long-term relationships in texts[17]. It creates feature embeddings
to help detect early signs of anorexia by recognising language patterns, from explicit mentions of body
image issues to under-expressed signs of distress.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Text Classification Methods</title>
        <p>Our approach to text classification combines engineered features with several machine learning
algorithms chosen for their ability to handle complex relationships within high-dimensional data. We use
Support Vector Machines (SVM) with linear and RBF kernels for their eficiency in high-dimensional
spaces and binary classification tasks, finding the optimal hyperplane to separate classes and maximising
the margin between them. This method is particularly efective for its robustness against overfitting
[22]. Logistic Regression (LR), a probabilistic model, is also employed with L1 and L2 regularisation due
1https://radimrehurek.com/gensim/models/logentropy_model.html
2https://huggingface.co/allenai/longformer-base-4096
to its interpretability and efectiveness for binary outcomes [ 23]. LR uses the probability of a particular
class by fitting a logistic function to the data, providing precise, interpretable coeficients for each
feature, which helps understand the influence of diferent features on the likelihood of anorexia.</p>
        <p>In addition, we have used Adaptive Boosting (AB), which improves classification accuracy by
combining multiple weak classifiers, such as decision trees, in our model. AB sequentially applies a weak
classifier to the data, adjusting the weights of misclassified instances so subsequent classifiers focus more
on dificult cases, enhancing performance on complex textual data [ 24]. Furthermore, we incorporate
Longformer, a Transformer-based model, which is pre-trained on large datasets and fine-tuned on
our specific dataset [ 21]. This model is used to understand more extended context and relationships
between words in a text to capture language patterns in user interactions.</p>
        <p>By training these classifiers on the features derived from the textual data, we aim to identify individual
posts indicating the risk of anorexia. We systematically analyse the performance of these models on
training data and validate them on unseen test data to develop a scalable and efective tool for the
early detection of anorexia. The SVM, Logistic Regression, and AdaBoost models were implemented
using Scikit-learn3, while the transformer-based models were fine-tuned using the Hugging Face
Transformers library [25]. In Section 3.3, we present the experimental results, showcasing the eficacy
of our frameworks for identifying signs of anorexia through social media interactions.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental Analysis</title>
      <p>The dataset provided by the eRisk 2024 challenge organisers includes 2,335 files, each containing a
collection of Reddit posts from an individual user. In XML format, these files include metadata such as
post dates and titles. Among these, 2,062 files are labelled as class "0" (no signs of anorexia, i.e., control
group), while 273 are labelled as class "1" (potential signs of anorexia).</p>
      <sec id="sec-3-1">
        <title>3.1. Data Preprocessing</title>
        <p>For our analysis, we extracted and merged all conversations of each subject from the XML files. It was
observed that the text field was empty in some files, but the title field had vital text, ensuring that either
had the data to be considered in such cases instead of just concentrating solely on the text field. This
preprocessed data was input for our feature engineering and classification pipelines.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Experimental Settings</title>
        <p>To assess the performance of our frameworks on the raining corpus, we utilised several evaluation
metrics commonly employed in binary classification tasks, including precision, recall and F1-score.
Stratified 5-fold cross-validation has been conducted on the training dataset to evaluate our models’
performance and tune the hyperparameters. Grid search4 and random search techniques were used to
identify the optimal hyperparameter settings for each model following the cross-validation scores. The
best-performing models from the cross-validation stage were then tested on the unseen test dataset.
The results of this evaluation are presented and discussed in Section 3.3.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Results and Discussion</title>
        <p>The BioNLP-IISERB team’s participation in the eRisk 2024 challenge aimed to detect anorexia early via
sequential text processing. The approach involved several experimental runs tailored to explore the
optimal parameters and methodologies.</p>
        <p>Table 2 shows the experimental comparison of various Bag of Words models across diferent classifiers
on a training corpus. Support Vector Machine (SVM) with TF-IDF feature selection emerged as the
top performer, achieving the highest precision (0.9639) and F1 score (0.8009). Logistic Regression (LR)
3http://scikit-learn.org/stable/supervised_learning.html
4https://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GridSearchCV.html
also performed strongly, particularly with TF-IDF, achieving the highest recall (0.6777) among all
models. The Adaboost classifier showed moderate performance, while Random Forest models generally
underperformed. Surprisingly, the LongFormer model showed remarkably poor performance across all
metrics and needed further investigation.</p>
        <p>A clear trend can be seen in the superiority of TF-IDF over Entropy-based feature selection across
all traditional classifiers, often by a significant margin. This suggests that TF-IDF’s ability to capture
both term importance and distinctiveness is precious for this classification task. The optimal number
of terms varied across classifiers, with SVM and LR performing best with 10,000 terms, while RF and
Adaboost achieved their best performance with fewer terms (1000 and 1500, respectively). These results
highlight the importance of classifier selection and feature engineering in text classification tasks, with
SVM and LR coupled with TF-IDF emerging as strong baseline models for similar problems.</p>
        <p>The performance exhibited in Table 3 with considerable variation across multiple experimental runs.
Run 1 (TF-IDF+BoW+LR) demonstrated superior performance with a high F1 score (0.62), indicating
an optimal balance between precision and recall. Run 3 and run 4 showed consistent performance,
both achieving an F1 score of 0.58 and 0.67, respectively. Notably, these runs also exhibited the lowest
ERDE50 values (0.08), suggesting enhanced early recognition capabilities.</p>
        <p>In contrast, run 0 (Entropy+BoW+LR) and run 2 (Longformer) obtained F1 scores of 0.32 and 0.25,</p>
        <p>BioNLP</p>
        <p>IISERB0
(Entropy +</p>
        <p>SVM)
0.10
0.19
0.06
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
0.00
respectively. These runs also had elevated ERDE50 values (0.10 and 1.00), indicating reduced efectiveness
in early correct decision recognition. Table 3 further demonstrates several strengths of the proposed
approaches, such as runs 1, run 3 and run 4 exhibited high precision balanced with satisfactory recall,
resulting in robust F1 scores. Moreover, run 3 and run 4 showed lower latencies (0.99 and 0.99), indicating
eficient decision-making processes. The consistent performance in Run 3 (TF-IDF+BoW+SVM) and 4
(TF-IDF+BoW+AB) suggests a reliable and stable approach.</p>
        <p>However, certain limitations were observed where run 0 and run 2 displayed notably lower recall
values (0.23 and 0.16), adversely afecting their overall F1 scores. Additionally, run 0 and run 2 exhibited
higher latency, potentially due to ineficiencies in the early stages of model deployment.</p>
        <p>The ranking-based evaluations in Table 4 depicted a decline in performance metrics as the number of
processed writings escalated. Initially, promising detection capabilities at smaller datasets experienced
a stark reduction in precision and NDCG scores, diminishing to zero as the data volume increased. This
suggests that while the initial models perform well under control, smaller datasets, the efectiveness
wanes significantly under larger scales, pointing to potential overfitting or the need for more robust
generalisation capabilities in the models.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>The BioNLP@IISERB team took part in the eRisk 2024 task 2 to develop robust text mining frameworks
for early prediction of anorexia by analysing social media texts. The classical bag-of-words models
and recent transformer-based methods were explored to generate potential features for identifying
nuances in the given texts. Subsequently, diferent text classifiers were trained using these features
to identify anorexia from the given social media texts. The experimental results suggests that these
frameworks are capable of identifying textual patterns indicative of anorexia, however, there are rooms
for further improvements. Some frameworks showed promising precision and recall scores on more
minor texts, but we encountered substantial challenges in maintaining consistent performance as the
volume of data increased. This suggests the need for continued research and refinement of our methods.
By processing and analysing large-scale social media data, we can extract valuable insights into the
onset and progression of conditions like anorexia, thus enabling timely and targeted support. In the
future, more robust and dynamic models will be developed that can eficiently handle the growing
volume of user-generated content while maintaining high accuracy in detection. Additionally, exploring
multi-modal approaches that combine textual data with other types of information, such as images
and social network data, could provide more comprehensive and nuanced insights into mental health
conditions. In conclusion, by successfully applying these techniques, we would have a significant
positive impact, providing early resilience and support for individuals dealing with anorexia and other
mental health challenges. Future work should focus on stabilising recall performance and reducing
latency to consistently enhance decision-making speed and accuracy.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>Prateek Sarangi and Tanmay Basu acknowledge the support of the seed funding
(INST/DSE/20232024/18) provided by the Indian Institute of Science Education and Research Bhopal, India.
prediction of self harm over social media., in: Proceedings of International Conference of CLEF
Association, 2021, pp. 928–937.
[17] H. Srivastava, N. S. Lijin, S. Sruthi, T. Basu, Nlp-iiserb@erisk2022: Exploring the potential of bag
of words, document embeddings and transformer based framework for early prediction of eating
disorder, depression and pathological gambling over social media, in: Proceedings of Experimental
IR Meets Multilinguality, Multimodality, and Interaction: 13th International Conference of the
CLEF Association, Bologna, Italy, 2022.
[18] S. Goswami, S. Pal, S. Goldsworthy, T. Basu, An efective machine learning framework for data
elements extraction from the literature of anxiety outcome measures to build systematic review,
in: Business Information Systems: 22nd International Conference, BIS 2019, Seville, Spain, June
26–28, 2019, Proceedings, Part I 22, Springer, 2019, pp. 247–258.
[19] T. Basu, C. Murthy, A supervised term selection technique for efective text categorization,</p>
      <p>International Journal of Machine Learning and Cybernetics 7 (2016) 877–892.
[20] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional transformers
for language understanding, arXiv preprint arXiv:1810.04805 (2018).
[21] I. Beltagy, M. E. Peters, A. Cohan, Longformer: The long-document transformer, arXiv preprint
arXiv:2004.05150 (2020).
[22] C. Cortes, V. Vapnik, Support-vector networks, Machine learning 20 (1995) 273–297.
[23] D. W. Hosmer, S. Lemeshow, R. X. Sturdivant, Applied logistic regression, volume 398, John Wiley
&amp; Sons, 2013.
[24] Y. Freund, R. E. Schapire, A decision-theoretic generalization of on-line learning and an application
to boosting, Journal of computer and system sciences 55 (1997) 119–139.
[25] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, T. Rault, R. Louf, M.
Funtowicz, et al., Transformers: State-of-the-art natural language processing, in: Proceedings of the
2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations,
2020, pp. 38–45.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>W.</given-names>
            <surname>Ragheb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Azé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bringay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Servajean</surname>
          </string-name>
          ,
          <article-title>Attentive multi-stage learning for early risk detection of signs of anorexia and self-harm on social media</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>M. De Choudhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Gamon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Counts</surname>
          </string-name>
          , E. Horvitz,
          <article-title>Predicting depression via social media</article-title>
          .,
          <source>ICWSM</source>
          <volume>13</volume>
          (
          <year>2013</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>M. De Choudhury</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Counts</surname>
          </string-name>
          , E. Horvitz,
          <article-title>Social media as a measurement tool of depression in populations</article-title>
          ,
          <source>in: Proceedings of the 5th Annual ACM Web Science Conference</source>
          , ACM,
          <year>2013</year>
          , pp.
          <fpage>47</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>L. G.</surname>
          </string-name>
          et al.,
          <article-title>Machine learning and natural language processing in mental health: systematic review</article-title>
          ,
          <source>Journal of Medical Internet Research</source>
          ,
          <volume>23</volume>
          (
          <year>2021</year>
          )
          <article-title>e15708</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Paul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Jandhyala</surname>
          </string-name>
          , T. Basu,
          <article-title>Early detection of signs of anorexia and depression over social media using efective machine learning frameworks</article-title>
          .,
          <source>in: CLEF (Working notes)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. W.</given-names>
            <surname>Eidem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cetta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Webb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. C.</given-names>
            <surname>Graham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Jay</surname>
          </string-name>
          ,
          <article-title>Early detection of cardiac dysfunction: use of the myocardial performance index in patients with anorexia nervosa</article-title>
          ,
          <source>Journal of adolescent health 29</source>
          (
          <year>2001</year>
          )
          <fpage>267</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Rodilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , Overview of erisk 2024:
          <article-title>Early risk prediction on the internet, in: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          ,
          <source>15th International Conference of the CLEF Association, CLEF 2024</source>
          , Springer International, Grenoble, France,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ranganathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Haritha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Thenmozhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Aravindan</surname>
          </string-name>
          ,
          <article-title>Early detection of anorexia using rnn-lstm and svm classifiers</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>F.</given-names>
            <surname>Galetta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Franzoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cupisti</surname>
          </string-name>
          , E. Morelli,
          <string-name>
            <given-names>G.</given-names>
            <surname>Santoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pentimone</surname>
          </string-name>
          ,
          <article-title>Early detection of cardiac dysfunction in patients with anorexia nervosa by tissue doppler imaging</article-title>
          ,
          <source>International journal of cardiology 101</source>
          (
          <year>2005</year>
          )
          <fpage>33</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Parapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Rodilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Losada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Crestani</surname>
          </string-name>
          , Overview of erisk 2024:
          <article-title>Early risk prediction on the internet (extended overview)</article-title>
          ,
          <source>in: Working Notes of the Conference and Labs of the Evaluation Forum CLEF</source>
          <year>2024</year>
          ,
          <article-title>CLEF 2024</article-title>
          , CEUR Workshop Proceedings, Grenoble, France,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Marion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lacroix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Caquard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Scherdel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. G. L.</given-names>
            <surname>Guen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Caldagues</surname>
          </string-name>
          , E. Launay,
          <article-title>Earlier diagnosis in anorexia nervosa: better watch growth charts!</article-title>
          ,
          <source>Journal of eating disorders 8</source>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Salton</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. J. McGill</surname>
          </string-name>
          , Introduction to Modern Information Retrieval,
          <source>McGraw Hill</source>
          ,
          <year>1983</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Goldsworthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Gkoutos</surname>
          </string-name>
          ,
          <article-title>A sentence classification framework to identify geometric errors in radiation therapy from relevant literature</article-title>
          ,
          <source>Information</source>
          <volume>12</volume>
          (
          <year>2021</year>
          )
          <fpage>139</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Selamat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Omatu</surname>
          </string-name>
          ,
          <article-title>Web page feature selection and classification using neural networks</article-title>
          ,
          <source>Information Sciences 158</source>
          (
          <year>2004</year>
          )
          <fpage>69</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sabbah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Selamat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Selamat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Al-Anzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Viedma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Krejcar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fujita</surname>
          </string-name>
          ,
          <article-title>Modified frequency-based term weighting schemes for text classification</article-title>
          ,
          <source>Applied Soft Computing</source>
          <volume>58</volume>
          (
          <year>2017</year>
          )
          <fpage>193</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. V.</given-names>
            <surname>Gkoutos</surname>
          </string-name>
          ,
          <article-title>Exploring the performance of baseline text mining frameworks for early</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>