<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>mood states from speech in bipolar disorder</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cristina Crocamo</string-name>
          <email>cristina.crocamo@unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aurelia Canestro</string-name>
          <email>a.canestro@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dario Palpella</string-name>
          <email>d.palpella@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Riccardo M. Cioni</string-name>
          <email>r.cioni1@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Nasti</string-name>
          <email>c.nasti@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Susanna Piacenti</string-name>
          <email>s.piacenti1@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandra Bartoccetti</string-name>
          <email>a.bartoccetti@campus.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Martina Re</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentina Simonetti</string-name>
          <email>valentinasimonetti@ab-acus.eu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chiara Barattieri di San Pietro</string-name>
          <email>cbarattieri@fatebenefratelli.eu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Bulgheroni</string-name>
          <email>mariabulgheroni@ab-acus.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Bartoli</string-name>
          <email>francesco.bartoli@unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Carrà</string-name>
          <email>giuseppe.carra@unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ab.Acus s.r.l. Milan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Workshop</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Medicine and Surgery, University of Milano-Bicocca</institution>
          ,
          <addr-line>Monza</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Regular monitoring is essential to efectively track mood fluctuations and assess ongoing treatment needs for mood disorders (e.g., identifying early signs of relapse, adjusting therapeutic interventions, and improving longterm outcomes). The current ongoing work aims at assessing the relationships between language and symptom severity in people with bipolar disorders, thus investigating potential mHealth mood detection mechanisms based on speech patterns. Acoustic features included conversational measures for nonverbal language and statistics for prosodic cues. Preliminary results, combining acoustic features and natural language processing (NLP) scores, were promising, somehow discriminating clinical conditions of people with BD when assessing their mood states. This approach may ofer potential benefits for individualized mental health care and early intervention approaches in real-world scenarios.</p>
      </abstract>
      <kwd-group>
        <kwd>speech</kwd>
        <kwd>signal analysis</kwd>
        <kwd>mood states</kwd>
        <kwd>mHealth</kwd>
        <kwd>remote assessment</kwd>
        <kwd>machine learning</kwd>
        <kwd>neural network</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Bipolar disorder (BD) is a lifelong episodic illness resulting in reduced psychosocial functioning. The
majority of BD cases onset in early adulthood and it is among the leading causes of disability in
workingage adults [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Community services often struggle in delivering regular monitoring of treatment needs,
contributing to a gap in care [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The assessment of mood states and potential variations is pivotal in
BD [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. Because of its chronicity, approaches for prediction and prevention of further episodes in
which the patient’s mood and activity levels are considerably disturbed are critical [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Mood states are
defined referring to the presence and severity of depressive symptoms as assessed by the
MontgomeryÅsberg Depression Rating Scale (MADRS), including items that measure sadness feelings [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and manic
symptoms as assessed by the Young Mania Rating Scale (YMRS), measuring elevated mood and increased
activity levels [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Higher scores indicate more severe depressive or manic symptoms, respectively.
      </p>
      <p>Traditionally, this evaluation has heavily relied on clinical interviews, including an analysis of
thoughts and their manifestation in language of people with BD. Focusing on meaning and
communication, language has a central role for diagnosis and treatment in BD with speech patterns being
LGOBE
(G. Carrà)</p>
      <p>CEUR</p>
      <p>
        ceur-ws.org
crucial when assessing current experiences, emotions, and thoughts. For instance, pressure of speech
encompassing a high number of words during phonation and a small number of pauses is likely to
be a sign of underlying manic symptoms. Conversely, mood states in depression are characterized by
poverty of speech and a monotone pitch [
        <xref ref-type="bibr" rid="ref10 ref7 ref8 ref9">7, 8, 9, 10</xref>
        ]. Therefore, we hypothesized that speech patterns
would relate to standard psychometric assessments in BD, discriminating clinical conditions when
predicting individual mood states.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        Progress in Machine Learning (ML) and Natural Language Processing (NLP) techniques may support
the development of automated systems assessing speech patterns as objective markers of mood states
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. A recent review highlighted favourable evidence about the use of audio data to monitor mood
disorders, despite some challenges [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. However, prior research emphasized the potential for the
use of speech mainly to distinguish between individuals with and without a variety of psychiatric
disorders, including BD [
        <xref ref-type="bibr" rid="ref13 ref9">9, 13</xref>
        ]. Alternatively, some studies focused on the correlation between acoustic
features and parameters from electroglottographic signal of voiced segments [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Acoustic features (e.g.,
jitter) were identified as likely reflecting a dysregulation of autonomous nervous system that influence
muscular tone and articulatory control. Consistently, specifically considering mood fluctuations among
people with BD, available evidence suggests that speech patterns impairments may be sensitive and
valid measures of mood states [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ]. Previous work on ecological speech signal analysis from phone
calls recordings showed model ability to diferentiate between hypomanic and euthymic as well as
between depressed and euthymic speech according to a support vector machine classifier (average AUC
of 0.81 for hypomania and 0.67 for depression) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Similarly, a more recent study based on phone
calls data trained ML models considering random forest classifiers to classify mood states according to
estimated voice features [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Although a varying accuracy (0.61 to 0.74) was estimated when classifying
a depressive or manic versus a euthymic state, these approaches seem promising to complement rating
scales with speech markers, thus possibly improving mood states monitoring in real-world settings.
      </p>
      <p>
        However, there are several barriers in the implementation of mood detection systems in real-world
applications, including high degree of heterogeneity between studies and the use of non-standardized
metrics reporting [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Moreover, several areas remain understudied, including the use of speech
spectrograms testing performance to remotely assess the individual clinical status among people with BD.
However, a few studies explored speech-based emotion recognition using spectrograms in related fields.
A recent study proposed a convolutional neural network for anger and stress detection using handcrafted
features and deep learned features [18]. A further study explored speech emotion recognition from
the utterances of interacting professional actors performing spontaneously, by exploiting a novel
convolutional neural network architecture to recognize speech emotions based on local correlations
and global contextual information from speech spectrograms [19]. These approaches have been proven
successful, with high accuracy for speech emotion recognition, emphasizing related feasibility when
processing speech segments. In addition, these systems can be integrated with other signals (i.e.,
linguistic and paralinguistic components of speech), thus implementing the analysis of a multimodal
signal [19].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Speech signal analysis and deep learning for mood prediction</title>
      <p>Referring to the process of examining and interpreting the characteristics of spoken output, speech
signal analysis was used to identify key patterns from acoustic signals generated during speech. Speech
signal analysis was proven efective to characterize mood states of people with BD, thus contributing to
individualized approaches including the estimate of various speech features based on signal’s frequency
and energy/amplitude [20]. Remarkably, deep learning techniques may significantly advance speech
signal processing, enabling more accurate recognition, analysis, and interpretation of individuals’
language, especially for mood detection. Indeed, deep learning encompasses ML techniques that can
automatically learn hierarchical representations from data. Core models embrace neural networks
(NN) including multiple interconnected layers of nodes (mimicking the structure and functioning of
the human brain) as well as convolutional neural networks (CNN), in which the nodes of each layer
are clustered [21]. In addition, considering spectrograms in speech analysis systems, the latter may
foster automatically learning representations from image data according to a benchmark performance
in image classification [ 19].
3.1. Proposed approach
Assessing mood fluctuations based on gold-standard assessments of mania and depression in BD, the
proposed approach aimed at exploiting deep learning algorithms for mood states prediction. Eligible
participants involved subjects with a diagnosis of BD, aged between 18 and 65 years old from both
inpatient and outpatient services. They were approached by specifically trained staf. Subjects unable
to provide informed consent, those with vocal or hearing issues were excluded. Speech data were
collected and processed through a mobile app that study participants accessed on their smartphones
using password-protected access to self-administer verbal performance tasks. In particular, the system
embedded in the smartphone was grounded on a cloud-based architecture hosting the system database,
the Representational State Transfer (REST) Application Programming Interfaces (API), and the backend
processing modules.</p>
      <p>Considering mood states variability and relevant speech signal segmentation, we aimed to combine
two diferent approaches for speech analysis. First, highlighting speaking segments from raw audio
data, speech was automatically processed through speech recognition and quantities representing voice
characteristics (i.e., acoustic features) were estimated. By leveraging Parselmouth module as a bridge
to speech-to-text preprocessing and related Praat’s built-in functions, basic acoustic features were
computed (e.g., fundamental frequency, harmonics-to-noise ratio, jitter and shimmer). Speech rate,
verbal task duration, and phonation duration were also considered. In addition, acoustic features were
integrated with both standard and novel NLP scores for linguistic components of speech according to
distributional semantic models (e.g., estimating information on processing speed and capturing both
lexical overlap and semantic similarity in the spoken output). As a whole, previous studies showed an
enhanced performance considering combined features [22]. With predictive accuracy as the primary
goal rather than understanding the exact contributions of individual features (both speech signal-derived
and NLP-extracted features), the models’ architecture was based on a feedforward neural network
with fully connected layers (Rectified Linear Unit -ReLU- activation in two hidden layers and sigmoid
activation in the output layer). The Adaptive Moment Estimation (adam) optimizer was used with the
model seeking to minimize cross-entropy loss. The model stops the training if the validation loss does
not improve for 10 epochs. A 5-fold cross-validation was used with each fold providing metrics that are
averaged for the final evaluation based on 80% of the data for training and 20% for testing.</p>
      <p>On the other hand, by considering the feasibility of performing acoustic signalling analysis as a
function of frequency and time, this ongoing work aimed to focus on potential speech segments from
spectrograms for image data classification. Higher energy against low energy regions (e.g., pauses) in
spectrograms may be distinguished by darker/lighter colours. Consistently, spectrograms may be able
to display the properties of a changing signal through a series of snapshots according to segment length
with speech corpora capturing tones, emotions, rhythms among signals beyond content of speech.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>Based on the caseload of the ASST Nord Milano Mental Health Care Trust, 37 subjects with BD were
involved, while enrolment is still active. Involved participants mainly lived alone or with family,
were globally educated, but unemployed. They were likely to report severe depressive symptoms
(MADRS score ≥19, 46%) and just a few had severe manic features (YMRS score ≥20; 24%). Therefore, we
preliminarily focused on MADRS assessment (Table 1), taking into account sex-specific discriminating
ability of NLP-based and acoustic features for mood states prediction.</p>
      <sec id="sec-4-1">
        <title>Item</title>
        <sec id="sec-4-1-1">
          <title>Apparent Sadness</title>
        </sec>
        <sec id="sec-4-1-2">
          <title>Reported Sadness</title>
        </sec>
        <sec id="sec-4-1-3">
          <title>Inner Tension</title>
        </sec>
        <sec id="sec-4-1-4">
          <title>Reduced Sleep</title>
        </sec>
        <sec id="sec-4-1-5">
          <title>Reduced Appetite</title>
        </sec>
        <sec id="sec-4-1-6">
          <title>Concentration Dificulties</title>
        </sec>
        <sec id="sec-4-1-7">
          <title>Lassitude</title>
        </sec>
        <sec id="sec-4-1-8">
          <title>Inability to Feel</title>
        </sec>
        <sec id="sec-4-1-9">
          <title>Pessimistic Thoughts</title>
        </sec>
        <sec id="sec-4-1-10">
          <title>Suicidal Thoughts</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Rating description</title>
        <sec id="sec-4-2-1">
          <title>Representing despondency, gloom and despair, (more than just ordinary transient low spirits) reflected in speech, facial expression, and posture. Rated by depth and inability to brighten up.</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>Representing reports of depressed mood, regardless of whether it is reflected in</title>
          <p>appearance or not. Includes low spirits, despondency or the feeling of being
beyond help and without hope. Rated according to intensity, duration and the
extent to which the mood is reported to be influenced by events.</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>Representing feelings of ill-defined discomfort, edginess, inner turmoil, mental</title>
          <p>tension mounting to either panic, dread or anguish. Rated according to intensity,
frequency, duration and the extent of reassurance called for.</p>
        </sec>
        <sec id="sec-4-2-4">
          <title>Representing the experience of reduced duration or depth of sleep compared to the subject’s own normal pattern when well.</title>
        </sec>
        <sec id="sec-4-2-5">
          <title>Representing the feeling of a loss of appetite compared with when well. Rated by loss of desire for food or the need to force oneself to eat.</title>
        </sec>
        <sec id="sec-4-2-6">
          <title>Representing dificulties in collecting one’s thoughts mounting to incapacitating lack of concentration. Rated according to intensity, frequency, and degree of incapacity produced.</title>
        </sec>
        <sec id="sec-4-2-7">
          <title>Representing a dificulty getting started or slowness initiating and performing everyday activities.</title>
        </sec>
        <sec id="sec-4-2-8">
          <title>Representing the subjective experience of reduced interest in the surroundings, or activities that normally give pleasure. The ability to react with adequate emotion to circumstances or people is reduced.</title>
        </sec>
        <sec id="sec-4-2-9">
          <title>Representing thoughts of guilt, inferiority, reproach, sinfulness, remorse and ruin.</title>
        </sec>
        <sec id="sec-4-2-10">
          <title>Representing the feeling that life is not worth living, that a natural death would be welcome, suicidal thoughts, and preparations for suicide. Suicidal attempts should not in themselves influence the rating.</title>
          <p>Model performance was comprehensively evaluated according to accuracy and Receiver Operating
Characteristic (ROC) Area Under the Curve (AUC) estimates to assess the ability of the models to
correctly classify mood states. In particular, two classes for symptom severity were considered for
classification (i.e., severe/not severe). Neural network models developed -including diferent sets of
speech features and considering a chance level of 0.5 - showed varying levels of performance (Table 2).
Notably, NLP features, such as mean intraword time and semantic similarity between words, provided
satisfactory results as compared to models relying on acoustic features only (e.g., fundamental frequency,
and jitter- and shimmer-related features). However, further analysis revealed that sex-based diferences
influenced the models’ ability to accurately discriminate between mood states, thus suggesting that sex
may modulate the expression of mood in both linguistic and acoustic features.</p>
          <p>Figure 1 shows sample speech spectrograms of two study participants with diferent levels of symptom
severity. While sampled spectrograms might have relatively limited representativeness, related visual
quality -based on the identification of key features- can indicate how well this captures the underlying
patterns in the data, thus providing a useful representation of the signal for further analyses. Indeed,
visual inspection of relevant spectrograms of people with BD revealed likely distinct acoustic patterns
when assessing mood states, possibly reflecting symptom severity. Specifically, sample patterns from
verbal tasks of participants were likely to exhibit a diferent number and duration of pauses with varying
speech rate and mean intraword time (i.e., presence/absence of pressure of speech) as well as potential
diferences in signal frequency and intensity. However, according to existing evidence, no standardized
feature framework is available. Therefore, considering the uncertainty about which features should be
extracted as well as the risk of bias due to potentially missing information, these results suggested the
need to focus on speech segments more in detail, by pre-processing and analyse image data directly
from speech spectrograms for image data classification purposes, possibly corroborating the role of
speech features as digital markers of mood states in people with BD.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and future research</title>
      <p>The current work explored the use of speech signal analysis to map symptom severity in people with
BD when assessing mood states using neural networks. Preliminary results showed relatively adequate
accuracy for prediction, though with varying model performance according both to features selected
and subgroups (e.g., sex). As a whole, combining acoustic signals and NLP can be a feasible, clinically
useful, application in mental healthcare with acoustic features representing novel markers for mood
states.</p>
      <p>Future work will focus on capturing a higher degree of complexity of the underlying data distribution,
by extending in a larger, more diverse sample of people with BD, accounting for potential confounders,
and exploring speech segments more in detail based on speech spectrograms. This would enable a better
understanding of how these methods can operate in real-world settings, particularly with regard to
their potential for integration into clinical practice. Indeed, feature information may be complemented
by feeding the spectrograms directly into the models as input data, using short voice segments to
develop deep learning algorithms from speech spectrograms for classification purposes (e.g., CNN). In
addition, speech signals are often mixed with other signals and both frequency and amplitude are likely
to change over time resulting in non-stationary and non-linear signals. Therefore, empirical mode
decomposition (EMD) related approaches, by breaking the signal down into components that reflect
relevant changes, could ofer valuable insights [ 23, 24, 25]. Consistently, smartphone-based approaches
for speech processing show potential for real-time monitoring (or detection) of mood states in BD
likely relying on ecological momentary assessments with NLP and artificial intelligence (AI) being
promising for smart mental healthcare over time [26, 27, 28]. Based on mHealth technologies, this
approach would help devising human-centered mood remote monitoring based on symptom patterns
from speech possibly with significant clinical impact.
Ecologically valid long-term mood monitoring of individuals with bipolar disorder using speech,
in: Proc. IEEE Int. Conf. Acoust. Speech Signal Process, 2014, pp. 4858–4862. doi:10.1109/ICASSP.
2014.6854525.
[18] S. Kapoor, T. Kumar, Fusing traditionally extracted features with deep learned features from the
speech spectrogram for anger and stress detection using convolution neural network, Multimed
Tools Appl 81 (2022) 31107–31128. doi:10.1007/s11042-022-12886-0.
[19] H. Meng, T. Yan, F. Yuan, H. Wei, Speech emotion recognition from 3D log-mel spectrograms with
deep learning network, IEEE Access 7 (2019) 125868–125881. doi:10.1109/ACCESS.2019.2938007.
[20] A. Z. Antosik-Wójcińska, M. Dominiak, M. Chojnacka, K. Kaczmarek-Majer, K. R. Opara,
W. Radziszewska, A. Olwert, Ł. Święcicki, Smartphone as a monitoring tool for bipolar
disorder: a systematic review including data analysis, machine learning algorithms and predictive
modelling, Int J Med Inform 138 (2020). doi:10.1016/j.ijmedinf.2020.104131.
[21] Y. LeCun, Y. Bengio, G. Hinton, Deep learning, Nature 521 (2015) 436–444. doi:10.1038/
nature14539.
[22] H. Naderi, B. H. Soleimani, S. Matwin, Multimodal deep learning for mental disorders prediction
from audio speech samples, in: 33rd Conference on Neural Information Processing Systems
(NeurIPS), Vancouver, Canada, 2019, pp. 4858–4862. arXiv 1909.01067v5.
[23] P. Marti-Puig, E. Gallego-Jutglà, G. Masferrer, J. Solé-Casals, A New Algorithm for Speech
Enhancement Based on Multivariate Empirical Mode Decomposition, Artificial Intelligence Research
and Development, IOS Press., 2018, pp. 247–255. doi:10.3233/978-1-61499-918-8-247.
[24] C. Sun, H. Li, L. Ma, Speech emotion recognition based on improved masking EMD and
convolutional recurrent neural network, Front Psychol 13 (2023) 1075624. doi:10.3389/fpsyg.2022.
1075624.
[25] U. Souza, J. Escola, T. Vedovatto, L. Brito, R. Lemos, Bidirectional EMD-RLS: Performance analysis
for denoising in speech signal, Journal of Computational Science 74 (2023) 102181. doi:10.1016/j.
jocs.2023.102181.
[26] E. Kerz, S. Zanwar, Y. Qiao, D. Wiechmann, Toward explainable AI (XAI) for mental health
detection based on language behavior, Front Psychiatry 14 (2023) 1219479. doi:10.3389/fpsyt.
2023.1219479.
[27] B. Zhou, G. Yang, Z. Shi, S. Ma, Natural language processing for smart healthcare, IEEE Rev</p>
      <p>Biomed Eng 17 (2024) 4–18. doi:10.1109/RBME.2022.3210270.
[28] W. Hinzen, L. Palaniyappan, The ’L-factor’: Language as a transdiagnostic dimension in
psychopathology, Prog Neuropsychopharmacol Biol Psychiatry 131 (2024) 110952. doi:10.1016/j.
pnpbp.2024.110952.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bolton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Warner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Harriss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geddes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. E. A.</given-names>
            <surname>Saunders</surname>
          </string-name>
          ,
          <article-title>Bipolar disorder: Trimodal age-at-onset distribution</article-title>
          ,
          <source>Bipolar Disord</source>
          <volume>23</volume>
          (
          <year>2021</year>
          )
          <fpage>341</fpage>
          -
          <lpage>356</lpage>
          . doi:
          <volume>10</volume>
          .1111/bdi.13016.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bauer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. A.</given-names>
            <surname>Andreassen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Geddes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. V.</given-names>
            <surname>Kessing</surname>
          </string-name>
          , U. Lewitzka,
          <string-name>
            <given-names>T. G.</given-names>
            <surname>Schulze</surname>
          </string-name>
          , E. Vieta,
          <article-title>Areas of uncertainties and unmet needs in bipolar disorders: clinical and research perspectives</article-title>
          ,
          <source>Lancet Psychiatry</source>
          <volume>5</volume>
          (
          <year>2018</year>
          )
          <fpage>930</fpage>
          -
          <lpage>939</lpage>
          . doi:
          <volume>10</volume>
          .1016/S2215-
          <volume>0366</volume>
          (
          <issue>18</issue>
          )
          <fpage>30253</fpage>
          -
          <lpage>0</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.</given-names>
            <surname>Grande</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Berk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Birmaher</surname>
          </string-name>
          , E. Vieta, Bipolar disorder,
          <source>Lancet</source>
          <volume>387</volume>
          (
          <year>2016</year>
          )
          <fpage>1561</fpage>
          -
          <lpage>1572</lpage>
          . doi:
          <volume>10</volume>
          . 1016/S0140-
          <volume>6736</volume>
          (
          <issue>15</issue>
          )
          <fpage>00241</fpage>
          -
          <lpage>X</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>McIntyre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Berk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Brietzke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. I.</given-names>
            <surname>Goldstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>López-Jaramillo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. V.</given-names>
            <surname>Kessing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Malhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Nierenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Rosenblat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Majeed</surname>
          </string-name>
          , E. Vieta,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vinberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Mansur</surname>
          </string-name>
          , Bipolar disorders,
          <source>Lancet</source>
          <volume>396</volume>
          (
          <year>2020</year>
          )
          <fpage>1841</fpage>
          -
          <lpage>1856</lpage>
          . doi:
          <volume>10</volume>
          .1016/S0140-
          <volume>6736</volume>
          (
          <issue>20</issue>
          )
          <fpage>31544</fpage>
          -
          <lpage>0</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Montgomery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Asberg</surname>
          </string-name>
          ,
          <article-title>A new depression scale designed to be sensitive to change</article-title>
          ,
          <source>Br J Psychiatry</source>
          <volume>134</volume>
          (
          <year>1979</year>
          )
          <fpage>382</fpage>
          -
          <lpage>389</lpage>
          . doi:
          <volume>10</volume>
          .1192/bjp.134.4.382.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Young</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Biggs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. E.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Meyer</surname>
          </string-name>
          ,
          <article-title>A rating scale for mania: reliability, validity and sensitivity</article-title>
          ,
          <source>Br J Psychiatry</source>
          <volume>133</volume>
          (
          <year>1978</year>
          )
          <fpage>429</fpage>
          -
          <lpage>435</lpage>
          . doi:
          <volume>10</volume>
          .1192/bjp.133.5.429.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Harvey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lobban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rayson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Warner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <article-title>Natural language processing methods and bipolar disorder: Scoping review</article-title>
          ,
          <source>JMIR Ment Health</source>
          <volume>9</volume>
          (
          <year>2022</year>
          )
          <article-title>e35928</article-title>
          . doi:
          <volume>10</volume>
          .2196/35928.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Arevian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Malandrakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. R.</given-names>
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. B. Wells</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          <string-name>
            <surname>Miklowitz</surname>
            ,
            <given-names>S. Narayanan,</given-names>
          </string-name>
          <article-title>Clinical state tracking in serious mental illness through computational analysis of speech</article-title>
          ,
          <source>PLoS One</source>
          <volume>15</volume>
          (
          <year>2020</year>
          )
          <article-title>e0225695</article-title>
          . doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0225695</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Faurholt-Jepsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Rohani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Busk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. L.</given-names>
            <surname>Tønning</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vinberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Bardram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. V.</given-names>
            <surname>Kessing</surname>
          </string-name>
          ,
          <article-title>Discriminating between patients with unipolar disorder, bipolar disorder, and healthy control individuals based on voice features collected from naturalistic smartphone calls</article-title>
          ,
          <source>Acta Psychiatr Scand</source>
          <volume>145</volume>
          (
          <year>2022</year>
          )
          <fpage>255</fpage>
          -
          <lpage>267</lpage>
          . doi:
          <volume>10</volume>
          .1111/acps.13391.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Mundt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Vogel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. E.</given-names>
            <surname>Feltner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. R.</given-names>
            <surname>Lenderking</surname>
          </string-name>
          ,
          <article-title>Vocal acoustic biomarkers of depression severity and treatment response</article-title>
          ,
          <source>Biol Psychiatry</source>
          <volume>72</volume>
          (
          <year>2012</year>
          )
          <fpage>580</fpage>
          -
          <lpage>587</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.biopsych.
          <year>2012</year>
          .
          <volume>03</volume>
          .015.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Cummins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. W.</given-names>
            <surname>Schuller</surname>
          </string-name>
          ,
          <article-title>Speech analysis for health: Current state-of-the-art and the increasing impact of deep learning</article-title>
          ,
          <source>Methods</source>
          <volume>151</volume>
          (
          <year>2018</year>
          )
          <fpage>41</fpage>
          -
          <lpage>54</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.ymeth.
          <year>2018</year>
          .
          <volume>07</volume>
          . 007.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Or</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Torous</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Onnela</surname>
          </string-name>
          ,
          <article-title>High potential but limited evidence: Using voice data from smartphones to monitor and diagnose mood disorders</article-title>
          ,
          <source>Psychiatr Rehabil J</source>
          <volume>40</volume>
          (
          <year>2017</year>
          )
          <fpage>320</fpage>
          -
          <lpage>324</lpage>
          . doi:
          <volume>10</volume>
          .1037/prj0000279.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>J. N. de Boer</surname>
            ,
            <given-names>A. E.</given-names>
          </string-name>
          <string-name>
            <surname>Voppel</surname>
            ,
            <given-names>M. J. H.</given-names>
          </string-name>
          <string-name>
            <surname>Begemann</surname>
            ,
            <given-names>H. G.</given-names>
          </string-name>
          <string-name>
            <surname>Schnack</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Wijnen</surname>
            ,
            <given-names>I. E. C.</given-names>
          </string-name>
          <string-name>
            <surname>Sommer</surname>
          </string-name>
          ,
          <article-title>Clinical use of semantic space models in psychiatry and neurology: A systematic review and meta-analysis</article-title>
          ,
          <source>Neurosci Biobehav Rev</source>
          <volume>93</volume>
          (
          <year>2018</year>
          )
          <fpage>85</fpage>
          -
          <lpage>92</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.neubiorev.
          <year>2018</year>
          .
          <volume>06</volume>
          .008.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Vanello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Guidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gentili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Werner</surname>
          </string-name>
          , G. Bertschy,
          <string-name>
            <given-names>G.</given-names>
            <surname>Valenza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lanata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Scilingo</surname>
          </string-name>
          ,
          <article-title>Speech analysis for mood state characterization in bipolar patients</article-title>
          ,
          <source>in: Annu Int Conf IEEE Eng Med Biol Soc</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>2104</fpage>
          -
          <lpage>2107</lpage>
          . doi:
          <volume>10</volume>
          .1109/EMBC.
          <year>2012</year>
          .
          <volume>6346375</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>K.</given-names>
            <surname>Matton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>McInnis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Provost</surname>
          </string-name>
          ,
          <article-title>Into the wild: Transitioning from recognizing mood in clinical interactions to personal conversations for individuals with bipolar disorder</article-title>
          , in: Interspeech, Graz, Austria,
          <year>2019</year>
          . doi:
          <volume>10</volume>
          .21437/Interspeech.2019-
          <volume>2698</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Faurholt-Jepsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Busk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Frost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Vinberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Christensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Winther</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Bardram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. V.</given-names>
            <surname>Kessing</surname>
          </string-name>
          ,
          <article-title>Voice analysis as an objective state marker in bipolar disorder</article-title>
          ,
          <source>Transl Psychiatry</source>
          <volume>6</volume>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1038/tp.
          <year>2016</year>
          .
          <volume>123</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Z. N.</given-names>
            <surname>Karam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Provost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Montgomery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Archer</surname>
          </string-name>
          , G. Harrington,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Mcinnis</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>