<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>H. Lee, N. S. Kim, Y. M. Ahn, Detection of minor and proof of concept, Healthcare Analytics</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.4088/JCP.15M10330</article-id>
      <title-group>
        <article-title>Prediction of Relapse in Adolescent Depression using Fusion of Video and Speech Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christopher Lucasius</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mai Ali</string-name>
          <email>maia.ali@mail.utoronto.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco Battaglia</string-name>
          <email>marco.battaglia@camh.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Strauss</string-name>
          <email>john.strauss@islandhealth.ca</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Szatmari</string-name>
          <email>peter.szatmari@camh.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Deepa Kundur</string-name>
          <email>dkundur@ece.utoronto.ca</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Depression relapse, Multimodality, Inception ResNet, LSTM</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical and Computer Engineering, University of Toronto</institution>
          ,
          <addr-line>Toronto</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Psychiatry, University of Toronto</institution>
          ,
          <addr-line>Toronto</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Division of Child and Youth Psychiatry, Centre for Addiction and Mental Health</institution>
          ,
          <addr-line>Toronto</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The Hospital for Sick Children</institution>
          ,
          <addr-line>Toronto</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Vancouver Island Health Authority</institution>
          ,
          <addr-line>Vancouver</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>2</volume>
      <issue>2022</issue>
      <fpage>2818</fpage>
      <lpage>2826</lpage>
      <abstract>
        <p>This article presents an innovative approach to predicting depression relapse in adolescents. Adolescentsíntensive use of video and voice-based smartphone apps presents a rich, multimodal dataset that can be utilized for this purpose. This work uses a dataset from the Depression Early Warning study conducted at the Center for Addiction and Mental Health. After using a pre-trained Inception ResNet to generate embeddings of video frames, the proposed framework integrates this with synchronized speech data. These embeddings are fused with audio features, resulting in a multimodal dataset. The combined features are processed through a Long Short-Term Memory model and a fully connected network to predict relapse of depression. An average accuracy of 0.80 highlights the efectiveness of the proposed multimodal approach and underscores its potential to efectively predict depression relapse in adolescents.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Depression is a worldwide, prevalent mental health
disorder among adolescents. The recognition and treatment
of adolescent depression hold paramount significance
due to its association with substantial risks, notably
suicide, which stands as the fourth leading cause of death
within this demographic [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Disturbingly, over half of
adolescents who commit suicide are reported to have
been struggling with a depressive disorder [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Beyond
this, depression in adolescents causes profound social
and educational impairments, underscoring the need for
timely intervention. The consequences extend to
heightened rates of smoking, substance misuse, and obesity,
accentuating the urgency of addressing this mental health
concern [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <sec id="sec-2-1">
        <title>Standard mental health diagnoses rely on clinical sur</title>
        <p>
          veys that may be subject to recall bias. This approach also
does not allow for timely interventions [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. To address
these limitations, diverse modalities have been proposed
in the literature for timely mental health assessment and
Machine Learning for Cognitive and Mental Health Workshop
(ML4CMH), AAAI 2024, Vancouver, BC, Canada
∗Corresponding author.
nEvelop-O
features such as heart rate and temperature, as well as
behavioral features such as voice, facial expression, and
gesture. Video chat and gaming are very popular among
youth with statistics reaching 87% in this population [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
However, despite the widespread engagement in these
activities, research exploring the use of video and speech
modalities for the assessment of depression and
prediction of relapse in youth is limited. This work investigates
the use of speech and video for depression relapse
prediction in adolescents. As far as the authors are aware,
it presents the first pipeline for predicting depression
relapse in adolescents using fusion of video- and
speechbased features.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Literature Review</title>
      <p>The use of speech and video analysis for depression
prediction represents an innovative and promising approach
in mental health research. Analyzing speech patterns
and facial expressions can provide valuable insights into
an individual’s emotional and mental state. Below is a
review on the use of speech and video for depression
prediction.</p>
      <sec id="sec-3-1">
        <title>2.1. Speech-based Depression Prediction</title>
        <sec id="sec-3-1-1">
          <title>Several studies demonstrated that voice quality contains</title>
          <p>
            © 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License information about the mental state of a person and vocal
Attribution 4.0 International (CC BY 4.0).
source features can be used as biomarkers of depres- in [11]. The framework combined spatial information
sion severity [
            <xref ref-type="bibr" rid="ref5 ref6">5, 6, 7</xref>
            ]. The work in [8] is based on extracted from the Inception-ResNet-v2 network with a
a cross-sectional and longitudinal study aimed to ex- volume local directional number (VLDN) based dynamic
plore the potential of voice acoustic features as objec- feature descriptor to capture facial motions. The VLDN
tive biomarkers for assessing depression severity and feature map was then fed into a CNN to obtain more
treatment efectiveness. The study identified 30 voice discriminative features. Temporal information was
obacoustic features to be associated with depression such tained using a multilayer Bi-LSTM which integrated the
as Mel-cepstral (MCEP), Mel-scale Frequency Cepstral temporal median pooling (TMP) approach on the
temCoeficients deltas (MFCC-deltas) and Harmonic Model poral fragments of spatial and temporal features. The
Phase Distortion Mean (HMPDM) among others. A neu- performance of this work was benchmarked against the
ral network model based on Neural Architecture Search AVEC2013 and AVEC2014 datasets, and it achieved an
(NAS) was developed for predicting depression severity. MAE of 7.04 and 6.86 on AVEC2013 and AVEC2014,
reGrid search was used to obtain the optimal model ar- spectively.
chitecture which consisted of 4 hidden layers with with Zhou et al. presented a deep regression network called
32 units each. The model achieved a Mean Absolute Er- DepressNet which aimed to learn a visually interpretable
ror (MAE) of 3.137 when predicting depression severity representation of depression from facial images [12].
based on Hamilton Depression (HAMD) Scale. Addi- Their model is based on a CNN with a global average
tionally, a longitudinal study investigated the changes pooling layer which is first trained with facial depression
in voice features after an Internet-based cognitive- data, for identifying salient regions of an input image
behavioral therapy (ICBT) program, revealing four fea- in terms of its severity score based on the generated
detures that significantly decreased: Peak2RMS_kurtosis, pression activation map (DAM). The authors proposed
MFCC_deltas_10_intercept, MFCC_delta_deltas_4_kur- a multi-region DepressNet that combines multiple local
tosis, and MFCC_delta_deltas_9_kurtosis. This indicated deep regression models for diferent face regions to
entheir potential correlation with treatment response and hance recognition performance. The method achieved
improvement in depression. In [9], Vázquez-Romero et.al. an MAE of 6.20 and 6.21 on AVEC 2013 and 2014 datasets,
proposed a method for automatic classification of depres- respectively.
sion using speech and ensemble learning with
Convolutional Neural Networks (CNNs). In the preprocessing 2.3. Speech and Video-based Depression
phase, speech files are transformed into sequences of
logPrediction
spectrograms and randomly sampled to ensure a balance
between positive and negative samples. For the classi- Physiological and psychological studies have identified
ifcation task, multiple CNNs are trained using diferent diferences in speech and facial expressions between
painitializations, and their individual predictions are com- tients with depression and healthy individuals, providing
bined using an ensemble averaging algorithm. The pre- potential cues for automatic depression detection [13].
dictions are then aggregated for each speaker to obtain a Another related work by [14] presented a depression
deifnal decision. The performance of the proposed model tection model that utilizes audiovisual features extracted
was evaluated on the DAIC-WOZ dataset and compared from video logs (vlogs) on YouTube. The model extracts
against the AVEC-2016 models that use support vector eight low-level acoustic descriptors, including loudness,
machine (SVM) classifiers and hand-crafted features, as fundamental frequency (F0), and spectral flux, using the
well as the DepAudionet architecture that consisted of OpenSmile toolkit. These features capture
characterisa 1D-CNN, Long Short-Term Memory (LSTM) cell, and tics such as voice intensity and pitch which have been
fully connected layers. The results demonstrated a rela- found to be relevant in detecting depression. For visual
tive improvement in F1-score of 58.5%, 30.0%, and 10.2% features, the model utilizes a pre-trained face expression
compared to the baseline, DepAudionet, and single 1D- recognition model (FER) to extract emotional information
CNN architecture, respectively. from the vlogs. The proposed eXtreme Gradient
Boosting (XGBoost) depression detection model achieved an
2.2. Video-based Depression Prediction overall performance with an accuracy of 75.85%, recall
of 78.18%, precision of 76.79%, and F1 score of 77.48%.
          </p>
          <p>Behavioral analysis of facial expressions has been stud- The model’s performance was further analyzed based on
ied as a source for eliciting the underlying emotional diferent modalities where the model trained with audio
state [10]. Computer vision methods have been used to features performed better than the model trained with
analyze facial expressions and gestures to predict the un- visual features. The best performance was achieved by
derlying mental health state of users [11]. A framework the model trained on the audiovisual features. The work
for estimating depression levels from video data using a of Othmani et.al. in [15] used deep learning techniques
two-stream deep spatiotemporal network was introduced to recognize depression and predict relapse from audio
and visual cues extracted from videos of clinical inter- interviewed by the coordinator during their initial visit
views. It involves a correlation-based anomaly detection and followup sessions. During recorded Zoom sessions,
framework that compares the audiovisual patterns of the coordinator asked them 10 open-ended questions
depression-free subjects to those of depressed individu- about their past activities and mood, resulting in 2-10
als. The correlation between the audiovisual encoding minutes of video data per session. This dataset was
colof a test subject and a deep audiovisual representation lected as part of an ongoing research study at CAMH and
of depression is computed to monitor depressed subjects is unavailable to the public.
and predict relapse. The approach achieves promising
results, with an accuracy of 80.99% and 82.55% for relapse 4.2. Definition of Relapse
depression prediction on the DAIC-Woz dataset.</p>
          <p>The existing landscape of research on adolescent de- While there are many definitions of relapse in depression,
pression has made significant strides in understanding a commonly accepted one is given by [16] which defines
the onset and symptoms of depression in this age group. a relapse in adolescents as observing a CDRS score of at
However, there is a notable gap in the ability to efec- most 28 during at least 12 weeks of treatment followed by
tively predict depression relapse from audio and video an increase in CDRS to at least 40 for at least two weeks.
modalities. By incorporating synchronized video and The first period of 12 weeks corresponds to a remission
speech data, this research captures a broader spectrum of stage where the depressed adolescent does not exhibit
behavioral and emotional cues that might signify impend- symptoms but has not yet completed treatment. The
ing relapse in adolescents. The synchronization ensures period of two weeks corresponds to a depressed episode.
that both modalities are aligned, allowing for a detailed In this study, there can be at least a three month break
examination of facial expressions, body language, and between followup visits. Hence, the timing aspect of
Kenspeech patterns simultaneously. nard’s definition must be accordingly modified to adhere
to the provided data. This work proposes a definition of
relapse as a period of at least one visit with a CDRS score
3. Problem Formulation of at most 40 followed by one visit with a CDRS score of
at least 40.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>4.3. Pipeline</title>
        <p>There are three main stages that make up the methods of
this pipeline. The first consists of preprocessing the video
and audio data and organizing them such that the two
modalities are aligned and the labels are balanced. The
second involves training models on random subsets of the
training data. In the final stage, the final model that was
trained on the training set is evaluated on multiple test
sets, and the performance metrics are averaged across
each set. A diagram summarizing the pipeline is shown
in Figure 1.
4.3.1. Stage 1: Data Preparation</p>
        <sec id="sec-3-2-1">
          <title>Each video interview is divided into segments where</title>
          <p>only the participant is speaking. Since the interviews are
conducted via Zoom, the videos are also cropped such
that only the participant’s face is visible. Several
spectral features are extracted from the audio data using the
Python package libRosa [17]. They include the MFCCs,
fundamental frequency, chromagrams, power spectral
density, and spectral rollof. These features are computed
over a rolling window that is applied across the video.
The amount of overlap is chosen such that the number of
windows matches that of the video frames and are evenly
spread out across the video.</p>
          <p>Our work aims to classify fused video and speech features
for the classification of data that is measured before a
relapse event. This entails a binary classification task
where the two classes include “relapse sometime in the
future” and “non-relapse”. This problem is significantly
diferent from detecting the presence of depression or
predicting a certain depression rating scale score. The
problem of relapse prediction is more complex since it
involves the direct prediction of a clinical event within
a population of adolescents who are already diagnosed
with Major Depressive Disorder.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methods</title>
      <sec id="sec-4-1">
        <title>This work uses a dataset that is collected as part of the</title>
        <p>depression early warning study that was run in the
Centre for Addiction and Mental Health (CAMH). It includes
80 video interviews collected from 52 adolescents aged
12-21 who were all diagnosed with Major Depressive
Disorder.</p>
        <sec id="sec-4-1-1">
          <title>4.1. CAMH Dataset</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>All participants had an initial baseline visit followed by</title>
        <p>up to 7 followup visits, each spaced apart by 3-12 months.</p>
        <p>During each visit, participants were assessed by a trained
research coordinator and psychiatrist, providing
psychiatric evaluations of their depressive states via the
Children’s Depression Rating Scale (CDRS). Participants were</p>
        <p>Spectral
Feature
Extraction</p>
        <p>Train Fold 1
Train Fold 2
Train Fold N
Test Fold
Inception
ResNet
LSTM</p>
        <p>In the provided dataset, there is a significant class ing an LSTM and a fully connected network. The LSTM
imbalance where the non-relapse data is heavily over- is used to process 16 consecutive frames of features at a
represented (96.25% non-relapse). In order to not bias time, and the resulting hidden state is then fed into the
the training of the models and the evaluation metrics (de- fully connected network to be classified as either relapse
scribed in the next two sections), several training folds are or non-relapse. During the training process, random
segprepared alongside a test fold. The folds are constructed ments of 16 frames are sampled from the training video
by first randomly selecting a proportion of relapse video clips in order to not bias the training of the network
clips to use in the test fold. This proportion is chosen towards a certain class.
to be 30%, and it is computed based on the number of The training process is applied to each train fold, and
frames within each video clip. A random selection of within a given fold, it is repeated for eight epochs.
Afnon-relapse clips are chosen to match the number of ter the AudioVisual Network is trained on a given fold,
frames of the relapse ones (rounded to the nearest whole its saved parameters are used to continue training the
number of clips). This completes the test fold which network on a new fold. This is repeated until all train
is reserved for Stage 3 of the pipeline. The rest of the folds are exhausted. This allows the network to train on
relapse subjects are assigned to be used by train folds the entire training dataset while still keeping the classes
in Stage 2. Non-relapse video clips are randomly sam- relatively balanced.
pled without replacement where the number of clips is
selected to match the number of frames of the relapse 4.3.3. Stage 3: Evaluation of Models
subjects. Each random sample of non-relapse clips makes
up another train fold, and this process is repeated until
all non-relapse clips are used.</p>
      </sec>
      <sec id="sec-4-3">
        <title>After the AudioVisual Network is trained, the architec</title>
        <p>ture (+InceptionResNet), is evaluated on the test fold that
was reserved in Stage 1. A receiver operating
character4.3.2. Stage 2: Training of Models istic (ROC) analysis is carried out on the predictions and
ground truth labels. The optimal threshold of the ROC
The video frames are fed into an InceptionResNet model curve is selected by choosing the point that maximizes
that was pre-trained on VGGFace2 [18], a large-scale face the diference between the true and false positive rats.
dataset. The resulting embeddings from this network are This threshold is used to compute the MAE.
then fused with the spectral features of the audio data. The entire process of training the models and
evaluatThe resulting fused features are then fed into a neural ing the final one on a test fold is carried out for 10 sets
network module (named AudioVisual Network) contain- of folds. This is to ensure that the reported metrics are
not biased toward a certain set of subjects. The resulting
performance metrics are averaged across all of the test
folds.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Significance</title>
      <p>Table 1 shows the results of evaluating the trained model
on the 10 test folds. Each accuracy and MAE measure
was reported after finding the optimal threshold of the
ROC curve.</p>
      <p>An average accuracy of 0.80 shows that video and
speech data are relatively promising in the prediction of
relapse in adolescent depression. In previous work by
Othmani et al. [15], the authors also predicted relapse
of depression using video and speech data. Similar to
our work, they also yielded accuracies at around 0.8. To
the best of our knowledge, this is the only other work
that used video and speech to predict relapse of
depression. Our work diferentiates from Othmani et al. in two
significant ways: 1) our study focuses on adolescents
and 2) the source of our data is from non-clinical
interviews. These interviews allow for more conversational
topics that may better mimic a real-life situation in an
adolescent’s everyday life. While the target population
for this work includes adolescents, this framework can
be extended to other depressed populations.
of gender-based analysis is a notable limitation,
potentially overlooking important nuances in how depression
manifests across diferent genders.</p>
      <p>Future work in predicting depression from audiovisual
features will prioritize the development of gender and
context aware models. Moreover, given the longitudinal
nature of the study, a promising avenue for future work
is to exploit the temporal nature of data to track changes
in audiovisual features over long periods of time.
Employing an overarching time series model could enhance
the understanding of the dynamic nature of depression,
allowing for the development of more adaptive and
personalized prediction models.</p>
      <p>Another way to extend this work is to combine other
objective sources of data that can be collected
simultaneously with video and speech. One such modality includes
wearable technologies, and there have been several
studies on using them for the prediction of depression [19].
Using similar techniques, it may be possible to fuse
audiovisual features and those derived from wearables to
create a more robust predictor of adolescent depression
relapse. Finally, we intend to evaluate our work using
publicly available audio/video depression datasets such
as AVEC.
Fold
1
2
3
4
5
6
7
8
9
10
Average</p>
      <p>MAE
0.077
0.28
0.21
0.23
0.12
0.26
0.17
0.23
0.17
0.27
0.21</p>
      <p>Accuracy
0.92
0.72
0.79
0.77
0.88
0.74
0.83
0.77
0.83
0.73
0.80</p>
    </sec>
    <sec id="sec-6">
      <title>6. Limitations and Future Work</title>
      <p>Predicting depression from audiovisual features
encounters various challenges. The subjectivity of depression
labels and the heterogeneous nature of this condition make
it dificult to develop a universally applicable model.
Additionally, there may be ethnic and cultural biases in the
data that may have impacted the model’s
generalizability. This work did not consider the context within which
interviews were conducted. Furthermore, the exclusion</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>World</given-names>
            <surname>Health</surname>
          </string-name>
          <string-name>
            <surname>Organization</surname>
          </string-name>
          , Suicide,
          <year>2023</year>
          . URL: https://www.who.int/news-room/fact-sheets/ detail/suicide.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Thapar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Collishaw</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Pine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Thapar</surname>
          </string-name>
          , Depression in adolescence,
          <source>The Lancet</source>
          <volume>379</volume>
          (
          <year>2012</year>
          )
          <fpage>1056</fpage>
          -
          <lpage>1067</lpage>
          . URL: https://www.ncbi.nlm.nih. gov/pmc/articles/PMC3488279/. doi:https://doi. org/10.1016/s0140-
          <volume>6736</volume>
          (
          <issue>11</issue>
          )
          <fpage>60871</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N. H.</given-names>
            <surname>Goldhaber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Hekler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fergerson</surname>
          </string-name>
          ,
          <article-title>Evaluating the mental health of physician-trainees using an sms text message-based assessment tool: Longitudinal pilot study</article-title>
          ,
          <source>JMIR Formative Research</source>
          <volume>7</volume>
          (
          <year>2023</year>
          )
          <fpage>e45102</fpage>
          -
          <lpage>e45102</lpage>
          . doi:https://doi.org/10.2196/ 45102.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Summerfield</surname>
          </string-name>
          ,
          <article-title>How many kids in canada are connecting with video games?</article-title>
          ,
          <year>2023</year>
          . URL: https://mediaincanada.com/
          <year>2023</year>
          /01/30/ how-many
          <article-title>-kids-in-canada-are-connecting-with-video-games/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.-Z.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Liu</surname>
          </string-name>
          , Y.
          <string-name>
            <surname>-X. Wu</surname>
            ,
            <given-names>Y.-L.</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Z.-X.</given-names>
          </string-name>
          <string-name>
            <surname>Tian</surname>
            ,
            <given-names>Z.-R.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.-L.</given-names>
          </string-name>
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>S.-P.</given-names>
          </string-name>
          <string-name>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Vocal acoustic features as potential biomarkers for identifying/diagnosing depression: A cross-sectional study</article-title>
          ,
          <source>Frontiers in Psychiatry</source>
          <volume>13</volume>
          (
          <year>2022</year>
          ). doi:https://doi.org/10.3389/fpsyt.
          <year>2022</year>
          .
          <volume>815678</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Shin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. I.</given-names>
            <surname>Cho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H. K.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Rhee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>