<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Multimodal Fusion Techniques to Enhance Voice Disorder Diagnoses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Politecnico di Torino</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Turin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qingqing Liu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabriele Ciravegna</string-name>
          <email>gabriele.ciravegna@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alkis Koudounas</string-name>
          <email>alkis.koudounas@polito.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tania Cerquitelli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elena Baralis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Medical AI</institution>
          ,
          <addr-line>Pathological voice disorder, Transformers, Modality Fusion, Multimodal learning</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>detection performance. The cross-attention technique</institution>
          ,
          <addr-line>in</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Voice disorders constitute a significant health concern, with an annual prevalence of approximately 7% among the adult population, adversely afecting patients' quality of life, encompassing both social and occupational functioning. Also, the majority of diagnostic methodologies continue to depend on invasive techniques, whereas non-invasive automated diagnostic approaches have not been extensively investigated yet. This study introduces a transformer-based method for detecting voice disorders aimed at enhancing detection eficacy through a multimodal fusion strategy. Specifically addressing two distinct types of voice recordings - extracted from sentences reading and vowels emissions -- we devised and assessed five multimodal fusion strategies across three stages: early, mid, and late. Our experimental findings indicate that the cross-attention mid-fusion method harnesses the benefits of both data types, and it achieves a detection accuracy of 0.885 and a macro F1 score of 0.843 on an internal dataset. These results represent an improvement of +.03 to +.06 in accuracy and +.02 to +.05 in macro F1 score when compared to unimodal models (trained on sentence or vowel data only). This study represents an advancement for an efective non-invasive detection of voice disorders and provides insights for clinical practice.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Voice disorders have a significant impact on people’s lives,
with 7% of adults sufering from them each year. They can
lead to communication dificulties, reduced work
productivity (7.4 work days lost per year on average), and even career
changes, with 4% of patients reporting a career change due
to voice problems [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. There are many types of voice
disorders, including but not limited to murmurs, vocal cord
dysfunction, and other voice problems caused by
neurological diseases, and their early and accurate diagnosis is crucial
for efective treatment [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Although traditional diagnostic techniques such as
laryngoscopy and speech assessment are widely used clinically,
they have significant limitations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. First, these diagnostic
methods are very invasive and may cause discomfort to
the patient, thus afecting the experience particularly for
patients requiring several investigations and recurrent
controls (e.g. cancer patients) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Secondly, these technologies
often rely on expensive equipment and highly specialized
operators, which limits their accessibility in resource-poor
settings. Thirdly, traditional methods rely on doctors’
subjective judgments and sufer from subjective bias in
evaluation results. Finally, these methods are mostly used for
diagnosis when symptoms are evident rather than as
proactive preventive screening tools, limiting their role in the
early detection of voice disorders [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        The development of artificial intelligence technology [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
especially the application of deep learning in the field of
audio and sound processing, provides new possibilities for
overcoming the above challenges [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. By enabling
automated, non-invasive, eficient diagnostics, deep learning
methods can lower diagnostic costs and reduce the need for
professionals, making detection more accessible and
accurate. In addition, these technologies can be integrated into
portable devices or mobile applications for active screening
Published in the Proceedings of the Workshops of the EDBT/ICDT 2025 Joint
Conference (March 25-28, 2025), Barcelona, Spain, as part of the DARLI-AP
workshop held in conjunction with the EDBT/ICDT 2025 conference.
∗Corresponding author.
tion in enhancing detection performance. These findings
highlight the feasibility of multi-modal transformer-based
models in clinical applications and lay a solid foundation for
further advancement of automatic voice disorder detection.
      </p>
      <p>The rest of the paper is organized as follows. In Section 2
we first review the relevant research on voice disorder
detection and analyze the main challenges of existing methods in
application. Section 3 describes the proposed
Transformerbased method in detail, focusing on diferent multimodal
CEUR</p>
      <p>ceur-ws.org
fusion strategies. In Section 4 we introduce the
experimental setup and evaluation methods, while in Section 5 we
provide a comprehensive analysis and interpretation of the
experimental results. Finally, in Section 6 we summarize the
significance of our research, explore potential limitations,
and provide suggestions for future research directions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>This section reviews the relevant research on voice disorder
detection and provides a theoretical basis for the tools and
methods used in subsequent sections. The discussion
focuses on the evolution of voice feature analysis techniques,
the application of classifiers in detecting voice disorders,
and the latest progress in data augmentation and fusion
models.</p>
      <sec id="sec-2-1">
        <title>2.1. Automatic Voice Disorder Detection</title>
      </sec>
      <sec id="sec-2-2">
        <title>Methods</title>
        <p>
          Traditional voice disorder detection methods rely on
artiifcial feature engineering, that is, extracting acoustic
features such as Mel-frequency cepstral coeficients (MFCC),
pitch jitter, and amplitude shimmer from speech signals
[
          <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
          ]. These features, rooted in digital signal
processing and speech science, have long been the cornerstone of
voice analysis. Using these manual features, researchers
rely on shallow learning models such as support vector
machines (SVMs) and multi-layer perceptrons (MLPs), which
perform well in voice disorder detection problems in
relatively simple or well-controlled environments [
          <xref ref-type="bibr" rid="ref14 ref15 ref16">14, 15, 16</xref>
          ].
However, the complexity of pathological voice features and
the diversity of real-world scenarios have revealed the
limitations of these traditional methods, particularly in terms
of adaptability and generalization [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]
        </p>
        <p>
          The advent of deep learning has transformed voice
disorder detection, as it can automatically extract features from
raw speech signals. Unlike traditional methods that rely on
handcrafted features, deep learning models such as
convolutional neural networks (CNNs) and recurrent neural
networks (RNNs) can learn more abstract and comprehensive
feature representations directly from data [
          <xref ref-type="bibr" rid="ref12 ref16 ref18 ref19 ref20">12, 16, 18, 19, 20</xref>
          ].
CNNs excel at capturing local patterns, while RNNs excel at
modeling temporally related patterns, making them more
suitable for voice pathology analysis, particularly when
employed together.
        </p>
        <p>
          Recently, transformer-based architectures have made
breakthroughs in automatic speech recognition and related
tasks [
          <xref ref-type="bibr" rid="ref21 ref22 ref23 ref24 ref25 ref26">21, 22, 23, 24, 25, 26</xref>
          ]. These models use self-attention
mechanisms to capture short and long-range dependencies
at the same time, thus performing well in processing
complex speech patterns [
          <xref ref-type="bibr" rid="ref11 ref27 ref8">8, 11, 27</xref>
          ]. Among them, Wav2Vec2’s
end-to-end modeling capability [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] combines a
convolutional encoder for extracting potential speech
representations, a transformer-based context network for capturing
long-distance dependencies, and a quantization module for
self-supervised learning, further simplifying the feature
extraction process. This architecture enables the eficient and
accurate analysis of voices under various conditions.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.2. Multimodal fusion</title>
        <p>
          In voice analysis, multi-modality refers to input data
extracted from diferent data sources or forms of information
[
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. For example, people chatting, singing, reading, or
performing particular sound patterns are all typical
modalities. The information provided by each modality may be
diferent and complementary, and a single modality often
cannot fully capture pathological features. Therefore, by
fusing data from diferent modalities, we can have a more
comprehensive understanding of the pathological condition,
thus improving the accuracy and robustness of detection.
        </p>
        <p>
          In multimodal fusion, there are three main strategies:
early fusion, mid-level fusion, and late fusion [
          <xref ref-type="bibr" rid="ref28 ref29">28, 29</xref>
          ]. Early
fusion combines features from diferent modalities into a
vector before feeding them into a model. Mid-level fusion
integrates data at an intermediate stage, allowing for more
lfexibility in capturing deeper correlations while
maintaining some distinctions between modalities. Late fusion trains
separate models for each modality and combines their
predictions via an aggregation function such as average voting,
weighted voting, or using a meta-classifier.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>2.3. Shallow approach to Multi-modal</title>
      </sec>
      <sec id="sec-2-5">
        <title>Fusion for Voice Disorder detection</title>
        <p>
          The research by Koudounas et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] proposed an
end-toend method based on a transformer, which directly
processes the original audio signal. To address the challenges
posed by diferent recording types (such as sentence
reading and sustained vowel utterances), they used a shallow
mixture of experts (MoE) [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] framework to optimize the
prediction alignment across recording types. Experimental
results show that the method improves the single-modality
approach in speech pathology detection and classification
tasks, and achieves good performance on public and private
datasets. However, this study mainly focuses on synthetic
data and the MoE framework, and lacks in-depth exploration
of multimodal fusion strategies.
        </p>
        <p>
          Building on this, our study introduces a systematic study
of multimodal fusion strategies in voice pathology
detection. We focus on early, mid, and late fusion methods,
especially mid and late fusion, because these two methods
have greater flexibility and can capture deeper correlations
between modalities. Compared with the method of [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], our
study explores fusion strategies in more detail and
demonstrates how mid-fusion strategies are the best multimodal
approach in this domain to improve model generalization.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <p>This section outlines our contributions to multimodal
fusion strategies, emphasizing the mathematical formulation
of the problem and the model architecture. Specifically, we
introduce early, mid-level, and late fusion strategies in
transformer architecture that integrate multiple modalities for
robust prediction.</p>
      <p>
        In this study, we used two speech-based modalities to
solve the voice pathology detection task, each capturing
voice characteristics. The first modality  1 represents the
original features extracted from the sentence reading
recording, while  2 represents the features extracted from the
second modality, the sustained vowel pronunciation recording.
Given a multi-modal architecture  , we input the raw audio
waveforms into the Wav2Vec2 model [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] to combine the
feature extraction for the diferent modalities. The model
then outputs the probabilities  ,̂ which are used to produce
the final classification result.
extraction, the extracted embeddings are normalized to
ensure consistency across modalities, then concatenated as
done in the early fusion approach (Eq.4), but at a deeper
feature level, as shown in Figure 2. Finally, the combined
feature vector goes through a dimension reduction layer to
ift the input size of the subsequent transformer encoder.
where:
ℎ1 =  1( 1),
      </p>
      <p>ℎ2 =  2( 2)</p>
      <p>= [ℎ1; ℎ2]
 =̂ softmax ((  ))
spectively.</p>
      <p>modalities.
• ℎ1 and ℎ2: high-dimensional embeddings extracted
from modalities  1 and  2 using CNN extractor,
re•   : Concatenated feature embeddings from both
•  represents the transformer encoder layers.
•  : a 1-second silence padding between the two</p>
      <sec id="sec-3-1">
        <title>3.2.2. Cross-Attention</title>
        <sec id="sec-3-1-1">
          <title>3.1. Early fusion strategies</title>
          <p>The early fusion strategy connects the raw features of the
two modalities into a unified input representation. In this
stage, we first truncate all audio samples to a uniform length
to standardize the input length and eliminate the bias caused
by the diference in sample length. Then, we directly
concatenate the raw audio features from the two modalities.
Specifically, the raw audio features from the two
modalities are concatenated in a fixed order: modality  1 ifrst,
then modality  2. To further distinguish the two of them,
a 1-second silence ( ) padding is inserted between the two,
providing a clear boundary for the model (Eq.2). After
concatenation, the generated unified features are fed as input
into the pre-trained Wav2Vec2 model for prediction. Fig.1
visually depicts this process.</p>
          <p />
          <p>= [ 1; ;  2]
 =̂ softmax ( ( 
))
(1)
(2)</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.2. Mid-level fusion strategies</title>
          <p>In the mid-level fusion strategy, feature fusion is performed
after CNN encoding but before the features are fed into the
transformer encoder. This approach combines
modalityspecific features in a shared representation space,
allowing the model to leverage interactions between modalities
for more robust predictions. We will analyze two
diferent fusion strategies: concatenated embedding and
crossattention.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2.1. Concatenated embeddings</title>
        <p>In the concatenated embedding strategy, features are first
extracted from each modality using a separate CNN layer
and mapped to the same vector space (Eq.3). We thus
decompose the network into the composition of two modules
 =  ∘</p>
        <p>
          , where  is the CNN-based feature extractor, while
 represents the transformer encoder layers. After feature
(3)
(4)
(5)
(6)
(7)
(8)
The cross-attention mechanism [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] dynamically captures
interactions between modalities by computing attention
weights based on the relationship between the Query ( ),
Key ( ), and Value ( ) matrices. This allows the model to
focus on important features across modalities.
        </p>
        <p>First, given input feature matrices ℎ1 and ℎ2 of the two
modalities, we generate  ,  , and  through linear
transformation,</p>
        <p>= ℎ 1  ,  = ℎ 2  ,  = ℎ 2 
Here,   ,   , and   are learnable weight matrices for the
query, key, and value, respectively. Next, we calculate the
attention matrix  between the Query ( ) and the Key ( ) by
measuring their similarity, then normalized using softmax.
The attention weight is used to perform a weighted sum of
the Value  to generate output features  :
 =
softmax (
 =</p>
        <p>√ 
⊤
)
where:
•  is the general attention matrix.</p>
        <p>tion factor used for scaling.</p>
        <p>•   is the dimension of the key, √  is the
normalizaAs illustrated in Figure 3, cross-attention is computed in
both directions to efectively capture interactions between
the two modalities.</p>
        <p>1. We use ℎ1 as the query and ℎ2 as the key and value
to compute the attention (Eq.9).
2. We reverse the roles of the modalities and use ℎ2 as
the query and ℎ1 as the key and value (Eq.10).</p>
        <p>Finally, the outputs of the cross-attentions from both
directions,  1→2 and  2→1, are concatenated to form a unified
representation,  fused, as shown in Eq.11.</p>
        <p>This concatenation helps to merge the information from
both modalities in a unified feature space. The fused features
are then processed through a shared fusion layer before
passing them to a transformer encoder for deeper feature
extraction and ultimately classification.</p>
        <p>1→2 =  1→2 2 = CrossAttention(ℎ1, ℎ2)
 2→1 =  2→1 1 = CrossAttention(ℎ2, ℎ1)
 fused = [ 1→2;  2→1]
 =̂ softmax ((  
)
(9)
(10)
(11)
(12)</p>
        <sec id="sec-3-2-1">
          <title>3.3. Late fusion strategies</title>
          <p>While mid-level fusion captures fine-grained interactions
between modality-specific embeddings, late fusion is
performed at the decision level, allowing each modality to be
optimized independently and then integrated into a unified
prediction. This approach allows each model to focus on its
specific modality before being integrated, although it
doubles the size of the final model. Two late fusion techniques
are employed in our study.</p>
          <p>Simple average In this approach, the outputs of the two
models,  1̂ and  2̂, are combined by taking their simple
average, as illustrated in the top part of Figure 4. This
strategy assumes that both models contribute equally to the
ifnal prediction. The combined output  l̂ate1 is computed as
follows:</p>
          <p>l̂ate1 = 21 ( 1̂ +  2̂) (13)
where  1̂ and  2̂ are the probability distributions produced
by the two individual models.</p>
          <p>This fusion method is simple and it is computationally
eficient as it avoids any extra parameters.</p>
          <p>Mixture of Expert As a second late fusion strategy, we
employ a shallow mixture of experts (MoE) to combine the
outputs of two independent models and improve the overall
performance of the system. Unlike the simple averaging
method, this approach assigns weights to each model’s
predictions based on how relevant they are to the final output.</p>
          <p>As shown in Figure 4, we use a simple multi-layer
perceptron (MLP) configured with a single hidden layer to predict
weights to combine the outputs of each model. The
input layer of the MLP is a probabilistic concatenation of the
two modalities ( 1̂,  2̂), and the output layer applies a
softmax function to ensure that the sum of all model weights
is 1 (Eq.14). During inference, the final prediction is
computed using the weights to combine the contributions of
both models (Eq. 15). This approach improves the system’s
performance on unseen data while maintaining a simple
architecture.</p>
          <p>[ 1,  2] = softmax (MLP([ 1̂ ;  2̂]))
 l̂ate2 =  1 ⋅  1̂,test +  2 ⋅  2̂,test
(14)
(15)
Here:
• [ 1,  2]: Weights learned from the concatenated
outputs  1̂ and  2̂ on the validation set.
•  1̂,test,  2̂,test: Predicted probabilities from the two
models on the test set.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>This section provides an overview of the datasets and
preprocessing methods used in our experiments, followed by a
detailed description of the training setup to ensure
reproducibility.</p>
      <p>All experiments were conducted in a cloud-based
environment equipped with a Tesla P100-PCIE-16GB GPU1.</p>
      <p>Details of the software environment can be found in the
project repository2.</p>
      <sec id="sec-4-1">
        <title>4.1. Dataset</title>
        <p>
          IPV The Italian Pathological Voice (IPV) dataset is a novel
and diverse resource designed specifically for voice
pathology research, currently unpublished and introduced in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
Collected from participants in Italian otolaryngology and
voice therapy clinics, the dataset includes both healthy
individuals and patients with varying degrees of voice disorders.
All recordings were conducted under strict standardization
protocols in quiet environments, ensuring high-quality
samples with a signal-to-noise ratio exceeding 30 dB and a fixed
microphone distance of 30 cm.
        </p>
        <p>
          The dataset comprises two modalities: sustained
phonation of the vowel /a/ (SV) and reading of five phonetically
balanced sentences (CS) adapted from the Italian version
of CAPE-V [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]. Each sample includes detailed health
condition notes and diagnoses from experienced physicians.
Table 1 provides a detailed summary of the dataset
characteristics, including sample distribution, record length, and
modal information.
        </p>
        <p>Audio Preprocessing To ensure the consistency of audio
duration and facilitate comparison, we cropped the samples
in the datasets to fixed lengths: CS samples were cropped
to 19 seconds, and SV samples were cropped to 18 seconds.
These lengths are designed to cover approximately 90% of
the samples in each modality, ensuring that most voice
1We gratefully acknowledge the computational resources provided by
Kaggle (https://www.kaggle.com/) for this research. We also appreciate
the early-stage support from HPC@Polito (http://www.hpc.polito.it).
2Github repository: github.com/multimodal_pathologies_prediction
information is preserved for efective model training while
reducing the impact of outlier samples that are too long. For
samples shorter than the fixed lengths, zero-padding was
applied to extend them to the required duration.</p>
        <p>Then, the audio data was standardized using the
predeifned processor provided by the Wav2Vec2 framework. The
processor first resamples the audio data to 16kHz to ensure
compatibility with the framework, and reduce
computational overhead. Then the converted feature representation
can not only efectively capture the key information in the
speech signal, but also provide consistent and eficient input
features for the model to support subsequent training tasks.</p>
        <p>In order to avoid issues with the imbalance of
pathological voice data (healthy samples are less than pathological
samples), a stratified sampling method was used in the data
division process to ensure proportional representation of
healthy and pathological samples across all splits. We
divided the data into training, validation, and test sets in a
ratio of 8:1:1 to ensure fair and reproducible evaluations.
The test set was first separated using a fixed random seed.
Subsequently, the training and validation sets were further
split using three diferent random seeds to create multiple
splits. The final results are calculated by averaging the
performance metrics over these splits to ensure the robustness
and reliability of the evaluation.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Baselines</title>
        <p>
          To verify the efectiveness of our proposed method and
provide a comparison, we designed a series of traditional
baseline models, including the classic multi-layer
perceptron (MLP) and a lightweight convolutional neural network
(MobileNetV2 [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]) based on transfer learning. These
baseline models are trained based on traditional audio features
to evaluate the performance of diferent model architectures.
In contrast, the unimodal model based on the Wav2Vec2
processor directly processes the audio waveform to extract
features, reflecting the advantages of end-to-end methods.
        </p>
        <p>In the feature extraction process of the baseline model,
the audio data is uniformly sampled to 16kHz and truncated
to a fixed maximum duration to ensure sample consistency.
We extract 40-dimensional MFCC features through librosa,
transpose them into a time-step sequence form, and
uniformly zero-fill the feature sequence. At the same time, a
padding mask is generated to distinguish between real data
and padding parts. The following is the specific design of
the two baseline models.</p>
        <p>
          MLP is designed with two fully connected layers
containing 50 hidden units, using the ReLU activation function to
extract high-dimensional features, aggregating the time
dimension information through the global average pooling
layer, and finally performing binary classification through
the Softmax output layer. The training process uses the
Adam optimizer with a learning rate of 0.01, a batch size of
16, and an early stopping strategy to prevent overfitting.
2D-CNN The audio features are converted to 2D images
by repeating a single channel to RGB three channels to fit
the input requirements of the pre-trained model. We load
the pre-trained weights (ImageNet [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]) of MobileNetV2
[
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], remove the top classification head, and add a global
average pooling layer, a 512-unit fully connected layer, and
a Softmax classification layer. Dropout is added to the top
network to improve generalization, and the pre-trained
feature extraction part is fine-tuned. Two fine-tuning strategies
are used: full fine-tuning and head-only fine-tuning. In full
ifne-tuning, all layers of MobileNetV2 are updated during
training to maximize performance optimization; while in
head-only fine-tuning, only the newly added classification
head is trained, while the pre-trained feature extraction
layer is frozen to retain the common features learned from
ImageNet. The training hyperparameters of both strategies
are consistent with the MLP model.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Training Procedure</title>
        <p>Our method is based on a pre-trained Wav2Vec2.0 model
(trained on the LibriSpeech 960-hour dataset) and evaluates
three fusion strategies on the IPV dataset: early fusion,
midlevel fusion, and late fusion.</p>
        <p>Early fusion is accomplished by directly concatenating the
original audio of CS and SV, and adding 1 second of silence
(38 seconds) after the total length of the audio to avoid
feature loss. The concatenation is performed on the same
individual. The concatenated audio signals are uniformly
processed in a Wav2Vec2.0 processor to ensure consistency
in feature extraction. Mid-level fusion is based on 2
finetuned Wav2Vec2.0 models, and global feature modeling is
achieved through a shared Transformer encoder (initialized
with sentence reading, SV to a single modality with vowel pronunciation. Values spanning both columns refers to modality
fusion methods. Bold values indicate the best performance for a given metric.</p>
        <p>Modality
Single
Multi</p>
        <p>Method
MLP
Wav2Vec2
Early Fusion
2D-CNN (Train all layers)
2D-CNN (Fine-tune classify head)
Mid (Concatenated Embeddings)
Mid (Cross Attention)
Late (Simple Average)
Late (MoE)</p>
        <p>Accuracy</p>
        <p>Macro F1</p>
        <p>CS
.801±.011
.667±.011
.789±.019
.859±.029</p>
        <p>SV
.750±.057
.673±.000
.782±.048
.827±.000</p>
        <p>CS
.767±.022
.400±.004
.765±.021
.837±.038</p>
        <p>SV
10 hidden nodes) that determines modality weighting based
on the probabilities from the training and validation sets.</p>
        <p>All experiments above were completed within 50 training
rounds (epochs), and using fixed random seed to ensure
the reproducibility of the results. The AdamW optimizer
(weight decay = 0.01) was used for all experiments. A
linear learning rate scheduler is used to optimize the learning
rate adjustment. The scheduler reduced the learning rate
linearly over the total number of training steps, with no
warm-up steps. Initial learning rates are optimized by
manual adjustment, using 1e-5 for single modality and
concatenated fusion and 6e-6 for cross-attention fusion. To address
class imbalance, a weighted cross-entropy loss function was
applied, with class weights computed based on the
training dataset’s label distribution. The batch size was set to
8, and an early stopping strategy with the patience of 10
epochs was used to terminate training when the validation
performance plateaued. More experimental details and
hyperparameter configurations can be found in the GitHub
repository of the article.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Evaluation Metrics</title>
        <p>To evaluate the performance of the model in the voice
disorder detection task, we used two key metrics:</p>
        <sec id="sec-4-4-1">
          <title>Accuracy</title>
          <p>Accuracy measures the proportion of correctly
predicted samples to the total number of samples, providing
an overall assessment of classification performance:
Accuracy =</p>
          <p>Number of Correct Predictions</p>
          <p>Total Number of Samples
(16)
While accuracy is a useful general metric, it can be less
informative in imbalanced datasets.</p>
        </sec>
        <sec id="sec-4-4-2">
          <title>Macro F1-Score</title>
          <p>To better evaluate performance across
imbalanced classes, we adopted the macro-average F1 score,
which calculates the F1 score for each class and then
averages them:
 1 = 2 ×</p>
          <p>Precision ⋅ Recall
Precision + Recall</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>In this section, we analyze and interpret the experimental
results by focusing on two key aspects: comparing
baseline models to assess their efectiveness as reference, and
evaluating diferent fusion models to explore their ability
to integrate multimodal information and improve
generalization to unseen data. By systematically studying these
factors, we aim to highlight the strengths and limitations
of the proposed approach and provide insights for future
improvements.</p>
      <sec id="sec-5-1">
        <title>Benchmark comparison</title>
        <p>of the performance of unimodal baseline models for voice
disorder detection on the IPV dataset.</p>
        <p>As expected, Wav2Vec2 achieved the best results among
the four baseline models, with accuracy of .859 and .827
in CS and SV modes, respectively, and .837 and .793 for F1
macro, respectively. The superior performance of Wav2Vec2
underscores the benefits of self-supervised pre-training on
large-scale audio data. This means that the model does not
need to be trained from scratch, but through pre-training
and transfer learning capabilities, it can have audio features
with good generalization capabilities, even with a small
amount of labeled data. Moreover, it benefits of the attention
mechanism which better extract relevant features from long
sequence of data.</p>
        <p>The MLP model performs well in CS mode with an
accuracy of .801 and F1 Macro of .767, but drops to .750 and
.686 in SV mode, highlighting its limitations in capturing
complex audio features with limited contextual information.
Compared to MLP, our method improves F1 Macro by +.07
in CS mode and +.10-.11 in SV mode, with corresponding
accuracy improvements of +.05-.06 and +.07-.08.</p>
        <p>For the 2D-CNN, fully fine-tuning all layers leads to poor
performance (.667 and .673 accuracy in CS and SV modes,
respectively; .400 and .402 F1 Macro), likely due to the
disruption of pre-trained features. Fine-tuning only the
classification head improves the performance to .789 and .782
accuracy in CS and SV modes, and .765 and .723 F1 Macro,
respectively. However, our method performs better than
the fine-tuned 2D-CNN, with +.07-.08 improvement in F1
Macro, +.07-.08 and +.04-.05 improvement in accuracy in
CS and SV modes.</p>
        <p>The above results show that fine-tuning the pre-trained
Wav2Vec2 model is an efective solution for small dataset
tasks, it highlights the necessity of carefully designed
optimization methods.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Fusion strategy vs. single modality performance The</title>
        <p>fusion method shows an advantage over the single modality
by efectively combining the complementary information
of CS and SV inputs. In particular, as shown in Table 2,
the proposed mid-level fusion pipeline shows significant
improvements over single modality models. Concatenated
Embeddings improves accuracy by +.02 and macro F1 by
+.001 on the CS model, and by +.05 and +.04 on the SV
model, respectively. Cross Attention performs even better,
with accuracy and F1 gains of +.03 and +.006 on the CS
model, and +.06 and +.05 on the SV model for accuracy and
macro F1, respectively. These results highlight the benefits
of leveraging complementary information from multiple
modalities.</p>
        <p>When compared to other fusion strategies, instead,
midlevel fusion consistently outperforms both early and late
fusion methods. The cross-attention method achieves the
best results with .885 accuracy and .843 macro F1, which
is +.02-.03 in accuracy and +.01-.02 in macro F1 compared
with early fusion. Similarly, it achieves +.01-.04
improvement in accuracy and +.01-.03 improvement in macro F1
compared to late fusion strategies such as Mixture of
Experts (MoE). These results demonstrate the efectiveness of
dynamically capturing inter-modality dependencies during
feature integration.</p>
        <p>Compared with early fusion that concatenates raw
features, the proposed mid-level fusion method can model
complex inter-dependencies, leading to robust feature
representation. In contrast, late fusion methods, while simpler to
implement, operate at the decision level and cannot fully
exploit the interactions between modalities.</p>
        <p>In summary, our proposed mid-level fusion strategy,
especially the cross-attention strategy, achieves the best
performance among all methods. The results show that it is
able to dynamically integrate complementary modality
information, leading to significant improvements in accuracy
and macro F1 performance.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>This study investigates the efectiveness of various models,
and fusion methods for speech impairment detection using
unimodal and multimodal approaches. We leverage
endto-end pre-trained models Wav2Vec2, which is once again
proven to be an efective model for solving audio tasks, even
with a limited dataset size. This not only reduces the steps of
manual feature extraction but also enables robust features
to be extracted from audio data through self-supervised
pre-training, showing good generalization ability.</p>
      <p>Among multimodal methods, our experiments show that
mid-level fusion strategies, especially the cross-attention
mechanism, outperform early and late fusion techniques.
The cross-attention mechanism dynamically captures
finegrained inter-modal dependencies, leading to the highest
performance. In contrast, early fusion methods, while
beneifcial for capturing joint features from the beginning, may
lack flexibility in handling complex interactions between
modalities. This often leads to inferior performance
compared to mid-level fusion. Late fusion methods are easier to
implement but have limited capabilities in modeling
complex feature interactions and a higher number of parameters.</p>
      <p>These findings provide valuable insights into the design of
voice disorder detection systems, especially with regard to
their potential applications in clinical diagnosis and health
monitoring.</p>
      <p>Future Work Although this study provides valuable
insights, there are still some limitations and directions for
improvement.</p>
      <p>First, the experiments are limited to a specific dataset,
IPV, which contains two homogeneous audio modalities and
cannot cover a wider range of scenarios. Future work can
explore larger and more diverse datasets, including datasets
collected in realistic noisy environments, or cross-lingual
datasets to evaluate the reliability of the model in the real
world. In addition, future work can integrate other medical
modalities (e.g. laryngoscope images + audio samples), to
expand audio beyond the audio domain for more
comprehensive voice disorder detection. Second, the current study
only focuses on voice disorder detection tasks. In the
future, it can be expanded to multi-classification tasks to more
comprehensively evaluate the efectiveness of the model
in practical applications, especially in the classification of
diferent types of pathologies that are at the root of the voice
disorder.</p>
      <p>Third, we only used the wav2vec2 model for feature
extraction and multi-modal fusion, and did not compare it on
other advanced Tansformer models (e.g. Hubert, WavLM,
etc). Future work can explore and evaluate their
efectiveness in the medical voice pathology analysis of these models.</p>
      <p>Data augmentation techniques can also be combined to
enhance the generalization ability of the model, so as to
maintain excellent performance in more diverse application
scenarios.</p>
      <p>By addressing these limitations, future research can build
on this study to develop more powerful, eficient, and
scalable voice disorder detection solutions, thereby bringing
greater social and technological impact for practical
applications.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          ,
          <article-title>The prevalence of voice problems among adults in the united states</article-title>
          ,
          <source>The Laryngoscope</source>
          <volume>124</volume>
          (
          <year>2014</year>
          )
          <fpage>2359</fpage>
          -
          <lpage>2362</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Merrill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. D.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <article-title>Voice disorders in the general population: prevalence, risk factors, and occupational impact</article-title>
          ,
          <source>The Laryngoscope</source>
          <volume>115</volume>
          (
          <year>2005</year>
          )
          <fpage>1988</fpage>
          -
          <lpage>1995</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Payten</surname>
          </string-name>
          , G. Chiapello,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Weir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Madill</surname>
          </string-name>
          ,
          <article-title>Frameworks, terminology and definitions used for the classification of voice disorders: a scoping review</article-title>
          ,
          <source>Journal of Voice</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Daraei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Villari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Rubin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Hillel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. R.</given-names>
            <surname>Hapner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Klein</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Johns</surname>
          </string-name>
          ,
          <article-title>The role of laryngoscopy in the diagnosis of spasmodic dysphonia</article-title>
          ,
          <source>JAMA Otolaryngology-Head &amp; Neck Surgery</source>
          <volume>140</volume>
          (
          <year>2014</year>
          )
          <fpage>228</fpage>
          -
          <lpage>232</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Ciravegna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Koudounas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fantini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cerquitelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Baralis</surname>
          </string-name>
          , E. Crosetti, G. Succo,
          <article-title>Non-invasive aipowered diagnostics: The case of voice-disorder detection-vision paper</article-title>
          ,
          <source>EDBT/ICDT Workshop 2348</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fantini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Ciravegna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Koudounas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cerquitelli</surname>
          </string-name>
          , E. Baralis,
          <string-name>
            <given-names>G.</given-names>
            <surname>Succo</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Crosetti,</surname>
          </string-name>
          <article-title>The rapidly evolving scenario of acoustic voice analysis in otolaryngology</article-title>
          ,
          <source>Cureus</source>
          <volume>16</volume>
          (
          <year>2024</year>
          )
          <article-title>e73491</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Rajpurkar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Banerjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Topol</surname>
          </string-name>
          ,
          <article-title>Ai in health and medicine</article-title>
          ,
          <source>Nature medicine 28</source>
          (
          <year>2022</year>
          )
          <fpage>31</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Collobert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Auli</surname>
          </string-name>
          , wav2vec:
          <article-title>Unsupervised pre-training for speech recognition</article-title>
          , arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>05862</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Koudounas</surname>
          </string-name>
          , G. Ciravegna,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fantini</surname>
          </string-name>
          , E. Crosetti, G. Succo,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cerquitelli</surname>
          </string-name>
          , E. Baralis,
          <article-title>Voice disorder analysis: a transformer-based approach</article-title>
          ,
          <source>in: Interspeech</source>
          <year>2024</year>
          ,
          <year>2024</year>
          , pp.
          <fpage>3040</fpage>
          -
          <lpage>3044</lpage>
          . doi:
          <volume>10</volume>
          .21437/ Interspeech.2024-
          <volume>1122</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M. La Quatra</surname>
            ,
            <given-names>M. F.</given-names>
          </string-name>
          <string-name>
            <surname>Turco</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Svendsen</surname>
            , G. Salvi,
            <given-names>J. R.</given-names>
          </string-name>
          <string-name>
            <surname>Orozco-Arroyave</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Siniscalchi</surname>
          </string-name>
          ,
          <article-title>Exploiting foundation models and speech enhancement for parkinson's disease detection from speech in real-world operative conditions</article-title>
          ,
          <source>in: Interspeech</source>
          <year>2024</year>
          ,
          <year>2024</year>
          , pp.
          <fpage>1405</fpage>
          -
          <lpage>1409</lpage>
          . doi:
          <volume>10</volume>
          .21437/Interspeech.2024-
          <volume>522</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>X.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Voice disorder classification using convolutional neural network based on deep transfer learning</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>7264</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L. W.</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. B.</given-names>
            <surname>Simões</surname>
          </string-name>
          , J. D. da
          <string-name>
            <surname>Silva</surname>
          </string-name>
          , D. da Silva Evangelista, A. C. d. N. e
          <string-name>
            <surname>Ugulino</surname>
            ,
            <given-names>P. O. C.</given-names>
          </string-name>
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>V. J. D.</given-names>
          </string-name>
          <string-name>
            <surname>Vieira</surname>
          </string-name>
          ,
          <article-title>Accuracy of acoustic analysis measurements in the evaluation of patients with diferent laryngeal diagnoses</article-title>
          ,
          <source>Journal of voice 31</source>
          (
          <year>2017</year>
          )
          <fpage>382</fpage>
          -
          <lpage>e15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Alhussein</surname>
          </string-name>
          , G. Muhammad,
          <article-title>Automatic voice pathology monitoring using parallel deep models for smart healthcare</article-title>
          ,
          <source>Ieee Access</source>
          <volume>7</volume>
          (
          <year>2019</year>
          )
          <fpage>46474</fpage>
          -
          <lpage>46479</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Leung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. T.</given-names>
            <surname>Chui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>O. de Pablos</surname>
          </string-name>
          ,
          <article-title>A support vector machine-based voice disorders detection using human voice signal</article-title>
          ,
          <source>in: Artificial Intelligence and Big Data Analytics for Smart Healthcare, Elsevier</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>197</fpage>
          -
          <lpage>208</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>X.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Voice disorder classification using convolutional neural network based on deep transfer learning</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>7264</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>U. K. Lilhore</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Dalal</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Faujdar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Margala</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Chakrabarti</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Chakrabarti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Simaiya</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Thangaraju</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Velmurugan</surname>
          </string-name>
          ,
          <article-title>Hybrid cnn-lstm model with eficient hyperparameter tuning for prediction of parkinson's disease</article-title>
          ,
          <source>Scientific Reports</source>
          <volume>13</volume>
          (
          <year>2023</year>
          )
          <fpage>14605</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Almasoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A. E.</given-names>
            <surname>Eisa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. N.</given-names>
            <surname>Al-Wesabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elsafi</surname>
          </string-name>
          , M. Al Duhayyim,
          <string-name>
            <surname>I. Yaseen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hamza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Motwakel</surname>
          </string-name>
          ,
          <article-title>Parkinson's detection using rnn-graph-lstm with optimization based on speech signals</article-title>
          ,
          <source>Comput. Mater. Contin</source>
          <volume>72</volume>
          (
          <year>2022</year>
          )
          <fpage>872</fpage>
          -
          <lpage>886</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>R.</given-names>
            <surname>Islam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Abdel-Raheem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tarique</surname>
          </string-name>
          ,
          <article-title>Voice pathology detection using convolutional neural networks with electroglottographic (egg) and speech signals</article-title>
          ,
          <source>Computer Methods and Programs in Biomedicine Update</source>
          <volume>2</volume>
          (
          <year>2022</year>
          )
          <fpage>100074</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <article-title>A voice disease detection method based on mfccs and shallow cnn</article-title>
          ,
          <source>Journal of Voice</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kanda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Yoshioka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xiao</surname>
          </string-name>
          , et al.,
          <article-title>Wavlm: Large-scale self-supervised pre-training for full stack speech processing</article-title>
          ,
          <source>IEEE Journal of Selected Topics in Signal Processing</source>
          <volume>16</volume>
          (
          <year>2022</year>
          )
          <fpage>1505</fpage>
          -
          <lpage>1518</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>W.-N.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bolte</surname>
          </string-name>
          , Y.
          <string-name>
            <surname>-H. H. Tsai</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Lakhotia</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mohamed</surname>
          </string-name>
          , Hubert:
          <article-title>Selfsupervised speech representation learning by masked prediction of hidden units</article-title>
          ,
          <source>IEEE/ACM transactions on audio, speech, and language processing 29</source>
          (
          <year>2021</year>
          )
          <fpage>3451</fpage>
          -
          <lpage>3460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Koudounas</surname>
          </string-name>
          , E. Pastor,
          <string-name>
            <given-names>G.</given-names>
            <surname>Attanasio</surname>
          </string-name>
          , L. de Alfaro, E. Baralis,
          <article-title>Prioritizing data acquisition for end-to-end speech model improvement</article-title>
          ,
          <source>in: ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>7000</fpage>
          -
          <lpage>7004</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP48485.
          <year>2024</year>
          .
          <volume>10446326</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A.</given-names>
            <surname>Koudounas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. La</given-names>
            <surname>Quatra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Siniscalchi</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Baralis,</surname>
          </string-name>
          <article-title>voc2vec: A foundation model for non-verbal vocalization</article-title>
          ,
          <source>in: ICASSP 2025 - 2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>M. La Quatra</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Koudounas</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Vaiani</surname>
            , E. Baralis,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Garza</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Cagliero</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          <string-name>
            <surname>Siniscalchi</surname>
          </string-name>
          ,
          <article-title>Benchmarking representations for speech, music, and acoustic events</article-title>
          ,
          <source>in: 2024 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW)</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>S. wen Yang</given-names>
            , P.
            <surname>-H. Chi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-S.</given-names>
            <surname>Chuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-I. J.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lakhotia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chang</surname>
          </string-name>
          , G.-T. Lin,
          <string-name>
            <given-names>T.-H.</given-names>
            <surname>Huang</surname>
          </string-name>
          , W.-C. Tseng,
          <string-name>
            <given-names>K.</given-names>
            tik
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.-R.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Watanabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          , H. yi Lee,
          <article-title>Superb: Speech processing universal performance benchmark</article-title>
          ,
          <source>in: Interspeech</source>
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>1194</fpage>
          -
          <lpage>1198</lpage>
          . doi:
          <volume>10</volume>
          .21437/Interspeech.2021-
          <volume>1775</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Auli</surname>
          </string-name>
          , wav2vec
          <volume>2</volume>
          .
          <article-title>0: A framework for self-supervised learning of speech representations</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>12449</fpage>
          -
          <lpage>12460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Stahlschmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ulfenborg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Synnergren</surname>
          </string-name>
          ,
          <article-title>Multimodal deep learning for biomedical data fusion: a review, Briefings in Bioinformatics 23 (</article-title>
          <year>2022</year>
          )
          <article-title>bbab569</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>L.</given-names>
            <surname>Ilias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Askounis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Psarras</surname>
          </string-name>
          ,
          <article-title>Detecting dementia from speech and transcripts using transformers</article-title>
          ,
          <source>Computer Speech &amp; Language</source>
          <volume>79</volume>
          (
          <year>2023</year>
          )
          <fpage>101485</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Audhkhasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <article-title>A mixture of experts approach towards intelligibility classification of pathological speech</article-title>
          ,
          <source>in: 2015 IEEE international conference on acoustics, speech and signal processing (ICASSP)</source>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>1986</fpage>
          -
          <lpage>1990</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>G. B.</given-names>
            <surname>Kempster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Gerratt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. V.</given-names>
            <surname>Abbott</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. BarkmeierKraemer</surname>
          </string-name>
          , R. E. Hillman,
          <article-title>Consensus auditoryperceptual evaluation of voice: development of a standardized clinical protocol (</article-title>
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sandler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Howard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zhmoginov</surname>
          </string-name>
          , L. Chen,
          <article-title>Inverted residuals and linear bottlenecks: Mobile networks for classification, detection and segmentation</article-title>
          , CoRR abs/
          <year>1801</year>
          .04381 (
          <year>2018</year>
          ). URL: http: //arxiv.org/abs/
          <year>1801</year>
          .04381. arXiv:
          <year>1801</year>
          .04381.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <article-title>Imagenet: A large-scale hierarchical image database, in: 2009 IEEE conference on computer vision and pattern recognition</article-title>
          , Ieee,
          <year>2009</year>
          , pp.
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>