<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linguistic and Gender Variation in Speech Emotion Recognition Using Spectral Features?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zachary Dair</string-name>
          <email>zachary.dair@mycit.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryan Donovan</string-name>
          <email>brendan.donovan@mycit.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ruairi O'Reilly[</string-name>
          <email>ruairi.orielly@mtu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Munster Technological University</institution>
          ,
          <addr-line>Cork</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This work explores the e ect of gender and linguistic-based vocal variations on the accuracy of emotive expression classi cation. Emotive expressions are considered from the perspective of spectral features in speech (Mel-frequency Cepstral Coe cient, Melspectrogram, Spectral Contrast). Emotions are considered from the perspective of Basic Emotion Theory. A convolutional neural network is utilised to classify emotive expressions in emotive audio datasets in English, German, and Italian. Vocal variations for spectral features assessed by (i) a comparative analysis identifying suitable spectral features, (ii) the classi cation performance for mono, multi and cross-lingual emotive data and (iii) an empirical evaluation of a machine learning model to assess the e ects of gender and linguistic variation on classi cation accuracy. The results showed that spectral features provide a potential avenue for increasing emotive expression classi cation. Additionally, the accuracy of emotive expression classi cation was high within mono and cross-lingual emotive data, but poor in multi-lingual data. Similarly, there were di erences in classi cation accuracy between gender populations. These results demonstrate the importance of accounting for population di erences to enable accurate speech emotion recognition.</p>
      </abstract>
      <kwd-group>
        <kwd>A ective Computing</kwd>
        <kwd>Speech Emotion Recognition</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Prosody Analysis</kwd>
        <kwd>Convolutional Neural Networks</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Speech emotion recognition (SER) is the classi cation of the emotional states of
a speaker from speech. These emotional states can either be discrete experiences,
such as the basic emotions (Anger, Disgust, Fear, Joy, Sadness, and Surprise),
or attributes of emotional states (Arousal, Valence, Dominance). The ability to
accurately classify these emotional states through SER enables the provision of
tailored services that can adapt to the psychological needs of user groups. SER
attempts to classify these emotional states via verbal and non-verbal components
of speech. Verbal components of speech are the speci c words a speaker says.</p>
      <p>
        Verbal SER involves classifying emotional expressions from the word choice or
associations, syntax, use of colloquialisms or sarcasm concerning the intent of the
speaker [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Non-verbal components of speech refer to how the speaker expresses
their speech. Non-verbal SER involves classifying emotional expressions from the
? This publication has emanated from research supported in part by a Grant from
      </p>
      <p>
        Science Foundation Ireland under Grant number 18/CRT/6222
Copyright 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0)
speaker's acoustic features (Pitch, Tone, Timbre, Number of pauses, Loudness,
Speech Rate) while speaking [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>Research has shown that non-verbal components provide valuable
information, distinct from verbal content, in accurately classifying emotional expressions.</p>
      <p>
        For example, non-verbal features are capable of indicating emotional signals
when the verbal content of the speech is neutral [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Many non-verbal
features have been shown to exist across multiple languages, enabling generalisable
analysis [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Additionally, non-verbal features have been shown to have a
direct impact on emotion recognition processes in the brain, indicating they are
vital to achieving accurate emotion classi cation [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. Overall, the larger amount
of features available for non-verbal SER, in comparison to verbal SER, enables
more accurate emotion classi cation.
      </p>
      <p>
        Traditional non-verbal SER1 approaches recruit human observers, who are
asked to classify emotions from non-verbal content collected from one or
several speakers. While traditional approaches are useful for relatively small-scale
datasets, as they demonstrate high levels of accuracy [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], they are
impractical for dealing with large and continuously growing datasets that
characterise modern human-computer interactions. In an attempt to scale the bene ts
of SER, research and industrial applications have been developed to conduct
SER automatically. These applications range from open-source research tools (for
example, Audeering ) to commercial tools (for example, Good Vibrations,
Vokaturi, and deep a ect API ) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], all of which employ machine learning-based
approaches. If SER can be accurately automated, it would facilitate increased
e cacy in human-computer interactions, user modelling, and personalisation
services.
      </p>
      <p>
        Despite the potential bene ts of automated SER, it has not been widely
implemented in everyday life settings [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This is despite reviews of the SER
literature demonstrating that automated SER approaches perform similarly or
even better than traditional approaches [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. One reason for the limited uptake of
automated SER approaches is that they are not as generalisable and adaptable as
human classi ers, who can use contextual information (for example, the gender
of the speaker and the language of the speech) when classifying speech. These
two factors, gender and language, have been shown to signi cantly a ect the
accuracy of automated SER approaches [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>
        SER models that account for gender di erences are more accurate than those
that do not [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. It is di cult, however, to accurately model for these di erences.
      </p>
      <p>
        While there exists many features of speech that are signi cantly in uenced by
gender, there exists considerable overlap between males and females on these
same features. This overlap makes it di cult to distinguish between the gender
and linguistic features of emotions that reliably indicate emotions and those that
do not. Concerning language, SER models trained on one language signi cantly
decrease in performance when tested on another language. This contrasts to
human observers in traditional SER, who are capable of detecting emotional
expressions in speech in languages they do not speak [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
1 Hereonafter, non-verbal SER is referred to as SER as shorthand.
      </p>
      <p>
        One potential mechanism for improving the accuracy of SER across di
erent gender and linguistic populations are spectral features. Spectral features of
audio are computed by converting a time-based signal into the frequency
domain. These features represent the distribution of energy across a frequency and
the harmonic components of sound. Components such as pitch changes in the
audio signal, can assist in gender di erentiation, and voice quality, can be
identi ed through voice level as either tense, harsh, breathy, or lax [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Aggregating
spectral features' continuous, qualitative, and spectral data aids di erentiation
between genders and linguistic populations thereby enabling more generalisable
and accurate emotive expression classi cations.
      </p>
      <p>This study assesses the e ect of spectral features on the performance of
automatic SER across di erent gender and linguistic populations by (i) conducting
a comparative analysis of spectral features to identify a suitable feature set,
(ii) comparing classi cation performance of SER approaches when using
monolingual, multi-lingual, and cross-lingual emotive data, and (iii) empirically
evaluating the performance of a machine learning (ML) model to assess the e ect of
gender and linguistic variations on SER classi cation accuracy.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Previous research has examined the di erences between gender populations in
acoustic features and their e ect on SER performance. Several acoustic
features have been described as \gender-dependent". Gender dependent in this
case means that gender signi cantly in uences the expression of that feature
[
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. Pitch is an example of a gender-dependent feature. Females score higher,
on average, than males on pitch [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. This information, however, is not su cient
to enable generalisable SER classi cation, because there is considerable overlap
between males and females on this feature [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. This is also the case for several
other acoustic features described as \gender-dependent".
      </p>
      <p>
        Similarly, previous research has examined the impact of language in
developing accurate SER models. Human observers can accurately classify emotions in
their native language, and non-native languages [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This result has been
replicated across various cultures, including indigenous tribes, demonstrating the
generalisability of traditional SER [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Similarly, automated SER models
achieve a high level of accuracy in mono-lingual settings, where the models are
trained and evaluated in the same language. The performance of these models,
however, signi cantly drop when they are tested across languages [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        If ML models can di erentiate between acoustic features that indicate
emotional expressions for each gender and linguistic population reliably, then one
should expect generalisability of SER performance between both populations
[
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. To achieve this, spectral features that reliably indicate emotional
experiences are rst required.
      </p>
      <p>
        A survey [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] analysed features across 17 distinctive datasets. These datasets
comprise 9 languages, both male and female speakers, professional and
nonprofessional acting sessions, recorded call centre conversations and speech recorded
under simulated and naturally occurring stress situations.
      </p>
      <p>
        A pool of acoustic features and SER were analysed in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Which can
be divided into several categories (Continuous, Qualitative, Spectral and
Nonlinear Teager Energy Operated (TEO). The results of both reviews indicated
that Continuous and Spectral features were related to emotional expressions
in speech. A weaker relationship was found between emotional expressions and
Qualitative and TEO features.
      </p>
      <p>
        To automatically classify relevant spectral features, a ML methodology is
proposed in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The methodology uses a convolutional neural network (CNN)
and SER analysis. The approach consists of using spectral features such as
Mel-frequency Cepstral Coe cients (MFCCs), Mel-scaled spectrogram,
Chromagram, Spectral contrast and Tonnetz representation extracted from emotive
audio belonging to three distinct datasets. This work, however, focused on a
combined gender classi cation of emotive expression, inhibits the analysis of vocal
variations between gender. This study extends this work to by analysing the
effect of gender and linguistic populations on the emotive expression classi cation
of spectral features.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3 Methodology</title>
      <p>To assess the e ect of spectral features for emotive expression classi cation across
di erent gender and linguistic populations a series of experiments were
conducted. Which required the identi cation of a suitable set of gender and/or
language-dependent spectral features. The classi cation was enabled through
the usage of the CNN described in 3.5, the performance of this CNN will be
compared across experiments that account for linguistic and gender di erences
and experiments that do not account for these di erences. The gender
populations are male and female and the linguistic populations are English, German,
and Italian.</p>
      <p>
        Discrete emotional models identify several distinct emotions to be classi ed.
This study uses Basic Emotion Theory (BET), as the discrete emotional model
for SER. In BET the distinct emotions are Anger, Disgust, Fear, Joy, Sadness,
and Surprise. These six emotions are considered basic as (i) they consistently
correlate with psychological, behavioural, and neurophysiological activity [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ],
which makes their objective measurement possible, and (ii) they interact to form
more cognitively and culturally mediated emotions, such as shame or guilt [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
The BET contrasts from dimensional models of emotions that focus on attributes
of emotions (for example, arousal, valence). BET models have been extensively
used in SER research as they capture a wider range of emotions and are intuitive
to label in comparison to dimensional models [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
3.1
      </p>
      <sec id="sec-3-1">
        <title>Emotive Speech Data</title>
        <p>Three distinct emotive audio datasets were used (see Table 1) to enable a
comparative analysis of variations between the data. These variations can originate
from many factors such as the gathering method, emotions exhibited, language,
speaker gender, or sampling rate.</p>
        <p>RAVDESS contains emotional speech data constituting statements and
songs in English. For the proposed work, only speech samples are considered.</p>
        <p>Datasets</p>
        <p>
          RAVDESS[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]
EMO-DB[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
Participants uttered two neutral statements across multiple trials. For each trial,
participants were asked to utter the statement in a manner that conveyed one
of the six basic emotions. The statements were controlled to ensure equal levels
of syllables, word frequency, and familiarity to the speaker.
        </p>
        <p>EMO-DB contains emotional speech data constituting statements in
German. Participants uttered ten statements, comprised of everyday language and
syntactic structure, in various lengths to simulate natural speech. Each
utterance was evaluated with regards to recognisability and naturalness of emotions
exhibited. The emotive reaction of surprise is not considered.</p>
        <p>EMOVO contains emotional speech data constituting statements in Italian.
The participants uttered fourteen distinct semantically neutral statements. Each
conveys a basic emotion and is spoken naturally. An important consideration in
the creation of this dataset was the presence of all phonemes of the Italian
language and a balanced presence of voiced and unvoiced consonants.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Feature Selection</title>
        <p>
          In order to identify spectral features indicative of emotive expression, feature
selection is performed. Which identi es information such as pitch, energy, voice
levels and energy distribution for classi cation. Spectral features were extracted,
using the audio analysis library Librosa [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>These features are as follows: Mel-Frequency Cepstral Coe cients (MFCC)
which represents the shape of a signal spectrum, and is achieved by collectively
representing a Mel-frequency cepstrum per frame. Chroma Energy Normalized
(CENS) which represents statistics indicative of normalised values quantifying
tempo, articulation, and pitch deviations. Zero-Crossing Rate (ZCR) which
represents the rate of change in the audio signal from positive to negative or negative
to positive, through zero. Chromagram which represents a transformed signal
spectrum built using the chromatic scale capturing the harmonic and melodic
characteristics. Melspectrogram which represents a spectrogram where the
frequency is converted from a linear scale to a Mel-scale, resulting in a
spectrogram re ective of how humans perceive frequencies, capturing the amplitude of
the signal. Spectral Contrast which represents the energy contrast computed by
comparing the peak energy and valley energy in each band converted from
spectrogram frames. Tonnetz which represents the tonal features in a 6-dimensional
format, capturing traditional harmonic relationships. (perfect fth, minor third
and major third).</p>
        <p>In order to identify a set of spectral features that captures su cient emotive
data, a classi er is trained on these features individually before being trained on
permutations of the features across the datasets to identify the best performing
feature set.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3 Linguistic Variation</title>
        <p>To evaluate the accuracy of a SER model in relation to linguistic variation three
experiments are conducted across three languages (English, German, Italian) as
follows:
1. Mono-Lingual - The CNN is used to classify spectral features indicative of
emotive expression from each language independently. The results will be
used as a baseline performance of the model's classi cation performance in
a single language.
2. Multi-Lingual - This experiment consists of three permutations. For each,
the CNN is rst trained on one of the chosen languages and then evaluated
for SER accuracy against the remaining languages. The intent is to
identify spectral features capacity to classify emotive expression independent of
language.
3. Cross-Lingual - The CNN is trained on an aggregation of all training data.</p>
        <p>This will provide insights into the performance of a CNN model when its
training corpus contains multiple languages. The intent is to identify the
generalisability of the chosen spectral features for emotive expression
classication.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4 Cross-Gender Emotion</title>
        <p>To assess the e ect of gender populations on SER, an experiment was
conducted using the spectral features identi ed in Section. 3.2, comprising of two
steps. In step one, the data is classi ed using the six basic emotions. This
initial classi cation provided a baseline accuracy where gender is not speci ed. In
step two, the same dataset is classi ed with gender-emotion labels (for
example, Male-Anger/Female-Anger; Male-Joy/Female-Joy). The intent is to evaluate
if the CNN can identify gender-dependent acoustic elements from the spectral
features.</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5 Classi cation Of Emotive Speech</title>
        <p>
          The architecture of the CNN used for classi cation is depicted in Figure 1. A
CNN was re-implemented, due to the high emotive expression classi cation
performance using spectral features as exhibited in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Optimal hyper-parameters
were identi ed from a comparative analysis against related approaches. The
model is trained on the extracted spectral features over 150 epochs using a
batch size of 16 and 5-Fold validation is performed. During each epoch a portion
of data is used to evaluate the model. The model performance is evaluated across
the metrics precision, recall, F1-score for each of the six basic emotions.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 Results</title>
      <p>Feature extraction - The comparative analysis of the spectral features in
isolation provides insights into the performance on a per feature basis. MFCC,
Melspectrogram and Spectral Contrast were the highest performing individual
spectral features across the datasets. When combined these features formed a
vector of 155 data-points, and the highest performing permutation of spectral
features as denoted in Table. 2</p>
      <p>Linguistic variation - The results for emotive analysis in mono, multi and
cross-lingual data are denoted in Table. 3, these highlight the considerations
for SER across multiple languages. Mono-lingual and Cross-lingual approaches
achieve high accuracies. A degradation in performance is experienced from
multilingual approaches.</p>
      <p>Cross-Gender Emotion - Table. 4 identi es discrepancies in the emotive
classi cation performance, across both gender and language. This indicates a
degree of vocal variance stemming from the population di erences, highlighting
considerations for gender-speci c SER approaches.</p>
    </sec>
    <sec id="sec-5">
      <title>5 Discussion</title>
      <p>Feature extraction - There are several notable results from the feature
extraction stage that are worth discussion. Firstly, the results of the comparative
analysis showed that the MFCC feature enabled the highest emotion
recognition accuracy within each dataset. Demonstrating the importance of modelling
for phonetic properties, found within the speech signal shape, in enabling
accurate emotion classi cation across languages. Secondly, the performance of the
Melspectrogram feature varied across the datasets. Melspectrogram performed
highly on EMO-DB and EMOVO and poorly on RAVDESS. This
inconsistency was caused by the signi cantly lower amplitude within the audio data
of RAVDESS compared to EMO-DB and EMOVO denoted in Table. 2.
Participants in the RAVDESS dataset were given instructions to exhibit emotions</p>
      <p>Spectral Features
MFCC
Melspectogram
Spectral Contrast
CENS
STFT
ZCR
Tonnetz
Spectral Feature Permutations
MFCC, Melspectogram, Spectral Contrast
Spectral Contrast, Melspectogram, MFCC
MFCC, STFT, Spectral Contrast
STFT, Melspectogram, Spectral Contrast, Tonnetz
Additional Characteristics
Average Absoulte Amplitude (dB)
0.69
0.44
0.39
0.31
0.22
0.23
0.22
RAVDESS EMO-DB EMOVO Mean
NA
with varying intensity normal and strong, additionally, post-processing
procedures are likely the cause for the lower level of amplitude. This has signi cant
consequence for classi cation when amplitude is utilised as a measure of emotive
expression. Thirdly, Spectral Contrast in isolation provides greater average
accuracy than the remaining individual features. Therefore, the contrast between
peak and valley energy is partially indicative of emotion in emotive speech. The
high performance of these features in isolation indicates their potential suitability
for emotive expression classi cation. Finally, the results showed that combining
these three features increased accuracies for each data set. This demonstrates
the importance of including such spectral features in SER classi cations.</p>
      <p>
        Linguistic variation - SER accuracy decreased signi cantly between the
di erent types of analysis. The high performance on mono-lingual analysis
indicates the importance of the selected spectral features for each language. The poor
performance in multi-lingual analysis, however, indicates a lack of universality
between spectral representations of emotions across the datasets. For example,
there was signi cant variation in the amplitude and pitch range across the
languages, as identi ed in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Additionally, these di erences likely stemmed from
di erences in data collection (equipment, volume, actor-microphone distance)
across the datasets, thereby decreasing the performance of the CNN in multi
and cross-lingual analysis. This highlights the need for standardised recording
lassi cation accuracy per language derived from 5-fold cross-validation, and train/test
data speci ed. (X represents data from the corresponding column name)
Method
Mono-Ling.
      </p>
      <p>Train/Test</p>
      <p>X/X
Multi-Ling.-Eng. RAVDESS/X
Multi-Ling.-Ger. EMO-DB/X
Multi-Ling.-Ital. EMOVO/X
Cross-Ling.</p>
      <p>All/All</p>
      <p>English</p>
      <p>German</p>
      <p>Italian
NA
NA</p>
      <p>NA
NA
NA
NA</p>
      <p>NA
procedures within SER research, and/or additional procedures to normalize
linguistic variances.</p>
      <p>
        Cross-Gender Emotion - There were substantial discrepancies in spectral
representations between the gender populations. As a result, SER performance
between the genders di ered signi cantly across the six basic emotions. These
di erences were varied across languages, indicating that there is a generalisable
e ect of gender on SER performance when using spectral features, however this
is impacted by linguistic variation. For example, high energy emotions (such
as Anger and Joy) were more clearly detected in German-speaking males than
in other linguistic or gender groups. The German language is characterised by
higher amplitude, energy and harsher speech when articulated within males [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
This likely contributed to higher accuracy of detecting those emotions from
German-speaking males. Disgust and Fear, in contrast, were more accurately
classi ed from the speech of females across each language. This may have resulted
from di erences in vocals, particularly pitch, between the gender groups [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
Additionally, female voices tend to articulate in a softer manner [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] conducive
to representing softer, lower amplitude emotions.
      </p>
      <p>These results likely contributed to the weaker performance of combined
gender classi cation in comparison to gender-speci c classi cations of emotion.
Combined approaches can account for variations between genders and languages,
however, in certain cases, a speci c gender may act as a limiting factor reducing
the overall accuracy.</p>
      <p>Limitations - The major limitations of this work concern sampling issues.
Firstly, the sample size across the three datasets is small (N = 40). Since sample
sizes can exaggerate the variances between populations, then the small sample
size in this study may have exaggerated real di erences in spectral features
between gender and linguistic populations. Secondly, the datasets were comparing
linguistic populations across an unequal number of speakers. Di erent sample
sizes will have higher or lower ranges of variability. Comparing di erent sample
sizes makes it di cult to determine whether the results stem from real
betweengroup di erences or are the result of a higher level of \noisy" data found in
small sample sizes. These limitations damage the generalisability of the work.
In a follow-up study, the sample size of the overall dataset will be increased and
the number of speakers per linguistic group will be controlled.</p>
    </sec>
    <sec id="sec-6">
      <title>6 Conclusion</title>
      <p>This work was an exploratory analysis of linguistic and gender variation on the
emotive classi cation of spectral features. The results showed that features
using the Mel-Scale and representations of amplitude and energy are important in
accurate SER across di erent gender and linguistic populations. Additionally,
the results showed that higher energy emotions such as Anger, Joy, Surprise
were easier to identify originating from male voices in a high amplitude, harsh
language such as German. Similarly lower energy emotions such as Disgust and
Fear were easier to identify from female voices in each language. These
observations highlight the importance of signal amplitude and energy when analysing
emotion across gender and language.</p>
      <p>The performance of emotive expression classi cation across mono, multi and
cross-lingual data provides insights into linguistic variation of emotive audio.
Mono-lingual approaches are suitable as baselines in comparison to multi and
cross-lingual approaches. A linguistic variance between emotive expression
represented by spectral features can be identi ed. To overcome these variances,
it is recommended that cross-linguistic approaches, that combine languages for
training, are implemented.</p>
      <p>Future work should explore whether classi cation accuracy is a ected by
(i) model optimization in terms of structure and/or (ii) comprehensive
hyperparameter tuning (epochs, batch size, optimizers) using a grid search technique.
Additionally, the relationships between linguistics, gender and emotion should
be explored from a vocal, psychological and technical classi cation perspective,
with a larger sample size to garner further insights into improving SER
generalisability.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Akcay</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oguz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classi ers</article-title>
          .
          <source>Speech Communication</source>
          <volume>116</volume>
          ,
          <issue>56</issue>
          {
          <fpage>76</fpage>
          (
          <year>2020</year>
          ). https://doi.org/10.1016/j.specom.
          <year>2019</year>
          .
          <volume>12</volume>
          .001
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ancilin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Milton</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Improved speech emotion recognition with Mel frequency magnitude coe cient</article-title>
          .
          <source>Applied Acoustics</source>
          <volume>179</volume>
          ,
          <issue>108046</issue>
          (
          <year>2021</year>
          ). https://doi.org//10.1016/j.apacoust.
          <year>2021</year>
          .108046
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Burkhardt</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paeschke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rolfes</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sendlmeier</surname>
            ,
            <given-names>W.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A database of German emotional speech</article-title>
          .
          <source>In: 9th ECSCT</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cordaro</surname>
            ,
            <given-names>D.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keltner</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tshering</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wangchuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flynn</surname>
            ,
            <given-names>L.M.:</given-names>
          </string-name>
          <article-title>The voice conveys emotion in ten globalized cultures and one remote village in Bhutan</article-title>
          .
          <source>Emotion</source>
          <volume>16</volume>
          (
          <issue>1</issue>
          ),
          <volume>117</volume>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Costantini</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iadarola</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paoloni</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Todisco</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>EMOVO corpus: an Italian emotional speech database</article-title>
          . LREC p.
          <volume>4</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dair</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Donovan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Reilly</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          :
          <article-title>Classi cation of emotive expression using verbal and non verbal components of speech</article-title>
          .
          <source>In: 2021 ISSC</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          (
          <year>2021</year>
          ). https://doi.org/10.1109/ISSC52156.
          <year>2021</year>
          .9467869
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ekman</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Emotions revealed</article-title>
          .
          <source>Bmj</source>
          <volume>328</volume>
          (
          <issue>Suppl S5</issue>
          )
          <article-title>(</article-title>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>El</given-names>
            <surname>Ayadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Kamel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.S.</given-names>
            ,
            <surname>Karray</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          :
          <article-title>Survey on speech emotion recognition: Features, classi cation schemes, and databases</article-title>
          .
          <source>Pattern Recongnition</source>
          <volume>44</volume>
          (
          <issue>3</issue>
          ),
          <volume>572</volume>
          {
          <fpage>587</fpage>
          (
          <year>2011</year>
          ). https://doi.org/10.1016/j.patcog.
          <year>2010</year>
          .
          <volume>09</volume>
          .020
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Eyben</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huber</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marchi</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuller</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuller</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Real-time robust recognition of speakers' emotions and characteristics on mobile platforms</article-title>
          .
          <source>In: 2015 ACII</source>
          . pp.
          <volume>778</volume>
          {
          <fpage>780</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Feraru</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuller</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , et al.:
          <article-title>Cross-language acoustic emotion recognition: An overview and some tendencies</article-title>
          .
          <source>In: 2015 International Conference on A ective Computing and Intelligent Interaction (ACII)</source>
          . pp.
          <volume>125</volume>
          {
          <fpage>131</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Garcia-Garcia</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Penichet</surname>
            ,
            <given-names>V.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lozano</surname>
          </string-name>
          , M.D.:
          <article-title>Emotion detection: a technology review</article-title>
          .
          <source>In: Proceedings of the XVIII international conference on human computer interaction</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Gobl</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chasaide</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The role of voice quality in communicating emotion, mood and attitude</article-title>
          .
          <source>Speech Communication</source>
          <volume>40</volume>
          ,
          <issue>189</issue>
          {
          <volume>212</volume>
          (04
          <year>2003</year>
          ). https://doi.org/10.1016/S0167-
          <volume>6393</volume>
          (
          <issue>02</issue>
          )
          <fpage>00082</fpage>
          -
          <lpage>1</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Goddard</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Semantic Analysis: A Practical Introduction</article-title>
          . Oxford University Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hsu</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>M.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.H.</given-names>
          </string-name>
          :
          <article-title>Speech emotion recognition considering nonverbal vocalization in a ective conversations</article-title>
          .
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          <volume>29</volume>
          ,
          <volume>1675</volume>
          {
          <fpage>1686</fpage>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Issa</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fatih</surname>
            <given-names>Demirci</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Yazici</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Speech emotion recognition with deep convolutional neural networks</article-title>
          .
          <source>BSPC 59</source>
          ,
          <issue>101894</issue>
          (
          <year>2020</year>
          ). https://doi.org/10.1016/j.bspc.
          <year>2020</year>
          .101894
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Latinus</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , M.J.:
          <article-title>Discriminating male and female voices: di erentiating pitch and gender</article-title>
          .
          <source>Brain topography 25(2)</source>
          ,
          <volume>194</volume>
          {
          <fpage>204</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lausen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schacht</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Gender di erences in the recognition of vocal emotions</article-title>
          .
          <source>Frontiers in psychology 9</source>
          ,
          <issue>882</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Livingstone</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Russo</surname>
            ,
            <given-names>F.A.</given-names>
          </string-name>
          :
          <article-title>The Ryerson audio-visual database of emotional speech and song (RAVDESS)</article-title>
          .
          <source>PloS one 13(5)</source>
          ,
          <year>e0196391</year>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Brian</surname>
            <given-names>McFee</given-names>
          </string-name>
          ,
          <article-title>Colin Ra el</article-title>
          , Dawen Liang, Daniel P.W. Ellis,
          <string-name>
            <surname>Matt</surname>
            <given-names>McVicar</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Eric</given-names>
            <surname>Battenberg</surname>
          </string-name>
          , Oriol Nieto:
          <article-title>Librosa: Audio and Music Signal Analysis in Python</article-title>
          .
          <source>In: Proceedings of the 14th Python in Science Conference</source>
          . pp.
          <volume>18</volume>
          {
          <issue>24</issue>
          (
          <year>2015</year>
          ). https://doi.org/10.25080/Majora-7b98e3ed-003
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Mennen</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , Schae er, F.,
          <string-name>
            <surname>Docherty</surname>
          </string-name>
          , G.:
          <article-title>Pitching it di erently : a comparison of the pitch ranges of German and English speakers</article-title>
          .
          <source>16th ICPS</source>
          pp.
          <volume>1769</volume>
          {
          <issue>1772</issue>
          (08
          <year>2007</year>
          ). https://doi.org/20.500.12289/42
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Neumann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , et al.:
          <article-title>Cross-lingual and multilingual speech emotion recognition on English and French</article-title>
          . In: 2018 ICASSP. pp.
          <volume>5769</volume>
          {
          <fpage>5773</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Pepiot</surname>
          </string-name>
          , E.:
          <article-title>Voice, speech and gender:: Male-female acoustic di erences and cross-language variation in English and French speakers</article-title>
          .
          <source>corela</source>
          (
          <year>2015</year>
          ). https://doi.org/10.4000/corela.3783
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Rajoo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aun</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          :
          <article-title>In uences of languages in speech emotion recognition: A comparative study using malay, english and mandarin languages</article-title>
          .
          <source>In: 2016 ISCAIE</source>
          . pp.
          <volume>35</volume>
          {
          <fpage>39</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Sandier</surname>
            ,
            <given-names>B.R.</given-names>
          </string-name>
          :
          <article-title>Women faculty at work in the classroom, or, why it still hurts to be a woman in labor</article-title>
          .
          <source>Communication Education</source>
          <volume>40</volume>
          (
          <issue>1</issue>
          ),
          <volume>6</volume>
          {
          <fpage>15</fpage>
          (
          <year>1991</year>
          ). https://doi.org/10.1080/03634529109378821
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Sauter</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calder</surname>
            ,
            <given-names>A.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Perceptual cues in nonverbal vocal expressions of emotion</article-title>
          .
          <source>Quarterly Journal of Experimental Psychology</source>
          <volume>63</volume>
          (
          <issue>11</issue>
          ),
          <volume>2251</volume>
          {
          <fpage>2272</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Sauter</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ekman</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Cross-cultural recognition of basic emotions through nonverbal emotional vocalizations</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>107</volume>
          (
          <issue>6</issue>
          ),
          <volume>2408</volume>
          {
          <fpage>2412</fpage>
          (
          <year>2010</year>
          ), publisher: National Acad Sciences
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Sauter</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eisner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ekman</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          :
          <article-title>Emotional vocalizations are recognized across cultures regardless of the valence of distractors</article-title>
          .
          <source>Psychological science 26(3)</source>
          ,
          <volume>354</volume>
          {
          <fpage>356</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Schirmer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gunter</surname>
          </string-name>
          , T.C.:
          <article-title>Temporal signatures of processing voiceness and emotion in sound</article-title>
          .
          <source>Social cognitive and a ective neuroscience 12(6)</source>
          ,
          <volume>902</volume>
          {
          <fpage>909</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Soleymani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jou</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuller</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>S.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pantic</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A survey of multimodal sentiment analysis</article-title>
          .
          <source>Image and Vision Computing</source>
          <volume>65</volume>
          ,
          <issue>3</issue>
          {
          <fpage>14</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Vogt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andre</surname>
          </string-name>
          , E.:
          <article-title>Improving automatic emotion recognition from speech via gender di erentiaion</article-title>
          .
          <source>In: LREC</source>
          . pp.
          <volume>1123</volume>
          {
          <issue>1126</issue>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , Meyer, P.,
          <string-name>
            <surname>Fingscheidt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>On the e ects of speaker gender in emotion recognition training data</article-title>
          .
          <source>In: Speech Communication; 13th ITG-Symposium</source>
          . pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          .
          <string-name>
            <surname>VDE</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>