<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop on Perspectivist Approaches to NLP
* Corresponding author.
$ quanqi.du@uegnt.be (Q. Du); sofie.labat@ugent.be (S. Labat);
thomas.demeester@ugent.be (T. Demeester);
veronique.hoste@ugent.be (V. Hoste)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Unimodalities Count as Perspectives in Multimodal Emotion Annotation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quanqi Du</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sofie Labat</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Demeester</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Veronique Hoste</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LT3, Language and Translation Technology Team, Ghent University</institution>
          ,
          <addr-line>Groot-Brittanniëlaan 45, 9000 Gent</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>T2K, Text-to-Knowledge Research Group, IDLab, Ghent University - imec</institution>
          ,
          <addr-line>Technologiepark-Zwijnaarde 126, 9052 Gent</addr-line>
          ,
          <country country="BE">Belgium</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Most datasets for multimodal emotion recognition only have one emotion annotation for all the modalities combined, which serves as a gold standard for single modalities. This procedure ignores, however, the fact that each modality constitutes a unique perspective that contains its own clues. Moreover, as in unimodal emotion analysis, the perspectives of annotators can also diverge in a multimodal setup. In this paper, we therefore propose to annotate each modality independently and to more closely investigate how perspectives between modalities and annotators diverge. Moreover, we also explore the role of annotator training on perspectivism. We find that for the diferent unimodal levels, the annotations made on text resemble most closely those of the multimodal setup. Furthermore, we see that annotator training has a positive influence on the annotator agreement in modalities with lower agreement scores, but it also reduces the variety of perspectives. We therefore suggest that a moderate training which still values the individual perspectives of annotators might be beneficial before starting annotations. Finally, we observe that negative sentiment and emotions tend to be annotated more inconsistently across the diferent modality setups.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Multimodal versus unimodal emotion annotation</kwd>
        <kwd>Annotator agreement</kwd>
        <kwd>Emotion analysis</kwd>
        <kwd>Perspectivism in NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        1. Within modalities: How often do unimodal and
multimodal emotion annotations of the same
video snippet share the same emotion states? Can
we discern any unimodalities that dominate oth- are dificult to annotate due to their subjectivity, leading
ers and lead to the multimodal emotion state? to low inter-annotator agreement (IAA). On the other
(RQ1) hand, valence is less subjective and in some sense, it
2. Beyond modalities: Perspectivists advocate that shares the same connotation with sentiment, where both
gold standards do not reflect the multiple perspec- of them are scaled into positive, negative, neutral and
tives through which annotations are collected. sometimes some intermediate values, e.g., lightly positive
What is the efect of annotator training on the or very negative. Therefore, in some datasets, only
sentisubsequent annotation behaviour for both uni- ment was annotated. CMU-MOSEI [21] is such a dataset
modal and multimodal annotation? (RQ2) annotated by three crowdsourced judges with Ekman’s
3. Features of inconsistency: Inconsistency in emo- emotions [11] on a [
        <xref ref-type="bibr" rid="ref3">0,3</xref>
        ] Likert-scale and sentiment on
tion annotations across modalities is expected. a [
        <xref ref-type="bibr" rid="ref3">-3,3</xref>
        ] Likert-scale. The 3,228 videos in CMU-MOSEI
Which tendencies can we discern in this incon- make it one of the largest datasets for sentiment analysis
sistency? (RQ3) and emotion recognition. Taking the same annotation
strategy as CMU-MOSEI, the multimodal emotion dataset
MELD [22] evolved from the textual emotional dataset
2. Related Research EmotionLines [23] and achieved higher agreement,
suggesting that the additional modalities were instrumental
      </p>
      <sec id="sec-1-1">
        <title>2.1. Multimodal Emotion Annotation for the annotation improvement.</title>
        <p>
          Emotion annotation is not trivial. Emotions are too com- While the previous datasets were annotated with
emoplex to have a universally accepted standard taxonomy tion labels at the multimodal level, the Chinese CH-SIMS
or annotation scheme. Normally, emotions are either an- [8] dataset was annotated both at unimodal and
mulnotated along categorical or dimensional frameworks. In timodal level, albeit with less fine-grained annotations
the categorical emotion description, anger, disgust, fear, than the previously mentioned corpora since only
senhappiness, sadness and surprise are usually seen as the six timent was annotated. The dataset contains 2,281 video
most basic universal emotions [11]. This model has been segments annotated by five independent students with
extended by Plutchik [12] to cover two more emotions integers from [
          <xref ref-type="bibr" rid="ref1">-1,1</xref>
          ] for negative, neutral and positive
(anticipation and trust). Dimensional models, on the other sentiment. For the purpose of regression and multiclass
hand, project emotions in a multidimensional space with classification tasks, the annotations were then averaged
usually three axes, namely valence, arousal and domi- and divided into five clusters as negative, weakly
neganance [13]. Sometimes less (e.g., valence and arousal only tive, neutral, weakly positive and positive. During a closer
[14]) or more dimensions (e.g., unpredictability [15] or examination of the annotation results, it was found that
appraisal dimensions [16]) can also be used. the sentiment diference between the modalities was not
        </p>
        <p>Most multimodal emotion corpora are annotated with distributed evenly, with audio and the multimodal setup
either categorical or dimensional labels, or both. IEMO- showing a minimal diference while video and text
exCAP [17] is one popular dataset with 10,039 turns of hibited maximal diference [ 8]. As a pioneering study
acted conversations. The recordings were manually seg- in unimodal and multimodal sentiment annotation,
CHmented into conversation turns and then annotated with SIMS ofers inspiring findings on quantified sentiment
both discrete categorical emotion labels (i.e., Ekman’s relationships among diferent modalities. Compared with
six basic emotions [11] complemented with frustration, sentiment, emotion annotations would give more
fineexcited, neutral), and continuous dimensional labels (i.e., grained insights in the variety of emotions expressed in
valence, arousal and dominance on a five-point Likert- diferent modalities. In our study, we aim to tackle this
like scale). It was argued that the combination of both challenge by annotating fine-grained emotions both at
emotion frameworks could provide complementary infor- the unimodal and multimodal level.
mation on how emotion was displayed in real life, since
the categorical level could not give insights in the inten- 2.2. Perspectivism in Emotion Analysis
sity level of emotions [17]. In MSP-IMPROV [18], the
same two approaches were adopted, but a diferent set
of discrete labels was used (i.e., happy, angry, sad,
neutral and other) to collect emotion annotations on 8,438
sentences through crowdsourcing and majority voting.</p>
        <p>Although annotating emotions along the diferent
dimensions could relieve the burden of choosing the
appropriate categorical framework, numerous studies [19, 20]
have shown that the dimensions arousal and dominance
Data perspectivism, as the name suggests, is a recently
popular paradigm for data annotation, which advocates
integrating the diversity of human subjects’ opinions
in annotations and in the knowledge representations
machine learning models [24]. Traditionally, it is quite
often the case that annotators have diferent opinions,
but this disagreement is usually resolved through some
aggregation methods, such as majority voting in MELD
[22] or averaging in CH-SIMS [8]. The aggregation
process is named ground truthing, where a ground truth outline the design of three annotation rounds.
or gold standard is constructed. However, in subjective
tasks such as sentiment analysis and emotion recognition, 3.1. Data Collection and Annotators
there are often cases where there is no ground truth but
just diferent perspectives, and the creation of a ground For our pilot study, we collected emotion-rich videos
truth results in a loss of subtle, but valuable nuances from YouTube. YouTube videos are more easily
availin annotations [25, 26]. Also, annotation aggregation able than dramas, soap operas and movies. Furthermore,
may unfairly cause an under-representation of certain rather than acted emotions, these videos contain natural
annotators’ perspectives [27]. expressions of emotion. The collected dataset consists</p>
        <p>To make full use of the diferent annotations and cap- of 94 video clips of reviews, each of which last for about
ture the contextual nuances, some machine learning re- 10 seconds, which is longer than the average length in
searchers proposed to use diferent annotations as soft popular datasets, e.g., CH-SIMS (3.67 seconds) [32],
CMUlabels [28], instead of using the ground truth as hard la- MOSEI (7.28 seconds) [21], M3ED (7.39 seconds)[33],
bels. In this way, improvements of accuracy on speech MELD (about 8 seconds) [22]. We believe that 10 seconds
emotion detection were reported by using soft labels is a suficient time length to allow annotators to detect
that incorporated knowledge collected from all annota- emotion states in each of the independent unimodalities.
tors [28]. While yielding performance improvement, it This pilot corpus covers 14.75 minutes in total. The clips
also helped to solve the paucity of training data by utiliz- were assigned to three annotators who each annotated
ing ambiguous emotional utterances without dominant the clips separately, eventually leading to three sets of
targets [29]. annotations for the full corpus. These three annotators</p>
        <p>Instead of considering multi-annotator modeling as a are students from Ghent University who are proficient
multi-label problem where each annotator’s labels are in English. Before annotation started, the annotators
reseen as a perspective on the same task, the multi-task ceived instructions on the chosen emotion framework
approach attempts to learn multiple perspectives as sepa- and the custom-designed annotation interface, as shown
rate classification tasks. In computer science, multi-task in Figure 1.1
learning aims to leverage useful information contained in
multiple related tasks to help improve the general
performance of all tasks [30]. In our case, diferent annotators’
labels can be considered as input of diferent tasks. When
there are enough annotations contributed by each
annotator, the multi-task model shows significantly better
performance and higher robustness with lower standard
deviation, but without a significant drop in eficiency
although the annotations as input are multiplied [31].</p>
        <p>In the field of multimodal emotion recognition, when
drawing an analogy between diefrent annotators’ per- Figure 1: The custom-designed annotation interface.
spectives and independent modality annotations, it is
reasonable to consider multimodal emotion recognition as a
multi-task learning problem containing as subtasks the 3.2. Annotation Method
detection of emotions in diferent modality combinations.</p>
        <p>In the experiments of Yu et al. [8] for sentiment analy- For most multi-modal emotion datasets, there is only one
sis, it was found that multi-task models outperformed unified emotion label for each video clip. Inspired by
single-task models for most of the evaluation metrics. ifndings from [ 8] regarding the diference in sentiment</p>
        <p>In what follows, we investigate how perspectives can annotations between diferent modalities, we decided
be leveraged for emotion analysis. More precisely, we to annotate each video clip on four levels, namely text,
look at perspectives between (i) diferent modality levels audio, video (without audio) and all (all three modalities
and (ii) diferent annotators. combined). To ensure other modalities did not interfere
in their judgements, annotators received the four setups
of raw materials in a shufled order, meaning that there
3. Method were time gaps between an annotator seeing diferent
setups of the same video clip. It should be noted that
for the audio modality, annotators were instructed to
To obtain high-quality data, we carefully designed a pilot
study in which we monitored the annotation process. In
this section, we motivate our selection of multimodal
data, discuss the fine-grained annotation framework and</p>
        <sec id="sec-1-1-1">
          <title>1The frame in the figure is taken from the following video: https:</title>
          <p>//www.youtube.com/watch?v=kbn7gOXedSo.
mark the emotion by focusing on the clues in the audio clips for the four modality levels. The annotation results
only, and by ignoring the words in the speech as much before annotator training were grouped into three
cateas possible. gories based on their valence score agreement, namely</p>
          <p>Two emotion frameworks were adopted in the anno- full agreement (3 annotators agreeing), little agreement (2
tation, namely a categorical and a dimensional frame- annotators agreeing) and disagreement (no one agrees).
work. In order to cover a wide diversity of emotions, Ten video clips in each category were randomly selected
we opted for the 25 fine-grained emotion taxonomy pro- as the test set for the after-training session, leading to a
posed by Shaver et al. [34], including anger, contentment, subset of 30 video clips to be annotated. These three
sesdisappointment, disgust, enthrallment, enthusiasm, envy, sions (i.e., before annotator training, joint gold standard
fear, frustration, irritation, joy, longing, love, lust, ner- annotation and after annotator training) serve unique
vousness, optimism, pity, pride, rejection, relief, remorse, functions. The annotator training session resulted in a
sadness, sufering , surprise, and torment. A neutral label set of fully agreed annotations and more insights into
was also added in case there is no emotion present. In the way how the other annotators perceive emotions; the
case there was an emotion that did not match any of sessions before and after the training could be seen as
the provided labels, the annotators were allowed to cus- annotation processes without and with knowledge of an
tomize an emotion label in their own words. For the adjudicated gold standard.
dimensional framework, valence and arousal were
annotated on an analogue-visual five-point Self-Assessment
Manikin (SAM) scale [35]. Dominance was not included, 4. Perspectivism analysis in
as multiple studies on emotion annotation [19, 20] had multimodal annotation
shown that annotator agreement on dominance was too
low to be useful.</p>
        </sec>
      </sec>
      <sec id="sec-1-2">
        <title>3.3. Multiple Rounds of Annotation</title>
        <p>Although we obtained rich annotations for valence,
arousal and categorical emotion labels, in the
following part we first focus on valence analysis, which is less
subjective than arousal and easier to quantify than
categorical emotion labels. While the analysis on annotator
level before and after training aims to explain the
relationships among modality levels (RQ 1), the investigation
on inter-annotator level tries to probe into the changes
of annotators’ perspectives after the training session (RQ
2). In the last part of this section, we focus on the
inconsistency in emotion annotation for the diferent modality
levels (RQ3).</p>
        <p>Annotations were obtained in three sessions divided by
the training of the annotators, which is the process of
gold standard production by taking into account
multiple perspectives. Before training, three annotators were
given a minimum set of initial guidelines, and were then
allowed to do the annotations freely on the 94 video clips,
without interference from each other. During the training
session, the annotators were gathered to jointly annotate
a subset of the 94 video clips, with discussion and
negotiation. This subset for training consisted of 11 video clips 4.1. Polarity annotation analysis before
for which the annotators had to annotate the four
modality levels, and 53 clips for which they only had to annotate training
the multimodal setup. While the former annotation setup After the first annotation round, four subsets of
annotafocused on agreement between diferent modality levels, tions (i.e., one for each modality setup) were obtained
the latter targeted agreement over the video clips as a from the three annotators separately. The three
annotawhole. It took four hours for the three annotators to go tors had no communication with each other before and
through this training session. During the first hour of the during the annotation. The independently obtained
antraining session, the first author of this study sat with the notations were intuitively made by each annotator, and
annotators to guide the discussion. To maximally reduce could serve as a proxy of their personal emotion model.
interference from external factors (e.g., the author), for Since the valence annotations are 5-scale scores
rangthe next three hours, each annotator in turn took on the ing from positive, weakly positive, neutral, weakly
negrole of discussion leader. During the training session, ative, to negative, we took into account these scores to
they were supposed to learn more about each other’s account for the fact that, for example, the diference
bedefinition of emotions and evaluation of the expressed tween positive and neutral should be greater than the
emotions. Since all annotators were involved and ex- diference between weakly positive and neutral. Inspired
plained their views in the discussion, we might consider by Yu et al. [8], we calculated the diference in valence
this process as a kind of weak perspectivism since all per- scores between the four modality levels, which is
formuspectives would be summarized into one single position lated as:
or gold standard [24]. After training, the three
annotators separately annotated another subset of the 94 video
on the video modality, we find that video has the lowest
average valence scores. These results, however, are not
(1) corroborated by annotator 2.
whereby ,  ∈ ℳ are the considered modalities, ℳ = Table 1
{, , , },  is the number of video clips, Average valence scores on 94 annotations before training from
and  and  represent the assigned valence for clip  three annotators.  represents the average scores for each
in modality  and , respectively. modality level, while  2 is the variance. The higher the
va</p>
        <p>With this formula, we obtain the confusion matrices lence scores, the more positive the modalities. The higher the
in Figure 2. We can see that for the second and third an- variance, the less similar the annotations are in valence.
notator, the maximal diference can be observed between text audio video all
the text and video modality. For the first annotator, the  / 2  / 2  / 2  / 2
diference between the audio and video modality annota- 1 3.13/0.97 3.40/0.86 3.07/1.08 3.23/0.97
tions is maximal. Averaging the diference scores across 2 3.13/0.82 3.37/0.72 3.30/0.60 3.16/0.70
the three annotators leads to a 1.10 diference between 3 3.06/1.26 3.10/1.09 2.93/1.17 2.97/1.10
text and video, a little higher than the diference score
of audio-video (1.05). When we consider the minimal
diferences between the four modality levels, we can ob- In order to gain more insights into the valence
annoserve that the text (2 annotators) and audio (1 annotator) tations across the four modality setups, we more closely
annotations are most in line with the "all" multimodal investigated the counts of negative (i.e., 1 and 2), neutral
annotations. (i.e., 3) and positive (i.e., 4 and 5) annotations for the</p>
        <p>To further investigate diferences in valence annota- three annotators, as shown in Table 2. A first interesting
tions between the four modality levels, we averaged the observation to be made is that positive valence is
domivalence score of each modality setup for each of the three nantly annotated across all modalities. While annotator
annotators, as shown in Table 1. A first observation 1 and annotator 2 by a large margin more often recognize
which can be made is that the annotations of annotator 2 positive sentiment in the audio setup, this result is not
show a lower variance compared to those of the other two supported by the annotations of annotator 3. In general,
annotators. The reason is that annotator 2 only once used we found that the annotators detected more positive
sena more extreme sentiment (viz., score 1), while the other timent in audio than in other setups. At the same time,
two annotators used the full scale of valence scores. If we we could observe that the lowest values of positiveness
rank the four modalities by average valence scores, we are associated with the video modality. This indicates
can discern that the average valence score for the audio that annotators detect less often positive emotion states
modality is consistently higher for all three annotators. in the silent video than in other setups. As for
negaAs an important medium for emotional communication, tive emotion states, the lowest values lie in audio and
audio can arouse the audience’s emotional resonance the highest values lie in video (except annotator 2, who
through acoustic features, such as pitch or tone, which always has the lowest values for negativeness among
might explain the overall higher average valence scores. the three annotators), suggesting that annotators tend
In comparison to audio, text as a form of literal expres- to detect less negativeness in the modality of audio than
sion may not be direct and authentic enough, resulting in others, and more negativeness in the modalities of silent
lower average valence scores. Furthermore, when taking video. When putting the annotations all together, it is
the average of the score of annotator 1 and annotator 3 found the all modality setup does not hold the maximum
Relative sentiment counts across three annotators. , , and  mean negative, neutral, and positive, respectively.
1 − 3 means annotator 1 to annotator 3, and  means the average relative count across the three annotators.
1
2
3</p>
        <p>text
neg / neu / pos
pact on the dependent variable, i.e., the multimodal level. diferent modalities. The higher the agreement score, the more
modal level than the other two unimodalities. Our results
notators reach the highest agreement on the text
modalthus seem to confirm insights from previous studies [ 37]
ity, no matter whether these scores are calculated with
which reported a significant drop (30%) in binary
accu</p>
        <sec id="sec-1-2-1">
          <title>Fleiss’ kappa [38] or Krippendorf’s alpha [ 39]. It is as</title>
          <p>racy when removing the text modality in multimodal
expected that annotators have the least agreement over
sentiment analysis, a phenomenon which is termed "text
silent videos, since emotion recognition on non-linguistic
annotations combined from three annotators.</p>
        </sec>
        <sec id="sec-1-2-2">
          <title>The regression results in Table 3 indicate that the vari</title>
          <p>ables text, audio, and silent video have a significant
imWith positive coeficients, the increase of valence scores
in each unimodality is associated with an increase in the
multimodality level. The t-value and the p-value (&lt;0.05)
indicate that these coeficient estimates are statistically
significant and unlikely to be due to chance. The
coeficient of text is bigger than that of audio, and also
two times bigger than that of video, meaning each
oneunit valence increase in the text modality is associated
with a higher average increase of valence at the
multipredominance" [32].</p>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>4.2. Polarity annotation analysis after</title>
        <p>,  and  modality setup, respectively. 1 −
3 means annotator 1 to annotator 3, 
means the average
of the scores,  2 means the variance of the scores. In each
column the two numbers refer to the results before/after
the
training session. The higher the diference score, the bigger
the sentiment gap between the modality setups.</p>
        <p>Anno1</p>
        <p>Anno2</p>
        <p>Anno3

 2
 −
 −
 −  0.93/1.21 0.71/0.73 1.02/0.91 0.89/0.95 0.02/0.04↑
 −  1.29/1.58 0.98/1.29 1.41/1.47 1.23/1.45 0.03/0.01↓
 0.88/1.20 0.75/0.95 1.20/0.86 0.94/1.00 0.04/0.02↓
 −  1.29/1.20 0.73/1.13 0.98/1.00 1.00/1.00 0.05/0.01↓
 0.88/0.95 0.58/0.61 0.86/0.66 0.77/0.74
ities before ( ) and after (  ) annotator training.   means
categorical  , and   means interval  . The higher the
agreement score, the more similar perspectives the annotators have.</p>
        <p>text
bf
af</p>
        <p>audio
bf
af</p>
        <p>video
bf
af</p>
        <p>all
bf
af</p>
      </sec>
      <sec id="sec-1-4">
        <title>4.3. Inconsistency Analysis</title>
        <p>In addition to investigating the valence annotations of
the diferent annotators for the four modality levels, we
also performed an inconsistency analysis at the video clip
level. The annotation inconsistency among modalities
is considered to contain useful information as it might
give more insights on what makes some video segments
more dificult to label than others. Furthermore, related
work [8] also suggests that the stronger the inconsistency,
the better the complementarity of intermodal fusion. In
this section, we briefly discuss the inconsistency
distribution in our corpus. In doing so, we not only focus on
valence, but also on more fine-grained emotions.
4.3.1. Inconsistency Distribution of Sentiment
For the inconsistency analysis, we had a closer look at
the inconsistent annotations across modalities in the full
corpus of 94 video clips; the full corpus was chosen for
this analysis in order to have a suficient amount of
annotations for the analysis. We calculated the inconsistency
score in a similar way to the diference score, formulated
as:
 =
√︃ 1 ∑︁
︀( 4)︀
2 (,)∈ℳ
( − )2,
 ∈</p>
        <p>(2)
whereby (︀ 4)︀ is the number of diferent ways to select</p>
        <p>2
two modalities from a set of four, ,  ∈ ℳ are the
considered modalities, ℳ = {, , , }, 
is the number of video clips, and  and  represent the
assigned valence for clip  in modality  and ,
respectively. The sentiment polarities are the labels assigned to
the multimodal setup.</p>
        <p>As shown in Figure 3, we can see there are some
annotations with an inconsistency score of zero, which means
they have the same valence score for the four modality
levels. These consistent annotations account for 25.5%
of the total number of 282 annotations across the three
annotators, and among these annotations, the positive
sentiment accounts for about 60%, while the remainder
of the annotations is fairly equally spread over the other
two sentiment states. The rest of the annotations were
divided in two groups with diferent degrees of
inconsistency: those with the top 25% (specifically 23.8%, as
this is the nearest group of consistency scores) greatest
inconsistency scores were considered as strongly
inconsistent annotations, and the remainder of the annotations
(50.7%) were considered as weakly inconsistent. The
cutof point for this division between the two groups was
an inconsistency score of 1.35.</p>
        <p>As shown in Table 7, when comparing the total
annotations and the consistent annotations, it is found that the
percentage of negative sentiment drops down from 24.8%
in total annotations to 19.4% in consistent annotations,
and the percentage of positive sentiment increases from
49.3% in total annotations to 58.3% in consistent
annotations, while the percentage of neutral sentiment
experiences nearly no change. Therefore, it is indicated that
positive sentiment is associated with more consistent
annotations across modalities. On the other hand, although
there are no significant changes in sentiment distribution
from total annotations to inconsistent annotations, we
recognize changes in the sentiment distribution when
following the above concept of strong inconsistency. It is
noticed that the positive sentiment is distributed nearly
equally in weakly and strongly inconsistent annotations,
while negative sentiment is more located in strong
inconsistency and neutral sentiment is more located in weak
inconsistency. We hypothesize that people tend to fully
show their positive sentiment across modalities, while
they are less encouraged to fully show the negative
sentiment, resulting in strong inconsistency across modalities.</p>
        <p>At the same time, we cannot rule out other possibilities
due to the small size of our dataset.
4.3.2. Inconsistency in Emotions</p>
        <p>The strong inconsistency in valence scores (as
illustrated in Figure 4) also gives insights into emotion
inconsistency, since varied polarities are linked to varied
emotions. The strongly inconsistent instances (as
introduced in Section 4.3.1) were taken as a starting point for
the investigation into the inconsistency in the emotion
labels. To this end, we took the multimodal setup as
standard and counted the number of times each unimodal
modality (text, audio, and video) received the same
categorical emotion annotation as the multimodal setup. The
Table 8 One predominant unimodality? More than fifty
The number of cases in which emotion annotations at the years ago, Albert Mehrabian [41] presented an equation
unimodal level correspond to annotations at the multimodal of feeling as Total feeling = 7% verbal feeling + 38%
volevel. cal feeling + 55% facial feeling, according to which facial
emotion multimodal text audio video expressions contribute the most to multimodal emotion.</p>
        <p>Recently, Liu et al. [42] verified that facial expressions
 3 2 1 1 indeed play a more dominant role than emotive markers
 221 120 110 111 tohfattexthteinmeomdoatliiotyn opfertceexptthioans.mHoorweeevfeecrt, oonurmreuslutilmtsosdhaolw
ℎ 4 3 3 0 emotion recognition than the modality of silent video.
 7 1 2 5 This diference in results might be attributed to the data
 1 1 0 0 used for the experiments. First, the images of speakers
 1 0 0 1 with facial expressions in Liu’s experiments [42] were
 4 3 3 1 reproduced from the Amsterdam Dynamic Facial
Expres 1 0 1 0 sion Set [43], in which the facial expression of a particular
 3 3 0 0 emotion was intentionally portrayed by actors. We can
 5 2 3 1 expect these acted emotions to be much more
outspo 34 18 15 12 ken than emotional expressions in genuine interaction in
real-life scenarios [44]. The facial expressions in our data
come from non-acted recordings and we can assume that
the genuine emotions conveyed in our data are far more
results of this procedure are shown in Table 8. We found complex and subtle than acted emotions. Furthermore, in
that in the 67 instances that were classified as strongly our pilot study, we used dynamic displays (silent videos)
inconsistent, only 34 instances have at least one match in of emotions instead of static pictures of facial expressions.
emotion label with the unimodal setups. In other words, It has been observed that the identification accuracy of
in nearly half of the cases, each unimodal setup has a facial expressions reaches near perfection when using
diferent emotion label than the multimodal setup. Upon images portraying fully developed and intense emotional
further investigation, we found that there were 18 cases states [45]. Finally, we also observed that, although the
where the annotations of the text modality were con- annotators were instructed to score the audio sentiment
sistent with the multimodal setup. The corresponding only with the acoustic clues rather than the language in
numbers for audio and video were 15 and 12, respectively. spoken form, it was inevitable that the spoken language</p>
        <p>Again, the text modality seems to share the highest was automatically transcribed in their mind and exerted
consistency with the all modality setup in terms of emo- influence on the annotation. The annotation of audio is
tion labels, which is in line with the results we obtained thus not totally independent from the modality of text
for the polarity annotations. Furthermore, some emo- and one of the possible reasons why the results of text
tions seem to be more associated with one specific modal- and audio are intertwined, as shown in Figure 2.
ity. For example, the emotion joy has a stronger
association with the video modality, as shown in Table 8: the
video modality has 5 consistent cases, while the other two
modalities have 1 and 2 instances, respectively. This kind
of association within inconsistent cases, though weak,
gives a clue on how to explain the emotion labels of the
multimodal setup.</p>
        <p>To train or not to train? The training session in our
annotation experiments could be seen as the
construction of an adjudicated gold standard, the use of which
is believed to have some negative efects as it ignores
the diversity of opinions [24] due to the reduction of
perspectives. To quantitatively measure this change of
annotators’ perspectives, we measured inter-annotator
5. Discussion agreement before and after training. Our results show
that after the training session, inter-annotator agreement
Each of the modalities in multimodal communication of- indeed generally increased and was especially beneficial
fers a unique perspective on the communication. Imagine for the annotation of silent video. Overall, however, IAA
that someone has a hearing impairment and sufers from rates remained modest for all unimodalities, safeguarding
prosopagnosia, meaning that this person can read but suficient diversity in annotation perspectives.
not hear nor recognize facial expressions, one blind who
cannot see but hear, and one deaf-mute who cannot lip- Any benefits to inconsistency? Emotion expression
read. When standing in their shoes, we can experience varies in diferent modalities, and speech and facial
exthe perspectives represented by the modalities of text, pressions are under the control of diferent muscles.
Faaudio, and silent video in isolation.
cial expressions involve the movement of facial muscles, multimodal emotion recognition.
while speech involves the movement of vocal cords and After a first annotation round, we organised a
trainmuscles in the mouth. These diferent muscle combina- ing session to build an adjudicated gold standard and
tions and physiological processes can lead to varying to make the three annotators more acquainted with the
rates of emotional changes across diferent modalities. other annotators’ perspectives. From the valence
annotaTherefore, by observing the rapid changes in emotional tions before training, we concluded that the text modality
expression in one modality, it may be inferred that there seems to dominate the other modalities and resembles
will be corresponding changes in the emotional expres- most closely the multimodal annotations. After
trainsion of another modality, even though these changes ing, perspective diversity was reduced on the annotator
might not be as apparent or may occur at a slower pace. level, as evidenced by the general decrease of standard
As is shown in Figure 5, sometimes the sentiment scores deviation, and on the inter-annotator level, but the fairly
of audio and video are at diferent changing rates. The modest IAA scores and more stability in IAA across the
ones in blue circles show the diferent changing rates in diferent modalities might advocate for a training session.
positive sentiment, while the ones in red circles represent An analysis of the emotional annotation inconsistency
the diference in negative sentiment. among modalities showed that the inconsistency has
almost equal distribution in positive and negative
sentiments, and that inconsistency is most often located in
the modality of silent video. We believe that this
inconsistency is of special interest for future studies in more
ifne-grained emotion modeling. In future studies, apart
from scaling up the dataset, we will also investigate in
more depth the other annotation layers, such as the
multilabel emotion annotations, the trigger of the emotions,
and the annotation time.</p>
        <sec id="sec-1-4-1">
          <title>Inconsistency might also give more insight into com</title>
          <p>plex emotions. For example, text-only annotations do not
allow to detect all types of irony, but an inconsistent
audio or facial expression could account for the irony efect.
Of course, basic emotions [11] can also be complicated
when emotions in diferent modalities contradict each
other, as shown in Figure 4. However, from a
computational perspective, this inconsistency might help models
to learn more diferentiated information and improve the
complementarity between modalities [8].</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>6. Conclusion</title>
      <p>In this paper, we presented a pilot study on the basis
of a newly collected video corpus which we annotated
both at the level of the single modalities (text, speech,
video) and the multimodal level. Both dimensional and
categorical emotion annotation were provided for these
four annotation setups, based on the assumption that
unimodalities can also serve as unique perspectives in</p>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgments</title>
      <sec id="sec-3-1">
        <title>This research received funding from the Flemish Govern</title>
        <p>ment under the Research Program Artificial Intelligence
(174Z05623) and from the Research Foundation Flanders
(FWO-Vlaanderen) with grant number 1S96322N. We
would also like to thank the anonymous reviewers for
their valuable and constructive feedback.
grammar of visual design, Psychology Press, 1996. An acted corpus of dyadic interactions to study
emodoi:10.5860/choice.34-1950. tion perception, IEEE Transactions on Afective
[6] S. K. D’Mello, J. K. Westlund, A review and Computing 8 (2016) 67–80. doi:10.1109/TAFFC.
meta-analysis of multimodal afect detection sys- 2016.2515617.
tems, ACM Computing Surveys 47 (2015) 1– [19] L. De Bruyne, O. De Clercq, V. Hoste,
Annotat36. URL: https://doi.org/10.1145/2682899. doi:10. ing afective dimensions in user-generated content:
1145/2682899. Comparing the reliability of best–worst scaling,
[7] J. Chen, C. Sun, S. Zhang, J. Zeng, Cross-modal dy- pairwise comparison and rating scales for
annotatnamic sentiment annotation for speech sentiment ing valence, arousal and dominance, Language
Reanalysis, Computers and Electrical Engineering sources and Evaluation (2021) 1–29. doi:10.1007/
106 (2023) 108598. doi:10.1016/j.compeleceng. s10579-020-09524-2.</p>
        <p>2023.108598. [20] S. Labat, T. Demeester, V. Hoste, Emotwics: A
cor[8] W. Yu, H. Xu, F. Meng, Y. Zhu, Y. Ma, J. Wu, J. Zou, pus for modelling emotion trajectories in dutch
K. Yang, CH-SIMS: A Chinese multimodal sen- customer service dialogues on twitter, Language
timent analysis dataset with fine-grained annota- Resources and Evaluation. Accepted (2022). URL:
tion of modality, in: Proceedings of the 58th An- http://hdl.handle.net/1854/LU-8769949.
nual Meeting of ACL, ACL, Online, 2020, pp. 3718– [21] A. Bagher Zadeh, P. P. Liang, S. Poria, E. Cambria,
3727. URL: https://aclanthology.org/2020.acl-main. L.-P. Morency, Multimodal language analysis in
343. doi:10.18653/v1/2020.acl-main.343. the wild: CMU-MOSEI dataset and interpretable
[9] K. L. O’Halloran, Multimodal discourse analysis, dynamic fusion graph, in: Proceedings of the 56th
Companion to Discourse. London and New York: Annual Meeting of ACL, ACL, Melbourne, Australia,
Continuum (2011). 2018, pp. 2236–2246. URL: https://aclanthology.org/
[10] C. E. Jewitt, The Routledge handbook of multimodal P18-1208. doi:10.18653/v1/P18-1208.</p>
        <p>analysis, Routledge/Taylor &amp; Francis Group, 2011. [22] S. Poria, D. Hazarika, N. Majumder, G. Naik, E.
Cam[11] P. Ekman, An argument for basic emotions, Cog- bria, R. Mihalcea, MELD: A multimodal multi-party
nition &amp; Emotion 6 (1992) 169–200. doi:10.1080/ dataset for emotion recognition in conversations,
02699939208411068. in: Proceedings of the 57th Annual Meeting of
[12] R. Plutchik, A general psychoevolutionary theory ACL, ACL, Florence, Italy, 2019, pp. 527–536. URL:
of emotion, in: Theories of emotion, Elsevier, 1980, https://aclanthology.org/P19-1050. doi:10.18653/
pp. 3–33. doi:10.1016/B978-0-12-558701-3. v1/P19-1050.</p>
        <p>50007-7. [23] S.-Y. Chen, C.-C. Hsu, C.-C. Kuo, T.-H. Huang,
[13] A. Mehrabian, J. A. Russell, An approach to envi- L.-W. Ku, EmotionLines: An emotion corpus
ronmental psychology, MIT Press, 1974. of multi-party conversations, in: Proceedings
[14] J. A. Russell, A circumplex model of afect, Journal of the Eleventh ICLRE, ELRA, Miyazaki, Japan,
of personality and social psychology 39 (1980) 1161. 2018, pp. 1597–1601. URL: https://aclanthology.org/
doi:10.1037/H0077714. L18-1252.
[15] J. R. Fontaine, K. R. Scherer, E. B. Roesch, P. C. [24] V. Basile, F. Cabitza, A. Campagner, M. Fell, Toward
Ellsworth, The world of emotions is not two- a perspectivist turn in ground truthing for
predicdimensional, Psychological science 18 (2007) 1050– tive computing, arXiv preprint arXiv:2109.04270
1057. doi:10.1111/j.1467-9280.2007.02024. (2021). doi:10.48550/arXiv.2109.04270.
x. [25] L. Aroyo, C. Welty, Crowd truth: Harnessing
dis[16] E. Troiano, L. Oberländer, R. Klinger, Dimen- agreement in crowdsourcing a relation extraction
sional modeling of emotions in text with appraisal gold standard, in: ACM Web Science 2013, 2015.
theories: Corpus creation, annotation reliability, doi:10.6084/M9.FIGSHARE.679997.V1.
and prediction, Computational Linguistics 49 [26] H. M. Fayek, M. Lech, L. Cavedon, Modeling
sub(2023) 1–72. URL: https://doi.org/10.1162/coli_a_ jectiveness in emotion recognition with deep
neu00461. doi:10.1162/coli_a_00461. ral networks: Ensembles vs soft labels, in: 2016
[17] C. Busso, M. Bulut, C.-C. Lee, A. Kazemzadeh, IJCNN, 2016, pp. 566–570. doi:10.1109/IJCNN.</p>
        <p>E. Mower, S. Kim, J. N. Chang, S. Lee, S. S. 2016.7727250.</p>
        <p>Narayanan, IEMOCAP: Interactive emotional [27] V. Prabhakaran, A. Mostafazadeh Davani, M. Diaz,
dyadic motion capture database, Language Re- On releasing annotator-level labels and
informasources and Evaluation 42 (2008) 335–359. doi:10. tion in datasets, in: Proceedings of the Joint
1007/s10579-008-9076-6. 15th LAW and 3rd DMRW, ACL, Punta Cana,
Do[18] C. Busso, S. Parthasarathy, A. Burmania, M. Abdel- minican Republic, 2021, pp. 133–138. URL: https:
Wahab, N. Sadoughi, E. M. Provost, MSP-IMPROV: //aclanthology.org/2021.law-1.14. doi:10.18653/
v1/2021.law-1.14. (1971) 378. doi:10.1037/h0031619.
[28] H. M. Fayek, M. Lech, L. Cavedon, Modeling sub- [39] K. Krippendorf, Content analysis: An introduction
jectiveness in emotion recognition with deep neu- to its methodology, Sage Publications, 2018.
ral networks: Ensembles vs soft labels, in: 2016 [40] E. G. Krumhuber, A. Kappas, A. S. Manstead,
EfIJCNN, 2016, pp. 566–570. doi:10.1109/IJCNN. fects of dynamic aspects of facial expressions: A
2016.7727250. review, Emotion Review 5 (2013) 41–46. doi:10.
[29] A. Ando, S. Kobashikawa, H. Kamiyama, R. Ma- 1177/1754073912451349.
sumura, Y. Ijima, Y. Aono, Soft-target training [41] A. Mehrabian, Silent messages, Wadsworth
Belwith ambiguous emotional utterances for dnn- mont, 1971.
based speech emotion classification, in: 2018 [42] M. Liu, J. Schwab, U. Hess, Language and face in
IEEE ICASSP, 2018, pp. 4964–4968. doi:10.1109/ interactions: Emotion perception, social meanings,
ICASSP.2018.8461299. and communicative intentions, Frontiers in
Psy[30] Y. Zhang, Q. Yang, A survey on multi-task learning, chology 14 (2023) 1146494. doi:10.3389/fpsyg.</p>
        <p>IEEE Transactions on Knowledge and Data Engi- 2023.1146494.
neering 34 (2021) 5586–5609. doi:10.1109/TKDE. [43] J. Van Der Schalk, S. T. Hawk, A. H. Fischer,
2021.3070203. B. Doosje, Moving faces, looking places: Validation
[31] A. M. Davani, M. Díaz, V. Prabhakaran, Dealing of the Amsterdam Dynamic Facial Expression Set,
with disagreements: Looking beyond the majority Emotion 11 (2011) 907. doi:10.1037/a0023853.
vote in subjective annotations, Transactions of [44] E. Douglas-Cowie, L. Devillers, J.-C. Martin,
ACL 10 (2022) 92–110. URL: https://aclanthology. R. Cowie, S. Savvidou, S. Abrilian, C. Cox,
Mulorg/2022.tacl-1.6. doi:10.1162/tacl_a_00449. timodal databases of everyday emotion: Facing up
[32] Y. Liu, Z. Yuan, H. Mao, Z. Liang, W. Yang, Y. Qiu, to complexity, in: Ninth ECSCT, 2005, pp. 813–816.</p>
        <p>T. Cheng, X. Li, H. Xu, K. Gao, Make acoustic doi:10.21437/Interspeech.2005-381.
and visual cues matter: CH-SIMS v2.0 dataset and [45] J. M. Carroll, J. A. Russell, Facial expressions
AV-Mixup consistent module, in: Proceedings of in hollywood’s protrayal of emotion, Journal of
the 2022 ICMI, ACM, New York, NY, USA, 2022, Personality and Social Psychology 72 (1997) 164.
p. 247–258. URL: https://doi.org/10.1145/3536221. doi:10.1037/0022-3514.72.1.164.
3556630. doi:10.1145/3536221.3556630.
[33] J. Zhao, T. Zhang, J. Hu, Y. Liu, Q. Jin, X. Wang, H. Li,</p>
        <p>M3ED: Multi-modal multi-scene multi-label
emotional dialogue database, in: Proceedings of the 60th
Annual Meeting of ACL (Volume 1: Long Papers),
ACL, Dublin, Ireland, 2022, pp. 5699–5710. URL:
https://aclanthology.org/2022.acl-long.391. doi:10.</p>
        <p>18653/v1/2022.acl-long.391.
[34] P. Shaver, J. Schwartz, D. Kirson, C. O’connor,
Emotion knowledge: Further exploration of a prototype
approach, Journal of Personality and Social
Psychology 52 (1987) 1061. doi:10.1037//0022-3514.</p>
        <p>52.6.1061.
[35] M. M. Bradley, P. J. Lang, Measuring emotion: The
self-assessment manikin and the semantic
diferential, Journal of Behavior Therapy and
Experimental Psychiatry 25 (1994) 49–59. doi:10.1016/
0005-7916(94)90063-9.
[36] R. A. Fisher, Statistical methods for research
work</p>
        <p>ers, Springer, 1992.
[37] X. Li, M. Chen, Multimodal sentiment analysis
with multi-perspective fusion network focusing on
sense attentive language, in: Proceedings of the
19th CNCCL, CIPSC, Haikou, China, 2020, pp. 1089–
1100. URL: https://aclanthology.org/2020.ccl-1.101.</p>
        <p>doi:10.1007/978-3-030-63031-7_26.
[38] J. L. Fleiss, Measuring nominal scale agreement
among many raters, Psychological Bulletin 76</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Picard</surname>
          </string-name>
          , Afective Computing, MIT Press,
          <year>1997</year>
          . doi:
          <volume>10</volume>
          .7551/mitpress/1140.001.0001.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Labat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Amir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Hoste</surname>
          </string-name>
          ,
          <article-title>An emotional journey: Detecting emotion trajectories in Dutch customer service dialogues, in: Proceedings of the Eighth WNUT</article-title>
          , ACL,
          <year>2022</year>
          , pp.
          <fpage>106</fpage>
          -
          <lpage>112</lpage>
          . URL: https://aclanthology.org/
          <year>2022</year>
          .wnut-
          <volume>1</volume>
          .12.pdf .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Akçay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Oğuz</surname>
          </string-name>
          ,
          <article-title>Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers</article-title>
          ,
          <source>Speech Communication</source>
          <volume>116</volume>
          (
          <year>2020</year>
          )
          <fpage>56</fpage>
          -
          <lpage>76</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.specom.
          <year>2019</year>
          .
          <volume>12</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>W.</given-names>
            <surname>Mellouk</surname>
          </string-name>
          , W. Handouzi,
          <article-title>Facial emotion recognition using deep learning: Review and insights</article-title>
          ,
          <source>Procedia Computer Science</source>
          <volume>175</volume>
          (
          <year>2020</year>
          )
          <fpage>689</fpage>
          -
          <lpage>694</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.procs.
          <year>2020</year>
          .
          <volume>07</volume>
          .101.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G. R.</given-names>
            <surname>Kress</surname>
          </string-name>
          ,
          <string-name>
            <surname>T. Van Leeuwen</surname>
          </string-name>
          , Reading images: The
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>