<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Transparent Framework towards the Context-Sensitive Recognition of Conversational Engagement</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexander Heimerl</string-name>
          <email>heimerl@hcm-</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tobias Baur</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elisabeth Andre´</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>7</fpage>
      <lpage>16</lpage>
      <abstract>
        <p>Modelling and recognising affective and mental user states is an urging topic in multiple research fields. This work suggests an approach towards adequate recognition of such states by combining state-of-the-art behaviour recognition classifiers in a transparent and explainable modelling framework that also allows to consider contextual aspects in the inference process. More precisely, in this paper we exemplify the idea of our framework with the recognition of conversational engagement in bi-directional conversations. We introduce a multi-modal annotation scheme for conversational engagement. We further introduce our hybrid approach that combines the accuracy of state-of-the art machine learning techniques, such as deep learning, with the capabilities of Bayesian Networks that are inherently interpretable and feature an important aspect that modern approaches are lacking - causal inference. In an evaluation on a large multi-modal corpus of bi-directional conversations, we show that this hybrid approach can even outperform state-of-the-art black-box approaches by considering context information and causal relations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Nowadays, machine learning approaches are most often purely
datadriven as they use so-called ”black-box” approaches that map
lowlevel features or decisions of previous classifiers onto abstract labels
following statistical methods. Here we usually have no transparent
concept of how the model is internally represented, e.g. how and why
weights on the nodes of artificial neural networks are related.</p>
      <p>In most research areas (e.g., in psychology, behaviour analysis,
but also physics), the goal of creating a model is to reason about
observations in the world, while creating and validating theories that
aim to find causation and explanations. Then, such models are often
validated in simulations, or collated with real-world observations.
That means on the one hand, we have data-driven models in
machine learning that do a decent job in creating predictions for a huge
amount of recognition problems, but deliver no transparent way to
understand their decisions and don’t necessarily have a theory behind
them. On the other hand, we have models that aim to explain
interrelations of observations of the world and/or of their inner states. Such
models are also called ”white-box” approaches.</p>
      <p>
        In this paper, we suggest a hybrid approach that combines
stateof-the-art ”black-box” recognition models with a transparent causal
inference model. Lately, the focus of research tends towards deep
end-to-end learning with artificial neural networks. While such
approaches deliver promising results on audio-visual data, they only
give little insight on how and why they predict behaviours the way
they do. In this work, we investigate the recognition of
”conversational engagement”. Especially in scenarios where it is essential to
know why a person’s behaviour is interpreted as, e.g., ”strongly
disengaged”, the idea is often to identify cues that led to this
interpretation, providing an additional abstraction layer. Here, the relevance
of a comprehensible model becomes very clear. Imagine a system
that gives feedback on how engaged a person appeared in a social
coaching scenario. A model should be able to give feedback on why
it decided a person appeared to be strongly engaged or disengaged,
so that a human can learn from the feedback. In order to infer
complex social signals with a transparent model, we combine
predictions of multiple high-precision classifiers with dynamic Bayesian
networks (DBN) [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. DBNs are probabilistic models that allow
expressing causal relationships between nodes in a network, while at
the same time considering previous observations. Even tough the
parameters for such nodes and even the overall network structure may
be learned with machine learning techniques, DBNs allow retracing
the decisions they are making for each node or layer of nodes
visually and are therefore inherently interpretable. While the structure of
a DBN may be modelled based on a theory and grounded in social
sciences, our framework allows to consider parallel observations, so
it can learn correlations between concurrent behaviours, context and
the complex phenomena of interest.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
    </sec>
    <sec id="sec-3">
      <title>Engagement in psychology</title>
      <p>
        Engagement is a complex social attitude. This becomes apparent
when being confronted by the mass of available definitions. In fact
Glas et. al [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] gave an overview of many different engagement
definitions, with some of them being very context specific. The definition
of Poggi coincides best with a general understanding of engagement.
She describes it as: “The value that a participant in an interaction
attributes to the goal of being together with the other participant(s)
and of continuing the interaction.” [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. As complex as it is to
acquire a fitting definition, equally complex is the manifestation of
engagement in conversations. There are multiple behaviours that are
strongly connected to it.
      </p>
      <p>
        In general, body language is an elemental part in expressing
conversational engagement. To be more precise, the alignment of the body
and the limbs play an important role on broadcasting the state of
engagement [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. Interlocutors, that are engaged during a conversation,
align their bodies to each other, as described in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], “to create a
frame of engagement”.
      </p>
      <p>
        However not only the body position and body movement relative to
each other is an important criteria, also the individual body behaviour
is of great interest. Lots of body movement may indicate some kind
of restlessness. This was found to be connected to boredom, which
is a manifestation of low engagement [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Also depending on the
level of engagement the body reacts with more subtle signals. Heart
rate, blood pressure, EEG and galvanic skin response are all potential
candidates to draw conclusions about engagement [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ].
      </p>
      <p>
        Moreover specific gestures may allow to draw conclusions about the
level of engagement. Lausberg [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] investigated, among other things,
the origin of self-touch gestures. She describes, that self-touch
gestures occur when people are emotionally engaged. Alongside
selftouch gestures there are also more complex gestures, that reflect
different affective states [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
      </p>
      <p>
        Another crucial part in human interaction are “Feedback /
Backchannels”. It describes a high-level behaviour that is related to
engagement. Backchannels are a kind of feedback. They occur between
interlocutors and are typically in the form of non-intrusive acoustic or
visual signals, e.g. a simple “Yes” or a headnod. Backchannels are a
tool, to not only signal the success of communication, but also
provide information about the level of engagement [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <p>
        A strong form of engagement manifestation is mirroring of
behaviours, be it acoustic or visual, from one interlocutor by the other.
Those go by the terms “Synchrony”, “Mimicry” or “Alignment”.
All of those represent a connection or bonding between
interlocutors [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Recognition of engagement</title>
      <p>Engagement has been investigated from various research angles, e.g.
how to define engagement, how to annotate engagement or how to
automatically predict engagement. Therefore it is no surprise that
there are many different systems available to automatically predict
engagement.</p>
      <p>
        Rich et al. [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ] introduced a reusable module for the recognition of
engagement in human-robot interaction. They identified four
connection events that they found to be tools for the maintenance of
engagement. The four events were, directed gaze, mutual facial gaze,
adjacency pairs, verbal and non-verbal backchannels. Those
concepts built the theoretical foundation for their engagement
recognition module.
      </p>
      <p>
        Sanghvi et al. [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ] predicted engagement based on body posture
features. All their features have been extracted from video signals. They
identified following important posture features: “Body lean angle”,
“Slouch factor”, “Quantity of motion” and “Contraction index”. For
the classification they used Weka [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and evaluated 63 different
classifiers. The best ones achieved a prediction accuracy of 82% on
the two classes “engaged” and “not engaged”.
      </p>
      <p>
        Roman Bednarik et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] focused on recognising conversational
engagement with gaze data. Further, they introduced an annotation
scheme for the different levels of conversational engagement. They
defined a total of six levels. In ascending order, the first being the
lowest level of engagement and the last being the highest level of
engagement: “No interest”, “Following”, “Responding”, “Conversing”,
“Influencing discussion discourse/topic” and “Governing/managing
discussion”. To ease down the classification task the authors decided
to reduce the six classes of engagement to a two-classes problem
low and high engagement. For the automatic estimation they
computed a total of 26 features from the raw eye gaze data, e.g. number
of fixations, number of saccades, minimal and maximal fixation
duration, minimal and maximal saccade amplitude, quantity of fixation
at the speakers’ face. Those features have been used to train a SVM.
Following this approach they achieved a prediction accuracy of 74%.
Yun et al. [
        <xref ref-type="bibr" rid="ref56">56</xref>
        ] proposed a convolutional neural network(CNN) to
automatically predict engagement of children. For training their CNN
they relied solely on facial images. However due to limited training
data they used CNNs that have been pre-trained on face recognition
tasks. Their network architecture includes a new layer combination to
model temporal dynamics in order to extract high-level features from
low-level features. For predicting engagement they distinguished
between four levels of engagement, high engagement, low engagement,
low disengagement and high disengagement. On the given task their
network architecture achieved a balanced accuracy of 0.7807.
      </p>
      <p>There is already plenty of research available that targets
recognising engagement. However most of the systems focus solely on
finding feasible features, either handcrafted or extracted from
convolutional layers to optimise prediction accuracy. Little attention is
payed to context, which is important when it comes to recognising
engagement in everyday scenarios. Depending on the environment
individuals are in it can affect how people behave and also what kind
of cues they are using during a conversation. Imagine a student
talking to his friend during a break in comparison to a student
attending an oral exam. However not only external factors can influence
the broadcasting of engagement. Also the very unique
psychological traits every person has can influence their behaviour. An
extrovert person in comparison to an introvert person can appear totally
different during a conversation. Those examples illustrate potential
context information that should be considered when recognising
engagement.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Bayesian networks</title>
      <p>
        Bayesian networks have been successfully applied in earlier work
in the area of high-level interpretation of social signals. One of the
pioneer studies is the work by Conati et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. They have
incorporated bio-feedback sensors into a complex emotion model, that was
based on a subset of the emotions proposed by OCC theory [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ].
They employed a dynamic decision network (a generalisation of a
dynamic Bayesian network) to capture many of the complex
phenomena associated with appraisal theories. In particular, their model
estimated student goals based on personality traits and events which
represent changes in the environment (e.g., progress in the system)
as well as evidence from physical feedback channels to support the
model’s prediction.
      </p>
      <p>
        Sabourin et al. [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ] focused, similar to Conati et al., on learners’
emotions, and employed multiple variations of Bayesian networks. More
specifically, they investigated the benefits of using cognitive models
of learner emotions, to guide the development of Bayesian networks
for prediction of student affect. Predictive models were empirically
trained on data, acquired from 260 students interacting with a
gamebased learning environment. As a dynamic Bayesian network turned
out to be the most successful model, they emphasised the importance
of temporal information in predicting learner emotions. They
concluded that predictive models may be used to validate theoretical
models of emotion.
      </p>
      <p>
        Wo¨llmer et al. [
        <xref ref-type="bibr" rid="ref55">55</xref>
        ] combined a hierarchical dynamic Bayesian
network to detect linguistic keyword features together with long
shortterm memory (LSTM) neural networks [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] which model phoneme
context and emotional history to predict the affective state of the user.
This way, they are combining acoustic, linguistic, and long-term
context information to continuously predict the current valence and
activation in a two-dimensional emotion space.
      </p>
      <p>
        Lugrin et al. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] used Bayesian networks to incorporate culture
into intelligent systems by combining theory-based and data-driven
approaches. Their network aims to generate non-verbal
culturedependent behaviours. While the model is structured based on
cultural theories and theoretical knowledge of their influence on
prototypical behaviour, the parameters of the model are learned from a
multi-modal corpus recorded in the German and Japanese cultures.
In their work, they aim to generate adequate behaviours for an agent
to show, based on its simulated culture.
      </p>
      <p>
        Finally, one could conclude that (dynamic) Bayesian networks have
been successfully employed for some predefined contexts and
applications. Especially when considering context, as it is essential in e.g.
appraisal emotion models, or in specific applications, DBNs turn out
to be a promising approach. In contrast to most other fusion
mechanisms their structure may be actively modelled, based on existing
theories, so that the structure contains valuable information
implicitly, allowing to include existing knowledge in the model. This is
especially useful when it is required to make assumptions why the
model predicted one outcome and not another. It is worth mentioning
that context information has only rarely been taken into account - or
in most cases, limited to aspects like temporal context in previous
research. Yet, in human communication multiple aspects of context [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
continuously influence our behaviours.
2.4
      </p>
    </sec>
    <sec id="sec-6">
      <title>Explainable AI Approaches</title>
      <p>
        The current trend in machine learning tends towards deep learning
and neural network architectures that in contrast to Bayesian
networks aren’t inherently interpretable. Therefore efforts are made to
provide explanations for such ”black-box” approaches. In general
we can distinguish between two kinds of systems providing
explanations: model-agnostic or model-specific. Model-agnostic systems
are capable of generating explanations independent of the
underlying model. Ribeiro et al. introduce in [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ] LIME, a model-agnostic
approach for the generation of explanations. LIME is able to provide
explanations for any given model by approximating an interpretable
model around the passed model.
      </p>
      <p>
        Alber et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] introduced a library named iNNvestigate that
provides implementations of common analysis methods for neural
networks, e.g. PatternNet and LRP. The generated explanations come
in the form of highlighted regions, that have been important for the
classification. The supported methods are in contrast to Lime
modelspecific.
      </p>
      <p>
        Same goes for SHAP developed by Lundberg et al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Their
framework generates explanations by assigning each feature a value, that
describes its importance in regard to the prediction.
      </p>
      <p>Figure 1 displays what visual explanations generated by LIME and
iNNvestigate could possibly look like. The images have been
generated within the scope of the presented work. While such visual
explanation systems are of great value in helping to better understand
which part of the input data was relevant for a decision, they don’t
provide causal explanations. The explanation generated by LIME
highlights areas that are important for predicting a specific class in
green colour, whereas the red coloured shapes describe areas that
speak against the predicted class. In the example provided in Figure 1
it is evident that a large part of the face including the smile of the
person is important for classifying happiness. However the other half of
the face is coloured red and even some areas in the background are
coloured green. With this information alone it is not easily
comprehensible what the exact reasoning to predict a particular class has
been. The explanations generated with iNNvestigate are even harder
to correctly interpret. In the provided examples several edges
outlining the facial features of the subject are marked being relevant for
predicting. Those explanations often leave the user guessing and
applying self made causal coherencies to further explain the prediction.
Rather these approaches help to get better insight on the decisions of
a network on a feature level. A big advantage of Bayesian networks
is that the structure of a network can be modelled to have intrinsic
meaning. Those causal coherencies might be used as a foundation
for generating human-interpretable textual explanations.
3</p>
    </sec>
    <sec id="sec-7">
      <title>The Role of Context</title>
      <p>
        In current systems for recognising human behaviours only little
attention is given to context (e.g. context that is represented by
surrounding frames when training a model). Yet there are behaviours
that are difficult to analyse and interpret correctly without further
information about the context of a situation. Context is a wide-ranging
term that has different meanings depending on the paradigm of
research, application and scenario. Duranti et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] noted that it
seems impossible to present a single, precise and technical
definition of context. Context information might appear as a single
impact factor on the interaction or as a combination of multiple types
of information. In addition to that, various challenges occur when it
comes to context in multimodal communication [
        <xref ref-type="bibr" rid="ref50">50</xref>
        ]. In this section
we approach different aspects of context:
Temporal context: In classical linguistics, context is ”a frame that
surrounds the event and provides resources for its appropriate
interpretation” [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. W o¨llmer et al. [
        <xref ref-type="bibr" rid="ref54">54</xref>
        ] considered context as the
temporal surroundings of an observation. In their work they
successfully applied bidirectional long-short-term memory (BLSTM)
neural networks to consider contextual long-range observations
for the prediction of emotions. They further investigated
algorithms such as multidimensional dynamic time wrapping (DTW)
and asynchronous hidden-markov models to fuse mutual
information from multiple modalities, while considering their temporal
alignment [
        <xref ref-type="bibr" rid="ref53">53</xref>
        ]. An overview on algorithmic approaches, such as
dynamic and canonical time wrapping in the context of facial
expression analysis is given in [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ].
      </p>
      <p>
        When analysing complex social signals and emotions, the
temporal order of behaviours is of vast importance. As an example,
Keltner [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] describes a typical time series of behaviours in
multiple modalities, that represent a typical instance for the complex
emotion ”embarrassment” in a social situation - a similar times
series of events as we consider here for recognising engagement.
Typically, the gaze shifts towards the bottom, the lips make slight
Gaze
Lips
Face
Head
1
movements that often turn into a smile followed by the gaze and
head shifting to the side and back. Considering such sequences
of social signals adds valuable information to the interpretation,
compared to the analysis of isolated single cues.
      </p>
      <p>
        Interaction dynamics context Analysing the dynamics in human
communication includes being able to investigate both, the
individual multi-modal dynamics (see temporal context) as well as
the interpersonal dynamics. Researchers consider interpersonal
dynamics on multiple abstractions. For example, Delaherche et
al. and Varni et al. [
        <xref ref-type="bibr" rid="ref13 ref46">13, 46</xref>
        ] consider the synchronicity of people
in dyadic interactions on a signal level. Therefore, they
developed a set of synchronicity measurements. Rich et al. [
        <xref ref-type="bibr" rid="ref40">40</xref>
        ]
defined state machines to automatically recognise the four
interpersonal cues ”mutual gaze”, ”directed gaze”, ”adjacency pairs” and
”backchannels”. In their work they counted the appearance of such
bi-directional cues and considered their appearance as an
indicator of a person’s engagement. Another aspect is the current role in
a conversation. Depending on whether the user is in the role of a
listener or a speaker, the same kind of behaviour might be
interpreted in a completely different way. The influence of the
interaction role is illustrated by the following example. Let us assume
we observe a person showing a high amount of gestural activity.
If the person is in the role of a listener, the observed activity could
be interpreted as restlessness. On the opposite, if the person is in
the role of a speaker, we might conclude that the person is actively
engaged in the conversation. Salam et al. [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ] classify multiple
aspects of context as parts of the relationship of a social robot and a
human during an interaction. More precisely, the interaction
context in their definition describes how a scenario relates multiple
interlocutors.
      </p>
      <p>
        Semantic context: The interpretation of detected social cues can be
entirely altered through the semantics of accompanying verbal
utterances. For example, a laughter in combination with an utterance
commenting a negative event would no longer be interpreted as a
sign of happiness, but rather be taken as sarcasm. By considering
the semantics of accompanying spoken content, detected social
cues could be interpreted more accurately. Studies further indicate
that humans use semantic context for the interpretation of facial
expressions [
        <xref ref-type="bibr" rid="ref37 ref48 ref8">8, 37, 48</xref>
        ].
      </p>
      <p>
        Environmental context: The location and environmental
surroundings may also influence the way we behave during an interaction.
As an example, Zimmermann et al. [
        <xref ref-type="bibr" rid="ref57">57</xref>
        ] argues that the
environmental surroundings directly influence our behaviours e.g. in the
way we breathe or speak. In human-computer interaction and
especially in ubiquitous computing, a system is called context-aware
when it understands the circumstances and conditions surrounding
the user. Abowd et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], define context as ”any information that
can be used to characterise the situation of an entity. An entity is
a person, place, or object that is considered relevant to the
interaction between a user and an application, including the user and
applications themselves”. They further state that context is highly
dependable on the current perspective.
      </p>
      <p>
        Social context: Another aspect of context is the so called ”social
context”. Riek et al. [
        <xref ref-type="bibr" rid="ref41">41</xref>
        ] stress the importance of considering
social context when creating automated behaviour analysis
systems. In their definition, social context is the ”environment where
a particular person is situated with four factors that may influence
(their) behaviour: situational context, cultural context, the person’s
social role context, and the environmental social norms”. Such
aspects may be addressed by the following questions: In what kind
of situation does the conversation happen? What is the setting of
the interaction? (situational context), How well do the
interlocutors know each other? Do they share common knowledge? What
culture or gender do they have? What is their personality like?
(cultural context). How is their relationship? How is their social
status? (the person’s social role). What are the social norms in the
location of the interaction? What are the social norms in the
community of the interlocutors? (environmental social norms).
Questions like these play an important role, especially when
interpreting non-verbal behaviour. Some of these aspects might be difficult
to retrieve in an automated manner during the interaction between
multiple interlocutors. However, if it is not possible to
automatically gather such context information, it could be collected
upfront.
      </p>
      <p>
        When humans interpret behaviours of other people, they
consciously or unconsciously include these and similar considerations
in their reasoning process. Machines that aim to correctly interpret
human behaviours should therefore consider contextual aspects in
their interpretation models as well. Yet, besides temporal context
(e.g. [
        <xref ref-type="bibr" rid="ref54">54</xref>
        ]), only little attention has been put to contextual aspects
in current social signal processing research.
4
      </p>
    </sec>
    <sec id="sec-8">
      <title>NoXi Database</title>
      <p>
        The data for the upcoming evaluation tasks has been gathered from
the NoXi Database [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. NoXi provides dyadic novice-expert
conversations. One participant took the role of the expert and the other
one the role of the novice. Experts were free to chose the topic they
wanted to talk about. Furthermore, the novices were evenly free in
choosing what to listen to. This resulted in conversations covering
a broad scope of different topics ranging from photography to
dementia. Both participants were placed in separate rooms during the
recording. They interacted remotely through TV screens and
microphones. An example for the setup can be seen in Figure 3.
      </p>
      <p>The database covers multiple languages and ethnicities, e.g.
English, French, German, Indonesian, Arabic, Spanish, Italian.
However, English, German and French have been the languages that
occurred the most. A total of 84 sessions have been recorded, providing
25 hours and 18 minutes of conversational data. Additionally,
demographic information of the participants have been collected, which
include gender, cultural identity, age and level of education. The range
of age has been from 21 to 50 years. We decided for the NoXi corpus
due to the fact that it contains multi-modal multi-person interaction
data and its transferability to social coaching scenarios. Moreover
the setup of the corpus allowed for both, engaging, as well as
nonengaging interactions.</p>
      <p>A total of 19 sessions of the NoXi corpus have been annotated
regarding conversational engagement. The annotators followed the
engagement definition of Poggi, which we introduced in subsection 2.1. For
most of the sessions novice and expert annotations have been created.
Of the 19 sessions twelve are associated with French, four with
English and three with German. For annotating, a continuous scheme
has been chosen. The engagement annotations were created on the
ratings of 4-7 different annotators. To measure the quality of the
created annotations, from every annotator, they are validated against
each other using the Pearson Correlation Coefficient (PCC). Based
on the PCC, a gold standard for the annotations has been created.
Whenever different annotators have scored a PCC value greater than
0.5 they have been merged to a gold standard annotation.
Depending on the definition a value greater than 0.5 is considered a strong
uphill (positive) linear relationship. However, at least two annotators
have to score higher than 0.5, otherwise no gold standard has been
created for the specific session and the session has been discarded.
The gold standard itself is calculated by averaging the corresponding
annotations.</p>
      <p>Figure 4 displays examples for very low, medium and very high
engagement, with the corresponding gold standard annotation. The
first image has been interpreted by the annotators as very low
engagement. This scene occurred, as the novice decided to answer his
phone, during the conversation (there were planned interruptions in
the NoXi corpus, e.g. by calls from the experimenters or walk-ins).
Answering the phone can be considered as a strong signal of the
individual not willing to maintain the interaction. The alignment of head
and body, away from the interlocutor, go along with a very low level
of engagement. The next picture displays a neutral body position of
the novice. This behaviour has been associated with a medium level
of engagement. He aligned his body towards the other participant
and is focusing the TV-screen. The last image represents very high
engagement. The novice is smiling and shows a very open body
posture, with the arms wide spread using a large gesture space. Again
his body is aligned towards his interlocutor.</p>
      <p>
        Engagement comes in various facets and sometimes the
determination of its degree is distinct, like the just presented examples for very
low and very high engagement. However, sometimes things are less
obvious and leave room for a different interpretation. During the
continuous annotation of conversational engagement we faced similar
problems, as the ones mentioned by Whitehill et al. in [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ]. They
faced the issue, that an annotator tends to classify the level of
engagement in the context of the currently annotated individual.
Furthermore, they argue this could lead to annotations that are not
comparable between different sessions. In fact, during the process of
annotating, we often caught ourselves with statements like, “For their type
of character, this should be considered as low/medium/high
engagement”. However, we figured out that this causal chain is not wrong.
It shows, that the way the level of engagement of an individual is
perceived, also depends on the psychological traits the annotator
attributes to the individual. Those traits can be considered as context
information, which could be modelled inside the Bayesian network.
5
      </p>
    </sec>
    <sec id="sec-9">
      <title>Engagement Model</title>
      <p>
        Based on the evidences presented in subsection 2.1 we developed an
annotation scheme that has been used to train our Bayesian networks.
We considered different modalities besides context information.
Audio: First of all we considered the general voice activity of the
interlocutors as valuable information. Even though it is very basic in
its nature it allows to draw a conclusion about the overall
involvement of the individuals regarding the conversation. An overall low
voice activity may imply a conversation with low engaged
interlocutors. On top of that we distinguished between different types
of voice activity. We considered speech, filler and silence. The
fillers are a particularly interesting type of voice activity as they
also cover audio backchannels. In subsection 2.1 we mentioned
that backchannels are a very common tool during conversation
and provide information about the level of engagement [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
Further Knapp et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] argue that emotions are reliably
transported by the voice. Therefore we trained a support vector
machine (SVM) to predict the arousal of the voice [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The output of
both SVM models (arousal, speech/filler/silence) is used to train
the Bayesian network.
      </p>
      <p>
        Face/Head: During conversations the face usually occupies most of
the interlocutors attention. A lot of important information
regarding the level of engagement can be extracted from the face
respectively the head. Therefore we aimed in our annotation scheme to
cover a general impression of the region, as well as looking for
specific behaviour that is strongly connected to engagement. We
defined features that represent the overall movement of the head
in regard to X,Y and Z-Axis. Those features were mainly inspired
by the research of Ryota Ooko et al. [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. They found that a
moderate positive correlation of head movement regarding the level
of conversational engagement is present. Further we considered
the individual gaze behaviour of the participants. There are
multiple studies present about the recognition of engagement solely
based on gaze data, with good recognition scores [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Finally
we trained a neural network on the facial action units (FACS)
extracted with Openface [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to predict valence [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. We used the
output of the neural network to train our Bayesian network.
Body: We mentioned earlier in subsection 2.1 that the alignment
and movement of the body play an important role in the
recognition of engagement. We followed an approach that has been
similar to the head features. We tried to cover the general behaviour of
the body, as well as specific gestures or poses that are connected
to engagement. Therefore we defined a group of features, called
body properties. They are mainly inspired by the coding system
introduced in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. It contains values for the distance between the
arms and the hips for X and Z-Axis. Moreover, the alignment of
the arms is covered, by calculating the rotation of the elbow joints.
Those values are supposed to describe a general level of openness.
Also the distances of each arm to the hip allow interpretation of
the symmetry of the arms. In addition to that, the standard
deviation of the distance travelled by the head during a frame and the
rotation of the head is calculated. Those values have been chosen
based on [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        In subsection 2.1 we identified restlessness to be connected to
low levels of engagement. This is the reason we decided to
calculate the continuous movement of the interlocutors. Continuous
movement is a cumulative value, which describes the overall body
movement. Lots of movement may indicate restlessness. In
addition to that we wanted to cover the amount of gesticulation an
individual performs. Gesticulation is mentioned in [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] and [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] as a
crucial nonverbal queue in communication. Therefore we mapped
the amount of movement done by both hands onto a real number
value, which represents a numeric value for gesticulation.
Furthermore we considered the crossed arms and head touch
gestures. The crossing of the arms is a common and often observed
gesture. In research it is often interpreted as the expression of a
negative emotional attitude by individuals [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] [
        <xref ref-type="bibr" rid="ref49">49</xref>
        ]. Based on this
we argue that a negative emotional state is bonded to low
engagement. In subsection 2.1 we mentioned self touches as a possible
signal of being emotionally engaged. Moreover, Gunes et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
were able to achieve good recognition rates for emotions, based
on face and body features. Their system associated the emotions
of fear, sadness and surprise mostly with gestures of the hands
touching the head.
      </p>
      <p>We believe that context plays an important role when it comes to
correctly identifying social behaviour. The same applies to
recognising conversational engagement. Depending on the context a specific
gesture or behaviour may have a different meaning. Recall the
example of the very actively moving engaged expert. His continuous
movement is not a sign of restlessness. Given the fact that he is
talking and gesticulating he should be considered as actively engaged in
the conversation. Based on the different types of context we defined
in section 3 we considered following context to predict
conversational engagement.</p>
      <p>
        Turn hold: During a conversation the interlocutors usually alternate
their speaking turns. Therefore we determine the interlocutor that
is currently holding the turn. Turn taking and vocal cues play an
important part during conversations [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. This kind of information
can be considered as interaction dynamics context.
      </p>
      <p>Role: In the used corpus two roles have been present: novice and
expert. The novice has been the one with little to no knowledge about
the topic presented by the expert. Accordingly, the expert has been
the one introducing and providing information about the topic to
the novice. Furthermore, it is in the nature of the expert to be more
talkative than the novice, therefore a rather silent expert tends to
be in a state of lower engagement, when compared to a similar
silent novice, who might be just interestedly listening. In terms
of context the information about the role covers multiple aspects.
As we just elaborated, most of the time novices and experts
operate differently during conversations. Therefore this can be seen
as interaction dynamics context. Besides that, the role also covers
social context. This is due to the fact, that specific expectations are
raised towards the expert. By putting themselves in the role of an
expert they signal the novice that they have sophisticated
knowledge about their topic. This may result in novices being rather
reserved regarding their interactions and comments. Moreover, it
is common for the expert to take the lead during the conversation,
which automatically results in more speaking time.</p>
      <p>
        Gender: There are differences in the behaviour during
conversations depending on the gender of the interlocutors [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. For
example, in same-gender conversation pairs females tend to have more
eye contact with each other then males do. Also, males are more
prone to decrease eye contact over time, while females have a
tendency to increase it [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. That is only one of many examples where
the different genders behave differently. Due to that we think that
not only gender itself, but also the constellation of interlocutor
pairs, e.g. male-male, male-female, female-female, will be
beneficial to the recognition of engagement. By considering the gender
we aim to cover another aspect of social context.
      </p>
      <p>Temporal context: In section 3 we argued that the temporal order
of behaviours is important when it comes to analysing complex
social signals, such as engagement. That means, time series and
patterns of behaviours have different meaning when performed
differently.</p>
      <p>Coming up with a suitable architecture for the Bayesian network
has been an incremental approach. This process included
systematically adding, removing and exchanging classifiers, because even
though specific characteristics for engagement are suggested in the
literature, it does not necessarily mean they will work for any given
context.</p>
      <p>To provide more insight about the actual architecture Figure 5
displays an excerpt of the multi person dynamic Bayesian network.
Basically the network is a graphical representation of the just presented
annotation scheme. However, a big advantage of Bayesian networks
is that the structure has intrinsic meaning compared to other
models (e.g. artificial neural networks). This way, we were able to take
knowledge about causal coherencies into account. Context nodes
such as the gender or role are represented by conditional nodes, so
that engagement is predicted ”given” the context information, while
social cues are ”symptoms” shown by the observed person. In other
words, social cues can be observed, given that a person has a certain
level of engagement. Most of the context information we considered
important is focused on a single interlocutor. However we also
identified interaction dynamics context as a key element in correctly
interpreting conversational engagement. Therefore we chose to model
a multi person Bayesian network that also takes the interaction
context and the interaction dynamics of the different interlocutors into
account when estimating conversational engagement. For the NoXi
Database this resulted in a network considering two persons -
expert and novice. Moreover, we modelled our network as a dynamic
Bayesian network. This way we were able to take temporal context
into account.
6</p>
    </sec>
    <sec id="sec-10">
      <title>Transparency</title>
      <p>
        Bayesian networks not only allow us to easily model context and
other causal coherencies, but also provide transparency by default
[
        <xref ref-type="bibr" rid="ref52">52</xref>
        ]. In subsection 2.4 we mentioned that machine learning models,
in the context of explainable AI, can be distinguished between
inherently interpretable models and black-box models. Bayesian networks
are inherently interpretable. This is due to the fact that for a given set
of variables a Bayesian network is a representation of the joint
probability distribution [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. Usually we want a trained Bayesian network
- given a set of observation - to predict what the most likely class of
our target node is. In our use case we want to know how engaged one
of the interlocutors is. However in a Bayesian network we are not
only able to find out how engaged a person is but also what are the
most important features for a specific class and what characteristics
      </p>
      <sec id="sec-10-1">
        <title>Context</title>
        <p>Interaction
context
Interaction
Context
[Topic,
Unexpected
Events]</p>
        <p>IC
Interaction
Dynamics
[Back
channels,
Interruptions,..]
E_S
E_VA
E_ID
N_ID
N_VA
N_S
N_HE
Gender
[Male,
Female]
N_G
N_EN</p>
      </sec>
      <sec id="sec-10-2">
        <title>Novice</title>
        <p>Social
context</p>
        <sec id="sec-10-2-1">
          <title>Engagement</title>
        </sec>
      </sec>
      <sec id="sec-10-3">
        <title>Expert</title>
        <sec id="sec-10-3-1">
          <title>Engagement</title>
          <p>does the feature have. In Figure 6 a schematic of a reduced Bayesian
network for the recognition of engagement is presented. The network
contains the features Hand Energy and Voice Activity, which can take
the characteristics low, medium and high. Moreover we have our
target node Engagement, which also can be low, medium and high.
Finally we considered some social context by adding the Role of the
interlocutors. The schematic displays the probability distribution of
the nodes given the person is highly engaged. This information tells
us that when a person is highly engaged they are most likely in the
role of the expert (70%) and show most likely high levels of Hand
Energy and Voice Activity. We could now apply the same approach
to find out more about low and medium engagement and get
extensive insight about the learnt representations of our network.
7</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>Evaluation</title>
      <p>Even though transparency is important in the context of machine
learning, there is little use for a transparent model that isn’t able to
accurately predict the task at hand. That is why we investigate in the
following the performance of the introduced architectures compared
to other state-of-the-art machine learning approaches.</p>
      <p>We split the acquired data into dedicated sets for training and
evaluation. The training set included 13 sessions and had a size of 616374
samples. The evaluation set consisted out of six sessions, with a total
of 328385 samples. So we ended up with the evaluation set having
roughly half the samples of the training set.</p>
      <p>To evaluate the different models, the Pearson correlation coefficient
has been calculated between the model’s prediction and the gold
standard annotation.</p>
      <p>[ht]</p>
      <p>As described earlier, developing a suitable Bayesian network has
been an incremental approach by adjusting the classifier
composition. An early Bayesian network (BN) based on multiple modalities
including some context information achieved promising results with
a PCC of 0.7373. By extending this network with temporal context
for selected nodes that are related to body and face movement as
well as voice activity we were able to further improve the
correlation score to 0.7443. During our tests the network that performed
best has been a multi-person dynamic Bayesian network (MDBN). It
incorporates interpersonal dynamics, like mutual gaze and turn
transitions between the novice and expert. The network achieved a PCC
of 0.768 which is significantly better (p&lt;0.001) than the best
singleuser DBN (0.7443).</p>
      <p>The (D)BNs we applied are created using a hybrid approach where
classification results for sub-recognition tasks, as well as threshold
based features are used to update the evidences in the network. This
makes it difficult to compare the multi-modal model with other
classification models that rely on low level features. In order to have a
baseline to evaluate our approach, we created an engagement
feature set that is heavily influenced by the previously introduced
engagement annotation scheme. It contains features on body
movement, body posture, head movement, facial expression and audio. We
trained a linear support vector machine (LSVM) on this feature set
and achieved a PCC of 0.6253. Moreover, we tested several neural
networks implemented in Keras. The best one has been a fully
connected deep recurrent neural network (RNN) and was able to score
a PCC of 0.6034 on the engagement feature set. Those results are
significantly (p&lt;0.001) worse than our introduced hybrid model.
8</p>
    </sec>
    <sec id="sec-12">
      <title>Discussion</title>
      <p>
        We were able to show that our hybrid approach using a
theorymodelled DBN can deliver comparable results to purely statistical
black-box approaches. This is in compliance with the research of
Rudin [
        <xref ref-type="bibr" rid="ref42">42</xref>
        ]. On our corpus it even slightly outperformed the other
classification methods. With the introduction of a multi person
dynamic Bayesian network architecture we were able to further
increase the prediction accuracy. We explain this with several aspects:
by employing the transparent DBN we could intuitively refine our
first assumptions on what influences engagement, which allowed us
to incrementally add classifiers, until the network achieved satisfying
correlations with our gold standard annotation. Further, through the
update mechanism on annotation/event abstraction we aimed to
simulate a decision making and reasoning process that’s similar to the
one of humans. To our understanding, humans will consciously or
unconsciously map abstractions of behaviours (e.g. smiles) on their
perception of the other person (e.g. happiness). Further, we conclude
that for our particular use-case of recognising conversational
engagement, considering different types of context information leads to
improvements in terms of the correct and adequate interpretation. In
fact the more context information we added the better our model
performed.
9
      </p>
    </sec>
    <sec id="sec-13">
      <title>Conclusion</title>
      <p>
        Deep learning can be considered as the current gold standard in
machine learning. Deep neural networks proved themselves on
various problem domains by performing exceptionally well [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However their biggest weakness is their lack of interpretability.
That is why efforts are made to provide additional insight to
otherwise ”black-boxes” (see subsection 2.4). Even though there are
approaches present that help in gaining additional insight on the
decision-making of neural network architectures, they rather
provide additional information on a feature-level basis. In contrast to
that there are models, like Bayesian networks that are inherently
interpreteable and can be modelled to have intrinsic meaning. This
enables a user to gather causal coherencies on why a model made
a specific prediction. Often this seems to come down to a
tradeoff between prediction performance and transparency. However, we
showed for the use case of multi-modal engagement recognition that
by applying a hybrid approach that fuses abstractions of multiple
social cues in a causal recognition model, accuracy and transparency do
not necessarily need to exclude each other. Moreover we were able to
improve the recognition rates of our model by incorporating social,
temporal and interaction dynamics context. The significant impact
of context on recognition scores stresses the importance of context
in correctly and adequately interpreting conversational engagement.
The proposed system has been implemented within the SSI
Framework [
        <xref ref-type="bibr" rid="ref47">47</xref>
        ], so that all social cue classification models, as well as the
overall BN inference step can be performed in a real-time system.
This allows to apply this approach in a variety of applications, such
as human-agent or human-robot scenarios.
      </p>
    </sec>
    <sec id="sec-14">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work has received funding from the DFG under project number
392401413, DEEP.</p>
      <p>Further this work presents and discusses results in the context of the
research project ForDigitHealth. The project is part of the Bavarian
Research Association on Healthy Use of Digital Technologies and
Media (ForDigitHealth), funded by the Bavarian Ministry of Science
and Arts.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Gregory</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Abowd</surname>
          </string-name>
          ,
          <string-name>
            <surname>Anind K. Dey</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter J. Brown</surname>
            , Nigel Davies,
            <given-names>Mark</given-names>
          </string-name>
          <string-name>
            <surname>Smith</surname>
          </string-name>
          , and Pete Steggles, '
          <article-title>Towards a better understanding of context and context-awareness'</article-title>
          ,
          <source>in Handheld and Ubiquitous Computing</source>
          , First International Symposium, HUC'99,
          <string-name>
            <surname>Karlsruhe</surname>
          </string-name>
          , Germany,
          <source>September 27-29</source>
          ,
          <year>1999</year>
          , Proceedings, ed.,
          <string-name>
            <surname>Hans-Werner</surname>
            <given-names>Gellersen</given-names>
          </string-name>
          , volume
          <volume>1707</volume>
          of Lecture Notes in Computer Science, pp.
          <fpage>304</fpage>
          -
          <lpage>307</lpage>
          . Springer, (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Igor</given-names>
            <surname>Aizenberg</surname>
          </string-name>
          and Gonzalez Alexander, '
          <article-title>Image recognition using mlmvn and frequency domain features'</article-title>
          ,
          <source>International Joint Conference on Neural Networks</source>
          , (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Maximilian</given-names>
            <surname>Alber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Lapuschkin</surname>
          </string-name>
          , Philipp Seegerer, Miriam Ha¨gele, Kristof T. Schu¨tt, Gre´goire Montavon, Wojciech Samek, KlausRobert Mu¨ller, Sven Da¨hne, and
          <string-name>
            <surname>Pieter-Jan</surname>
            <given-names>Kindermans</given-names>
          </string-name>
          , '
          <article-title>innvestigate neural networks!'</article-title>
          , CoRR, abs/
          <year>1808</year>
          .04260, (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>T. Baltrusˇaitis</surname>
          </string-name>
          , P. Robinson, and
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Morency</surname>
          </string-name>
          , '
          <article-title>Openface: An open source facial behavior analysis toolkit'</article-title>
          ,
          <source>in 2016 IEEE Winter Conference on Applications of Computer Vision (WACV)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          , (
          <year>March 2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Basu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bag</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mahadevappa</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>J. Mukherjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Guha</surname>
          </string-name>
          , '
          <article-title>Emotion recognition based on physiological signals using valence-arousal model'</article-title>
          ,
          <source>in 2015 Third International Conference on Image Information Processing (ICIIP)</source>
          , pp.
          <fpage>50</fpage>
          -
          <lpage>55</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tobias</given-names>
            <surname>Baur</surname>
          </string-name>
          , Dominik Schiller, and Elisabeth Andre´, '
          <article-title>Modeling user's social attitude in a conversational system'</article-title>
          ,
          <source>in Emotions and Personality in Personalized Services</source>
          ,
          <fpage>181</fpage>
          -
          <lpage>199</lpage>
          , Springer, (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Roman</given-names>
            <surname>Bednarik</surname>
          </string-name>
          , Shahram Eivazi, and Michal Hradis, '
          <article-title>Gaze and conversational engagement in multiparty video conversation: An annotation scheme and classification of high and low levels of engagement'</article-title>
          ,
          <source>in Proceedings of the 4th Workshop on Eye Gaze in Intelligent Human Machine Interaction, Gaze-In '12</source>
          , pp.
          <volume>10</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          :
          <fpage>6</fpage>
          , New York, NY, USA, (
          <year>2012</year>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Vicki</given-names>
            <surname>Bruce</surname>
          </string-name>
          and
          <article-title>Andy Young, In the eye of the beholder: the science of face perception</article-title>
          ., Oxford University Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Cafaro</surname>
          </string-name>
          , Johannes Wagner, Tobias Baur, Soumia Dermouche, Mercedes Torres Torres, Catherine Pelachaud, Elisabeth Andre´, and Michel Valstar, '
          <article-title>The noxi database: Multimodal recordings of mediated novice-expert interactions'</article-title>
          , ICMI'
          <fpage>17</fpage>
          , (
          <year>November 2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Dan</given-names>
            <surname>Cires</surname>
          </string-name>
          <article-title>¸an, Ueli Meier, and J u¨rgen Schmidhuber, 'Multi-column deep neural networks for image classification'</article-title>
          ,
          <source>arXiv preprint arXiv:1202.2745</source>
          , (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Cristina</given-names>
            <surname>Conati</surname>
          </string-name>
          and Heather Maclaren, '
          <article-title>Modeling user affect from causes and effects', in User Modeling, Adaptation, and</article-title>
          <string-name>
            <surname>Personalization</surname>
            , 17th International Conference,
            <given-names>UMAP</given-names>
          </string-name>
          <year>2009</year>
          ,
          <article-title>formerly UM and AH</article-title>
          , Trento, Italy, June 22-26,
          <year>2009</year>
          . Proceedings, pp.
          <fpage>4</fpage>
          -
          <lpage>15</lpage>
          , (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Nele</surname>
            <given-names>Dael</given-names>
          </string-name>
          , Marcello Mortillaro, and
          <string-name>
            <surname>Klaus R. Scherer</surname>
          </string-name>
          , '
          <article-title>The body action and posture coding system (bap): Development and reliability'</article-title>
          ,
          <source>Journal of Nonverbal Behavior</source>
          ,
          <volume>36</volume>
          (
          <issue>2</issue>
          ),
          <fpage>97</fpage>
          -
          <lpage>121</lpage>
          , (
          <year>Jun 2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Emilie</surname>
            <given-names>Delaherche</given-names>
          </string-name>
          , Mohamed Chetouani, Ammar Mahdhaoui, Catherine Saint-Georges,
          <string-name>
            <given-names>Sylvie</given-names>
            <surname>Viaux</surname>
          </string-name>
          , and David Cohen, '
          <article-title>Interpersonal synchrony: A survey of evaluation methods across disciplines'</article-title>
          ,
          <source>IEEE Trans. Affective Computing</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ),
          <fpage>349</fpage>
          -
          <lpage>365</lpage>
          , (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Sidney S D'Mello</surname>
            ,
            <given-names>Patrick</given-names>
          </string-name>
          <string-name>
            <surname>Chipman</surname>
          </string-name>
          , and Art Graesser, '
          <article-title>Posture as a predictor of learner's affective engagement'</article-title>
          ,
          <source>in Proceedings of the Annual Meeting of the Cognitive Science Society</source>
          , volume
          <volume>29</volume>
          , (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Duranti</surname>
          </string-name>
          and
          <string-name>
            <given-names>Charles</given-names>
            <surname>Goodwin</surname>
          </string-name>
          ,
          <article-title>Rethinking context: Language as an interactive phenomenon, number 11 in Studies in the Social and Cultural Foundations of Language</article-title>
          , Cambridge University Press,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Stephen R Garner</surname>
          </string-name>
          et al., 'Weka:
          <article-title>The waikato environment for knowledge analysis'</article-title>
          ,
          <source>in Proceedings of the New Zealand computer science research students conference</source>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>64</lpage>
          , (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.</given-names>
            <surname>Glas</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Pelachaud</surname>
          </string-name>
          , '
          <article-title>Definitions of engagement in human-agent interaction'</article-title>
          ,
          <source>in 2015 International Conference on Affective Computing and Intelligent Interaction (ACII)</source>
          , pp.
          <fpage>944</fpage>
          -
          <lpage>949</lpage>
          , (
          <year>Sept 2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Hatice</given-names>
            <surname>Gunes</surname>
          </string-name>
          and Massimo Piccardi, '
          <article-title>Affect recognition from face and body: early fusion vs. late fusion'</article-title>
          ,
          <source>in Systems, Man and Cybernetics</source>
          , 2005 IEEE International Conference on, volume
          <volume>4</volume>
          , pp.
          <fpage>3437</fpage>
          -
          <lpage>3443</lpage>
          . IEEE, (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <article-title>Ju¨rgen Schmidhuber, 'Long short-term memory'</article-title>
          ,
          <source>Neural Computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ),
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          , (
          <year>1997</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Ryo</given-names>
            <surname>Ishii</surname>
          </string-name>
          and
          <string-name>
            <surname>Yukiko I. Nakano</surname>
          </string-name>
          , '
          <article-title>An empirical study of eye-gaze behaviors: Towards the estimation of conversational engagement in humanagent communication'</article-title>
          ,
          <source>in Proceedings of the 2010 Workshop on Eye Gaze in Intelligent Human Machine Interaction, EGIHMI '10</source>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          , New York, NY, USA, (
          <year>2010</year>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Dacher</surname>
            <given-names>Keltner</given-names>
          </string-name>
          , '
          <article-title>Signs of appeasement: Evidence for the distinct displays of embarrassment, amusement, and shame'</article-title>
          ,
          <source>Journal of personality and social psychology</source>
          ,
          <volume>68</volume>
          (
          <issue>3</issue>
          ),
          <fpage>441</fpage>
          , (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Mardi</surname>
            <given-names>Kidwell</given-names>
          </string-name>
          , 'Framing, grounding, and
          <article-title>coordinating conversational interaction: Posture, gaze, facial expression, and movement in space'</article-title>
          ,
          <source>in Body - Language - Communication. An International Handbook on Multimodality in Human Interaction</source>
          ,
          <fpage>100</fpage>
          -
          <lpage>113</lpage>
          , De Gruyter Mouton, (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Knapp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            and
            <surname>Judith Hall</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          ,
          <source>Nonverbal Communication in Human Interaction, Harcourt Brace</source>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Alex</surname>
            <given-names>Krizhevsky</given-names>
          </string-name>
          , Ilya Sutskever, and Geoffrey E Hinton, '
          <article-title>Imagenet classification with deep convolutional neural networks'</article-title>
          ,
          <source>in Advances in Neural Information Processing Systems</source>
          <volume>25</volume>
          , eds.,
          <string-name>
            <given-names>F.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J. C.</given-names>
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bottou</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          ,
          <volume>1097</volume>
          -
          <fpage>1105</fpage>
          , Curran Associates, Inc., (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Hedda</surname>
            <given-names>Lausberg</given-names>
          </string-name>
          , '
          <article-title>Neuropsychology of gesture production'</article-title>
          ,
          <source>in Body - Language - Communication. An International Handbook on Multimodality in Human Interaction</source>
          ,
          <fpage>168</fpage>
          -
          <lpage>182</lpage>
          , De Gruyter Mouton, (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Birgit</surname>
            <given-names>Lugrin</given-names>
          </string-name>
          , Julian Frommel, and Elisabeth Andre´, '
          <article-title>Combining a data-driven and a theory-based approach to generate culture-dependent behaviours for virtual characters'</article-title>
          ,
          <source>in Advances in Culturally-Aware Intelligent Systems and in Cross-Cultural Psychological Studies</source>
          ,
          <fpage>111</fpage>
          -
          <lpage>142</lpage>
          , Springer, (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Scott</surname>
            <given-names>M Lundberg</given-names>
          </string-name>
          and
          <string-name>
            <surname>Su-In</surname>
            <given-names>Lee</given-names>
          </string-name>
          , '
          <article-title>A unified approach to interpreting model predictions'</article-title>
          ,
          <source>in Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , eds., I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Garnett</surname>
          </string-name>
          ,
          <volume>4765</volume>
          -
          <fpage>4774</fpage>
          , Curran Associates, Inc., (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Marwa</given-names>
            <surname>Mahmoud</surname>
          </string-name>
          and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Robinson</surname>
          </string-name>
          , '
          <article-title>Interpreting hand-over-face gestures'</article-title>
          ,
          <source>in Affective Computing and Intelligent Interaction</source>
          ,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          , Springer, (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Albert</surname>
            <given-names>Mehrabian</given-names>
          </string-name>
          , Nonverbal Communication, AldineTransaction,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Mitchell</surname>
          </string-name>
          , Tom,
          <source>Machine Learning</source>
          ,
          <fpage>177</fpage>
          -
          <lpage>197</lpage>
          , MacGraw-Hill,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31] Cornelia Mu¨ller, Alan Cienki, Ellen Fricke, Silva Ladewig,
          <string-name>
            <surname>David McNeill</surname>
            ,
            <given-names>and Sedinha</given-names>
          </string-name>
          <string-name>
            <surname>Tessendorf</surname>
          </string-name>
          ,
          <string-name>
            <surname>Body - Language - Communication</surname>
          </string-name>
          . An International Handbook on Multimodality in Human Interaction, De Gruyter Mouton,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Patrick</surname>
          </string-name>
          Murphy and Stuart Russell, '
          <article-title>Dynamic bayesian networks: representation, inference and learning'</article-title>
          ,
          <source>Ph.D Thesis</source>
          , (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Ryota</surname>
            <given-names>Ooko</given-names>
          </string-name>
          , Ryo Ishii, and
          <string-name>
            <surname>Yukiko</surname>
            <given-names>I. Nakano</given-names>
          </string-name>
          , '
          <article-title>Estimating a user's conversational engagement based on head pose information'</article-title>
          , in Intelligent Virtual Agents, eds., Hannes Ho¨gni Vilhja´lmsson, Stefan Kopp, Stacy Marsella, and Kristinn R. Tho´risson, pp.
          <fpage>262</fpage>
          -
          <lpage>268</lpage>
          , Berlin, Heidelberg, (
          <year>2011</year>
          ). Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Ortony</surname>
          </string-name>
          , Gerald L Clore, and Allan Collins,
          <article-title>The cognitive structure of emotions</article-title>
          , Cambridge university press,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Yiannis</surname>
            <given-names>Panagakis</given-names>
          </string-name>
          , Ognjen Rudovic, and Maja Pantic, '
          <article-title>Learning for multi-modal and context-sensitive interfaces', The Handbook of Multimodal-Multisensor Interfaces</article-title>
          , Volume
          <volume>2</volume>
          :
          <string-name>
            <given-names>Signal</given-names>
            <surname>Processing</surname>
          </string-name>
          , Architectures, and
          <article-title>Detection of Emotion and Cognition, 2</article-title>
          , in press, (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Isabella</surname>
            <given-names>Poggi</given-names>
          </string-name>
          ,
          <article-title>Mind, hands, face and body: a goal and belief view of multimodal communication</article-title>
          ,
          <source>Weidler</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Carl</surname>
            <given-names>Ratner</given-names>
          </string-name>
          , '
          <article-title>Back to dr. ratner's home page journal of mind and behavior,</article-title>
          <year>1989</year>
          ,
          <volume>10</volume>
          ,
          <fpage>211</fpage>
          -
          <lpage>230</lpage>
          a
          <article-title>social constructionist critique of naturalistic theories of emotion'</article-title>
          ,
          <source>Journal of Mind and Behavior</source>
          ,
          <volume>10</volume>
          ,
          <fpage>211</fpage>
          -
          <lpage>230</lpage>
          , (
          <year>1989</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Tulio</surname>
          </string-name>
          <string-name>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          , '
          <article-title>”why should I trust you?”: Explaining the predictions of any classifier'</article-title>
          ,
          <source>in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          , San Francisco, CA, USA,
          <year>August</year>
          13-
          <issue>17</issue>
          ,
          <year>2016</year>
          , pp.
          <fpage>1135</fpage>
          -
          <lpage>1144</lpage>
          , (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ponsler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Holroyd</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. L.</given-names>
            <surname>Sidner</surname>
          </string-name>
          , '
          <article-title>Recognizing engagement in human-robot interaction'</article-title>
          ,
          <source>in 2010 5th ACM/IEEE International Conference on Human-Robot Interaction (HRI)</source>
          , pp.
          <fpage>375</fpage>
          -
          <lpage>382</lpage>
          , (
          <year>March 2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>Charles</surname>
            <given-names>Rich</given-names>
          </string-name>
          , Brett Ponsleur, Aaron Holroyd, and Candace L. Sidner, '
          <article-title>Recognizing engagement in human-robot interaction'</article-title>
          ,
          <source>in Proceedings of the 5th ACM/IEEE International Conference on Human Robot Interaction, HRI</source>
          <year>2010</year>
          , Osaka, Japan, March 2-
          <issue>5</issue>
          ,
          <year>2010</year>
          , eds.,
          <string-name>
            <surname>Pamela</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hinds</surname>
          </string-name>
          , Hiroshi Ishiguro, Takayuki Kanda, and
          <string-name>
            <surname>Peter H. Kahn</surname>
          </string-name>
          Jr., pp.
          <fpage>375</fpage>
          -
          <lpage>382</lpage>
          . ACM, (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <surname>Laurel</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Riek</surname>
            and
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Robinson</surname>
          </string-name>
          , '
          <article-title>Challenges and opportunities in building socially intelligent machines [social sciences]', IEEE Signal Process</article-title>
          . Mag.,
          <volume>28</volume>
          (
          <issue>3</issue>
          ),
          <fpage>146</fpage>
          -
          <lpage>149</lpage>
          , (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>Cynthia</surname>
            <given-names>Rudin</given-names>
          </string-name>
          , '
          <article-title>Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead'</article-title>
          ,
          <source>Nature Machine Intelligence</source>
          ,
          <volume>1</volume>
          (
          <issue>5</issue>
          ),
          <fpage>206</fpage>
          -
          <lpage>215</lpage>
          , (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <surname>Jennifer</surname>
            <given-names>Sabourin</given-names>
          </string-name>
          , Bradford W. Mott, and James C. Lester, '
          <article-title>Modeling learner affect with theoretically grounded dynamic bayesian networks'</article-title>
          ,
          <source>in Affective Computing and Intelligent Interaction - 4th International Conference, ACII</source>
          <year>2011</year>
          ,
          <article-title>Memphis</article-title>
          ,
          <string-name>
            <surname>TN</surname>
          </string-name>
          , USA, October 9-
          <issue>12</issue>
          ,
          <year>2011</year>
          , Proceedings,
          <string-name>
            <surname>Part</surname>
            <given-names>I</given-names>
          </string-name>
          , pp.
          <fpage>286</fpage>
          -
          <lpage>295</lpage>
          , (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>Hanan</given-names>
            <surname>Salam</surname>
          </string-name>
          and Mohamed Chetouani, '
          <article-title>A multi-level context-based modeling of engagement in human-robot interaction'</article-title>
          ,
          <source>in 11th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition</source>
          ,
          <string-name>
            <surname>FG</surname>
          </string-name>
          <year>2015</year>
          , Ljubljana, Slovenia, May 4-
          <issue>8</issue>
          ,
          <year>2015</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . IEEE Computer Society, (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <surname>Jyotirmay</surname>
            <given-names>Sanghvi</given-names>
          </string-name>
          , Ginevra Castellano, Iolanda Leite, Andre´ Pereira,
          <string-name>
            <surname>Peter W. McOwan</surname>
          </string-name>
          , and Ana Paiva, '
          <article-title>Automatic analysis of affective postures and body motion to detect engagement with a game companion'</article-title>
          ,
          <source>in Proceedings of the 6th International Conference on Human-robot Interaction, HRI '11</source>
          , pp.
          <fpage>305</fpage>
          -
          <lpage>312</lpage>
          , New York, NY, USA, (
          <year>2011</year>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <surname>Giovanna</surname>
            <given-names>Varni</given-names>
          </string-name>
          , Marie Avril, Adem Usta, and Mohamed Chetouani, '
          <article-title>Syncpy: a unified open-source analytic library for synchrony'</article-title>
          ,
          <source>in Proceedings of the 1st Workshop on Modeling INTERPERsonal SynchrONy</source>
          And infLuence,
          <source>INTERPERSONAL@ICMI</source>
          <year>2015</year>
          , Seattle, Washington, USA, November
          <volume>13</volume>
          ,
          <year>2015</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>47</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <surname>Johannes</surname>
            <given-names>Wagner</given-names>
          </string-name>
          , Florian Lingenfelser, Tobias Baur, Ionut Damian, Felix Kistler, and Elisabeth Andre´, '
          <article-title>The social signal interpretation (ssi) framework: Multimodal signal processing and recognition in real-time'</article-title>
          ,
          <source>in Proceedings of the 21st ACM International Conference on Multimedia, MM '13</source>
          , p.
          <fpage>831</fpage>
          -
          <lpage>834</lpage>
          , New York, NY, USA, (
          <year>2013</year>
          ).
          <article-title>Association for Computing Machinery</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <surname>Harald</surname>
            <given-names>G Wallbott</given-names>
          </string-name>
          ,
          <article-title>'In and out of context: Influences of facial expression and context information on emotion attributions'</article-title>
          ,
          <source>British Journal of Social Psychology</source>
          ,
          <volume>27</volume>
          (
          <issue>4</issue>
          ),
          <fpage>357</fpage>
          -
          <lpage>369</lpage>
          , (
          <year>1988</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [49]
          <string-name>
            <surname>Harald</surname>
            <given-names>G Wallbott</given-names>
          </string-name>
          , 'Bodily expression of emotion',
          <source>European journal of social psychology</source>
          ,
          <volume>28</volume>
          (
          <issue>6</issue>
          ),
          <fpage>879</fpage>
          -
          <lpage>896</lpage>
          , (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [50]
          <string-name>
            <surname>Rebekah</surname>
            <given-names>Wegener</given-names>
          </string-name>
          ,
          <source>Studying Language in Society and Society through Language: Context and Multimodal Communication</source>
          ,
          <fpage>227</fpage>
          -
          <lpage>248</lpage>
          ,
          <string-name>
            <given-names>Palgrave</given-names>
            <surname>Macmillan</surname>
          </string-name>
          <string-name>
            <surname>UK</surname>
          </string-name>
          , London,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [51]
          <string-name>
            <surname>Jacob</surname>
            <given-names>Whitehill</given-names>
          </string-name>
          , Zewelanji Serpell,
          <string-name>
            <surname>Yi-Ching</surname>
            <given-names>Lin</given-names>
          </string-name>
          , Aysha
          <string-name>
            <surname>Foster</surname>
          </string-name>
          , and Javier R Movellan, '
          <article-title>The faces of engagement: Automatic recognition of student engagementfrom facial expressions'</article-title>
          ,
          <source>IEEE Transactions on Affective Computing</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <fpage>86</fpage>
          -
          <lpage>98</lpage>
          , (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [52]
          <string-name>
            <surname>Wim</surname>
            <given-names>Wiegerinck</given-names>
          </string-name>
          , Willem Burgers, and Bert Kappen, Bayesian Networks,
          <source>Introduction and Practical Applications</source>
          ,
          <fpage>401</fpage>
          -
          <lpage>431</lpage>
          , Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [53]
          <string-name>
            <surname>Martin</surname>
            <given-names>Wo¨llmer</given-names>
          </string-name>
          , Marc Al-Hames, Florian Eyben, Bjo¨rn
          <string-name>
            <given-names>W.</given-names>
            <surname>Schuller</surname>
          </string-name>
          , and Gerhard Rigoll, '
          <article-title>A multidimensional dynamic time warping algorithm for efficient multimodal fusion of asynchronous data streams'</article-title>
          ,
          <source>Neurocomputing</source>
          ,
          <volume>73</volume>
          (
          <issue>1-3</issue>
          ),
          <fpage>366</fpage>
          -
          <lpage>380</lpage>
          , (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          [54]
          <string-name>
            <surname>Martin</surname>
            <given-names>Wo¨llmer</given-names>
          </string-name>
          , Angeliki Metallinou, Florian Eyben, Bjo¨rn
          <string-name>
            <given-names>W.</given-names>
            <surname>Schuller</surname>
          </string-name>
          , and Shrikanth S. Narayanan, '
          <article-title>Context-sensitive multimodal emotion recognition from speech and facial expression using bidirectional LSTM modeling'</article-title>
          ,
          <source>in INTERSPEECH</source>
          <year>2010</year>
          ,
          <article-title>11th Annual Conference of the International Speech Communication Association</article-title>
          , Makuhari, Chiba, Japan,
          <source>September 26-30</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>2362</fpage>
          -
          <lpage>2365</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          [55]
          <string-name>
            <surname>Martin</surname>
            <given-names>Wo¨llmer</given-names>
          </string-name>
          , Bjo¨rn W. Schuller, Florian Eyben, and Gerhard Rigoll, '
          <article-title>Combining long short-term memory and dynamic bayesian networks for incremental emotion-sensitive artificial listening'</article-title>
          ,
          <source>J. Sel. Topics Signal Processing</source>
          ,
          <volume>4</volume>
          (
          <issue>5</issue>
          ),
          <fpage>867</fpage>
          -
          <lpage>881</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          [56]
          <string-name>
            <given-names>W.</given-names>
            <surname>Yun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          , '
          <article-title>Automatic recognition of children engagement from facial video using convolutional neural networks'</article-title>
          ,
          <source>IEEE Transactions on Affective Computing</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>1</lpage>
          , (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          [57]
          <string-name>
            <surname>Heinz</surname>
            <given-names>Zimmermann</given-names>
          </string-name>
          , Speaking, listening, understanding,
          <source>SteinerBooks</source>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>