<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Emotion-Recognition from Speech-based Interaction in AAL Environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>B. De Carolis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>S. Ferilli</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>G. Palestra</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Redavid</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dipartimento di Informatica, Università di Bari 70126 Bari</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In Ambient Assisted Living environments assistance and care are delegated to the intelligence embedded in the environment that, in our opinion, should provide not only a task-oriented support but also an interface able to establish a social empathic relation with the user. To this aim social assistive robots are being employed as a mediator interface and, in order to achieve a relation with the user, they should be endowed with the capability of recognizing the user affective state. Since a natural way to interact with a robot is speech, spoken user's input can be used to give to the robot the capability of recognizing the emotions and attitude of the user, thus providing more detail information about the user state. This paper focuses on this topic and proposes an approach based on the dimensional model of emotions in which the valence and arousal of user's spoken input are recognized. The experimental analysis shows the performance in terms of accuracy of the proposed approach on an Italian dataset. In order to show its application in the context of Ambient Assisted Living, an example is provided.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <sec id="sec-1-1">
        <title>A Smart Environment should support people in their daily activities by assisting and</title>
        <p>
          facilitating users when interacting with environment services in a natural and easy
way. The required assistance may be provided to the user through different devices.
The choice of an assistive robot agent as an interface is supported by several
considerations. First of all, the robot has a physical presence and it may participate in the
user’s daily life. Assistive robots can move around and perform actions, follow and
observe the user in the environment, which is fundamental when designing supportive
technologies for elderly people [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. In addition to typical service-oriented features,
assistive robots can be equipped with social and conversational capabilities, thus
improving the naturalness and effectiveness of the interaction between users and smart
environment services.
        </p>
        <p>
          Speech is a natural way for humans to interact with robots [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Moreover, speech
based interaction is seen as an effective interface for smart environments because it is
natural, hands-free and it enables different types of users with different capabilities,
and disabilities, to interact with systems. Since elderly people are an important user
group for smart environments, spoken interaction is of particular benefit for them
since it is natural and does not require particular skills. In addition, spoken user’s
input can be used not only to issue commands, but also to give to the robot the
capability of recognizing the emotions and attitude of the user and this is very important for
establishing a social relation and to personalize service execution. Indeed, providing
personalized services requires taking into account several factors, which are related to
the nature of the service, to user’s preferences and to context-related features such as
user’s emotional state.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>In this paper we focus on the latter issue and we will present an acoustic analyzer</title>
        <p>
          for the recognition of the emotion. This module is able to extract the prosodic features
of user’s spoken input and, starting from them is able to recognize the two dimension
of emotions: valence and arousal [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ]. Then, the module has been used by a social
assistive robot embodied in NAO. The robot acts as Interactor Agent in a smart home
environment implemented as a Multi Agent System (MAS) [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>The experimental analysis shows the performance in terms of accuracy of the proposed approach on an Italian dataset. The obtained results also show which combination of features assures a satisfying recognition rate allowing a better understanding of the user’s affective state.</title>
      </sec>
      <sec id="sec-1-4">
        <title>The paper is structured as follows. In Section 2 the motivations and technical</title>
        <p>background for this work are presented. Section 3 briefly describes the MAS
architecture implementing the smart home environment. Section 4 describes how VOCE has
been developed and Section 5 shows an example on how it can be applied in the
context of AAL. Conclusions and future work directions are illustrated in Section 6.
2</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Background and Motivations</title>
      <p>Interaction with services provided by a smart environment may be provided to the
user in a seamless way (i.e. by combining smart home technologies based on sensors
and effectors embedded in the appliances of the environment), or using an embodied
companion as an interface, or combining both approaches. In all cases, research
emphasizes the need of natural and user-friendly interfaces for accessing the services
provided by the environment. Moreover, research on social and affective computing
suggests that such an assistive environment should provide not only a task-oriented
support but also an interface able to establish a social empathic relation with the user.</p>
      <p>
        Several studies report successful results on how social assistive robots can be
employed as interface in the assisted living domain. For instance, projects ROBOCARE
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], Nursebot [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], Care-o-bot [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], CompaniAble [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and Ksera [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] aim at creating
assistive intelligent environments for the elderly in which robots offer support to the
user at home. However, to be accepted and integrated in the user’s daily life, interaction
with robots must be spontaneous and natural, and to provide a friendly environment
robots must exhibit social capabilities and learn how to react according to the human
emotional state. Since speech provides a natural and intuitive way for people to
interact with robots, automatic emotional speech recognition will expand the possibilities
of interaction.
      </p>
      <sec id="sec-2-1">
        <title>Emotions are expressed through various communicative signals in humans: facial</title>
        <p>
          expressions [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], vocal features [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], body movements and postures [
          <xref ref-type="bibr" rid="ref10 ref11">10,11</xref>
          ], or a
combination of some of them [
          <xref ref-type="bibr" rid="ref13 ref14 ref15">13,14,15</xref>
          ]. In this paper we focus on speech features and
how it is possible to use them to recognize emotions in communication with humans.
        </p>
        <p>
          Recognizing emotions in speech through several features has been a key research
issue in robotics, because by recognizing emotional factors the robot can handle social
situations. In emotional classification from speech, a multitude of different features
denoting prosodic cues have been used. Prosodic features, like pitch, loudness,
speaking rate, durations, pause and rhythm were proven to have strong correlations between
them, providing emotional information. In the case of the analysis of an entire
segment of voice, statistical functions like mean, median, minimum, maximum, standard
deviation are applied to the fundamental frequency (F0) base contour [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Taking
advantage of research work in Music Information Retrieval, Mel Frequency Cepstral
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Coefficients (MFCCs) are also used with great accuracy in emotion recognition [17].</title>
      </sec>
      <sec id="sec-2-3">
        <title>These features can be used to train a classifier and the learned model can be used to detect emotion in real-time situations.</title>
      </sec>
      <sec id="sec-2-4">
        <title>Several classifiers have been used in this field. Each of them has advantages and</title>
        <p>
          disadvantages in order to deal with the speech emotion recognition problem. The
more common group includes Hidden Markov Models (HMM) [
          <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
          ] regarded as
the simplest dynamic Bayesian networks, artificial neural networks (ANN) [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ],
support vector machines (SVM) [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], k-NN [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] and Decision Trees [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ].
        </p>
        <p>The majority of emotion recognition systems from speech have employed a
highdimensional speech grouped in a big vector of features. In this paper, the most
commonly used features in several researches for capturing emotional speech
characteristics in time and frequency were selected. The performance of different well known
classifiers was compared in order to select the best result to predict the emotion, based
on speech emotional data.
3.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Overview of the MAS</title>
      <sec id="sec-3-1">
        <title>In [24] we propose an approach based on software agents able to provide what we call</title>
      </sec>
      <sec id="sec-3-2">
        <title>Smart Services. A smart service can been defined as an integrated, interoperable and</title>
        <p>personalized service, accessible through several interfaces available on various
devices present in the environment in the optic of pervasive computing.</p>
        <p>The objective of the proposed approach is to recognize the users goal starting from
percepts (sensors data, user actions, etc.) and provide them with a smart service that
integrate elementary services according to the situation. In order to achieve this aim,
the environment has to be able to reason on the situation of the user so as to
understand which are his/her needs and goals through the composition of the most
appropriate smart service. The idea underlying our approach is the metaphor of the butler in
grand houses, who can be seen as an household affairs manager with duties of a
personal assistant, able to organize the housestaff in order to satisfy the needs of the
house inhabitants. To this aim, taking into account the results of a previous project, we
have developed a MAS in which the butler agent has to recognize the situation of the
user, based on interaction with Sensor Agent, in order to infer possible user’s goals.</p>
      </sec>
      <sec id="sec-3-3">
        <title>The recognized goals are then used to select the most suitable workflow among a set</title>
        <p>of available candidates representing a smart service. Such a selection is made by
matching semantically the goals, the current situation features and the effects expected by
the execution of the workflow. Once a workflow has been selected, its actions are
executed by the effector agents.</p>
        <p>One important feature of this architecture is the presence of an agent designed to
take care of the interaction with the users. In completely proactive approach, in fact,
users may feel a loss of control over the system actions. Therefore we adopt a
semiautomatic approach composition of services. The butler proactively propose smart
services and leaves, at the same time, the control over proposed service composition to
the user to select alternative services, to provide more preference information in order
to get a better personalization, to ask for explanation about the proposed services and
so on.</p>
        <p>The MAS is constituited by the following classes of agents:
- Sensor Agents (SA) are used for providing information about context parameters
and features (e.g., temperature, light level, humidity, etc.) at a higher abstraction
level than sensor data.
- Butler Agent (BA) reasons on the user’s goals and devises the workflow to
satisfy them (see Figure 1).
- Effector Agents (EA) each appliance and device is controlled by an EA that
reasons on the opportunity of performing an action instead of another in the current
context.
- Interactor Agent (IA) is in charge of handling interaction with the user in order
to carry on communicative tasks. In this case the IA is embodied in the NAO
Robot.
- Housekeeper Agent (HA) acts as a facilitator since it knows all the agents that
are active in the house and also the goal they are able to fulfill.</p>
        <p>Fig. 1. The MAS architecture</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. VOCE: VOice Classifier of Emotions</title>
      <p>
        Emotions can be classified using two main approaches. Discrete emotion models
focus on a defined set of labels denoting emotions (e.g. anger, fear, disgust, happiness,
surprise and sadness to name the most common ones). The discrete emotion model
has the advantage of clearly distinguishing categories of emotions, however the the
labels and their number differ a lot from one model to another. By contrast,
dimensional models describe the affective space within a limited amount of dimensions. For
instance in the circumplex model of emotions [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] only two dimensions are used to
represent an emotional state: the valence (from positive to negative or pleasant vs.
unpleasant) and the arousal (from high to low or aroused vs. relaxed) dimensions. In
contrast to discrete emotions, each emotion can be mapped within this space and this
model can be used to determine mixtures of different emotions because they are
represented by points in a space.
      </p>
      <sec id="sec-4-1">
        <title>In VOCE we decided to adopt the dimensional model to classify emotions. There</title>
        <p>fore, the analysis of the prosody user’s spoken utterance is made by two classifiers:
one for recognizing the valence dimension and the other for the arousal one.</p>
      </sec>
      <sec id="sec-4-2">
        <title>To this aim we developed a web-service called VOCE 2.0 (VOice Classifier of</title>
      </sec>
      <sec id="sec-4-3">
        <title>Emotions ver. 2.0) that classifies the valence and arousal of the voice prosody with an</title>
        <p>approach very similar to the one described in [26]. The major steps in speech
emotion recognition are audio segmentation, feature extraction and the actual
classification of the feature vectors into valence and arousal values.</p>
        <p>VOCE can be used in two ways: offline for creating and analysing the emotional
speech corpus (Figure 2) and, being a web-service, online for tracking the affect in
voice in real-time. While the off-line version allows to build the classifier, the online
emotion recognition just outputs the recognised emotions valence and arousal and
maps the combination of these values into one of the basic emotions by providing the
emotion label.</p>
        <p>Let’s see now how the classifier has been trained and how we use it in real-time.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Dataset</title>
      <p>Albeit our approach is based on the dimensional model, since we could not found any
corpus, among the few available for Italian, in which emotions were annotated
according to their valence and arousal we used the ∈motion dataset [27]. Among the
three availble ones the ∈motion dataset has been used for EVALITA ERT challenge1
and therefore we could compare our results with other research works in this domain.</p>
      <p>
        The emotional speech characteristics were extracted from the Italian subset of
∈motion contains 220 audio files corresponding to sentences for the 6 basic emotions
(joy, anger, surprise, sadness, disgust, fear) and the neutral one recorded by
professional actors. In order to use the dimensional approach on this dataset we mapped
each emotion to the correspondent valence (negative, neutral and positive) and
arousal (low, medium, high) using the approach explained in [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. For instance, “anger” is
mapped into negative valence and high arousal, while “sadness” is mapped into
negative valence and low arousal.
      </p>
      <sec id="sec-5-1">
        <title>3.1.2 Features extraction and classification</title>
      </sec>
      <sec id="sec-5-2">
        <title>In developing VOCE we exploited different combinations of features and several</title>
        <p>classification algorithms. For this task we used Praat [28]. In particular, besides pitch
and energy related features, we extracted features related to the spectrum, harmonicity
and the Mel-Frequency Cepstral Coefficients (MFCCs) that is used to describe a
spectrum frame, its first and second derivative in time are used to reflect dynamic
changes.</p>
        <p>Table 1 shows the features extracted from each audio file using Praat functions.</p>
        <p>In order to find the best set of features we tested three conditions with several
classification algorithms:
- Support Vector Machines (SVM), which offers robust classification to a very
large number of variables and small samples.
- Decision trees, that work with simple classification rules that are easy to
understand. The rules represent the information in a tree based on a set of
features.
- Artificial Neural Network (ANN), and in particular the Multilayer Perceptron
algorithm.
- k-Nearest Neighbors (kNN) is one of the simplest of classification algorithms
available for supervised learning. It classifies unlabeled examples based on
their similarity with examples in the training set.</p>
      </sec>
      <sec id="sec-5-3">
        <title>The three set of features were:</title>
        <p>- ALL: all the attributes in Table 1;
- MFCC: MFCC features only;
- No MFCC: all the features except MFCC.</p>
      </sec>
      <sec id="sec-5-4">
        <title>1 http://www.evalita.it/2014/tasks/emotion</title>
        <p>Feature  
Pitch  
PitchMin  
PitchMed  
PitchMax  
PitchMinMaxDiffLog  
PitchMinLog  
PitchMedLog  
PitchMaxLog  
PitchDevSta  
PitchSlope  
Energy  
EnergyMin  
EnergyMed  
EnergyMax  
EnergyMinMaxDiff  
EnergyDevSta  
Spectrum  
SpectrumCentralMoment  
SpectrumDevSta  
SpectrumGravityCentre  
SpectrumKurtosis  
SpectrumSkewness  
Harmonicity  
HarmonicityMin  
HarmonicityMed  
HarmonicityMax  
HarmonicityDevSta  
MFCC  
MFCCnMin  
MFCCnMed  
MFCCnMax  
MFCCnDevSta    
Description  
Minimum  value  
Average  value  
Maximum  value  
Logarithmic  differentiation  
Minimum  Logarithmic  
Average  logarithmic  
Maximum  logarithmic  
Standard  Deviation  
Slope  
Minimum  value  
Average  value  
Maximum  value  
Logarithmic  differentiation  
Standard  Deviation  
Central  moment  
Standard  Deviation  
Central  tendency  
Degree  of  centralization  
Degree  of  asymmetry  
Minimum  value  
Average  value  
Maximum  value  
Standard  Deviation  
Minimum  of  nth  MFCC  
Average  of  nth  MFCC  
Maximum  of  nth  MFCC  
Standard  Deviation  of  nth  MFCC  </p>
      </sec>
      <sec id="sec-5-5">
        <title>From the analysis of the performance of the most commonly used algorithms for</title>
        <p>classification starting from numeric features the most accurate one were MLP
(Multi</p>
      </sec>
      <sec id="sec-5-6">
        <title>Layer Perceptron) and SMO (Sequential Minimal Optimization) algorithm for trai</title>
        <p>ning a support vector classifier in Weka2.</p>
      </sec>
      <sec id="sec-5-7">
        <title>The accuracy was validated using a 10 Fold Cross Validation technique. A k-fold</title>
        <p>cross-validation with k = 10 was used to make validations over the classifiers. This
technique allowed the evaluation of the model facing an unknown dataset. Results of
the classification of valence, arousal and derived emotion labels are shown in Table 2.</p>
        <p>Results show that, for both algorithms, using the complete set of features improves
accuracy, however using only MFCC related features we get an accuracy comparable</p>
      </sec>
      <sec id="sec-5-8">
        <title>2 http://www.cs.waikato.ac.nz/ml/weka/</title>
        <p>with the one obtained using the all set of features. The worst setting is when MFCC is
not considered. As far as the choice of algorithm is concerned, even if MLP had a
slight better accuracy, the time to create the model and classify a vocal input is higher
(100:1). Since VOCE has to be employed in real-time classification tasks we selected
SMO.</p>
        <p>Since the arousal dimension is related to the importance of the goal and the
valence dimension is related to the achievement vs. the threatening of the goal, our speech
classifier performs well in recognizing negative states, like those related to anger, and
allows us to distinguish positive from negative attitudes. However, as expected, some
emotions are easier to recognize than others. For example, humans are much better at
recognizing anger than happiness; therefore, our results can be considered acceptable
under this view.</p>
        <p>Features
ALL</p>
        <p>Valence
Arousal</p>
        <p>Emotion
MFCC</p>
        <p>Valence
Arousal</p>
        <p>Emotion
No_MFCC</p>
        <p>Valence
Arousal
Emotion</p>
        <p>MLP
70,45
80,90
71,36
69,09
80,00
70,45
64,09
74,09
53,18</p>
        <p>SMO
71,36
77,27
68,63
64,54
75,00
68,18
55,45
69,09
53,18</p>
      </sec>
      <sec id="sec-5-9">
        <title>Comparing our results with other works based on the same dataset [29] we can say that our approach has a comparable accuracy over the set of basic emotions.</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>4 An Example of Application in the Context of AAL</title>
      <sec id="sec-6-1">
        <title>VOCE has been used in real-time as a web service with the NAO Robot for ena</title>
        <p>bling emotion recognition during speech-based interaction (see Figure 3).</p>
      </sec>
      <sec id="sec-6-2">
        <title>We have designed this architecture to endow the Aldebaran NAO robot with this</title>
        <p>capability. The system is composed by two fundamental units: the NAO humanoid
robot and the workstation connected with NAO robot. Audio files in wav format,
recorded from 4 microphones located in the head of the NAO, are collected by the</p>
      </sec>
      <sec id="sec-6-3">
        <title>Application Programming Interface (API) provided with NAO Software Development</title>
      </sec>
      <sec id="sec-6-4">
        <title>Kit (SDK). Captured audio files are sent to the Speech-based Interface module in</title>
        <p>order to allows the understaning of vocal commands and to recognize the valence and
arousal. An Automatic Speech Recognition Module (ASR) performs the first task
whereas the second task is accomplished by Voice Classifier for Emotions (VOCE)</p>
      </sec>
      <sec id="sec-6-5">
        <title>Module. Then, the results are sent to Behavior Decision Module that choose the appropriate behavior and send it to the robot to be executed. Communication between the robot and the workstation has been performed using the NAOqi API.</title>
        <p>Speech-Based
Interface</p>
      </sec>
      <sec id="sec-6-6">
        <title>Let us now provide an example of how the proposed approach can be integrated in an ambient assisted living environment.</title>
      </sec>
      <sec id="sec-6-7">
        <title>The scenario depicts a situation in which an old man lives in a house equipped</title>
        <p>with some sensors (to gather data about the house situation) and some effectors (to
control appliances in the environments). Moreover the house is equipped with the</p>
      </sec>
      <sec id="sec-6-8">
        <title>NAO robot acting as a natural and social interface between the user and the smart home environment, by implementing an easy conversational access to the (digital or physical) services of the environment.</title>
        <p>It’s friday evening and Nicola, a 73 y.o. man, is at home alone. He has an
appointment with his friends downtown to play cards like he does almost every
friday. Nicola is sitting on the bench in the living room that is equipped with
sensors and effectors for controlling appliances in the room and with the
social robot that acts as a mediator interface between the environment services
and the user (see Figure 4). Nicola received a message saying that his
doughter cannot accompany him downtown and this makes him a bit angry.
Nicola calls NAO to try to find a solution.</p>
      </sec>
      <sec id="sec-6-9">
        <title>In the following we provide an example of the interaction.</title>
      </sec>
      <sec id="sec-6-10">
        <title>R: ‘Hi Nicola, what can I do for you?’</title>
      </sec>
      <sec id="sec-6-11">
        <title>Nicola: ‘Damn … (with high arousal and negative valence) I need to go downtown to play cards with my friends and my doughter cannot come to bring me there …I cannot miss it tonight there is a tournament !’ (with high arousal and negative valence).</title>
      </sec>
      <sec id="sec-6-12">
        <title>R: ‘Don't’ be angry for this … we will find a solution’ Do you want me to call your daughter to ask for the permission to call a taxi to bring you there and take you back at a certain time?”.</title>
      </sec>
      <sec id="sec-6-13">
        <title>Nicola: ‘Yes … but you know I need a bit of assistance in walking from the car to the bar’ (with medium arousal and negative valence).</title>
      </sec>
      <sec id="sec-6-14">
        <title>Nicola: I will not play card with my friends tonight, I feel so lonely. (with low arousal and negative valence).</title>
      </sec>
      <sec id="sec-6-15">
        <title>R: ‘Oh, I’m sorry to hear that you are sad. I will ask the taxi driver to help you. OK?</title>
      </sec>
      <sec id="sec-6-16">
        <title>Nicola: OK.</title>
      </sec>
      <sec id="sec-6-17">
        <title>R: the robot send a message to the daughter that accepts and then calls the taxi.</title>
        <p>In this scenario initially the voice classifier recognizes a negative valence with a
high arousal from the prosody of the spoken utterance (Figure 2). This is interpreted
as anger and the robot besides expressing empathy (it show to understand the user’s
feelings) asks the reasons for it. When the robot understands that the user’s goal is to
go downtown it finds a workflow satisfying this goal by matching the constraints of
the situation (the daughter is not available). According to the selected workflow the
dialog goes on in order to get some information that are necessary (preconditions) to
execute some of its step (like the permission to call the taxi given from Nicola’s
daughter).</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5 Conclusions and Future Work</title>
      <p>We presented a preliminary work towards implementing a speech-based interface
between a Social Robot and a user for handling interaction in a smart environment. In
our opinion, besides assisting the elderly user in performing tasks, the robot has to
establish a social long-term relationship with the user so as to enforce trust and
confidence. The underlying idea of our work, in fact, is that the analysis of user’s spoken
utterances can be used for both issuing requests to the robot and understaning his
emotional state, which is important when the interaction happens in everyday life
environments. To this aim we developed VOCE 2.0 a module to classify emotions from
features extracted from the speech signal according to the dimensional model (valence
and arousal). The recognition accuracy results are comparable with other research
works in which the same dataset was used [29]. We are aware that an improvement is
needed and to this aim we plan to collect a new dataset possibly with example of
elderly voices, which may have different range of features from those used in the
∈motion dataset. Information on user’s emotions coupled to context and activity
recognition may give the robot the capability to infer which is or will be the most
probable user’s mood in a given context.
26. W.E. Bosma and E. André, “Exploiting Emotions to Disambiguate Dialogue Acts”, in Proc.
2004 Conference on Intelligent User Interfaces, January 13 2004, N.J. Nunes and C. Rich
(eds), Funchal, Portugal, pp. 85-92, 2004.
27. Vincenzo Galata’. 2010. Production and perception of vocal emotions: a
crosslinguistic and cross-cultural study. Ph.D. thesis, University of Calabria
28. www.praat.com
29. Antonio Origlia, Vincenzo Galatà e Bogdan Ludusan. Automatic classification of emotions
via global and local prosodic features on a multilingual emotional database. In: Speech
Prosody. 2010.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Thrun</surname>
          </string-name>
          ,
          <article-title>Towards a framework for human-robot interaction</article-title>
          ,
          <source>Human Computer Interaction</source>
          .
          <volume>19</volume>
          (
          <issue>1</issue>
          &amp;2), pp.
          <fpage>9</fpage>
          -
          <lpage>24</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Drygajlo</surname>
            ,
            <given-names>P.J.</given-names>
          </string-name>
          <string-name>
            <surname>Prodanov</surname>
            , G. Ramel G.,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Meisser</surname>
            , and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Siegwart</surname>
          </string-name>
          , “
          <article-title>On developing a voice-enabled interface for interactive tour-guide robots”</article-title>
          .
          <source>Journal of Advanced Robotics</source>
          , vol.
          <volume>17</volume>
          ,
          <issue>nr</issue>
          . 7,p.p.
          <fpage>599</fpage>
          -
          <lpage>616</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>3. B. De Carolis</surname>
            , G. Cozzolongo ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Pizzutilo</surname>
            ,
            <given-names>V. L.</given-names>
          </string-name>
          <string-name>
            <surname>Plantamura</surname>
          </string-name>
          ,
          <article-title>Agent-Based home simulation and control</article-title>
          ,
          <source>Proceedings of the 15th international conference on Foundations of Intelligent Systems, May 25-28</source>
          ,
          <year>2005</year>
          ,
          <string-name>
            <given-names>Saratoga</given-names>
            <surname>Springs</surname>
          </string-name>
          , NY
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cesta</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Cortellessa</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Pecora</surname>
            and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Rasconi</surname>
          </string-name>
          ,
          <article-title>Supporting Interaction in the RoboCare Intelligent Assistive Environment</article-title>
          , AAAI 2007
          <string-name>
            <given-names>Spring</given-names>
            <surname>Symposium</surname>
          </string-name>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Pineau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Montemerlo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pollack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Roy</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Thrun</surname>
          </string-name>
          ,
          <article-title>Towards Robotic Assistants in Nursing Homes: Challenges and Results</article-title>
          ,
          <source>Robotics and Autonomous Systems</source>
          <volume>42</volume>
          (
          <issue>3-4</issue>
          ), pp.
          <fpage>271</fpage>
          -
          <lpage>281</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Graf</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hans</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schraft</surname>
            <given-names>RD</given-names>
          </string-name>
          (
          <year>2004</year>
          )
          <article-title>Care-O-bot II - development of a next generation robotics home assistant</article-title>
          .
          <source>Auton. Robots</source>
          <volume>16</volume>
          ,
          <fpage>193</fpage>
          -
          <lpage>205</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. CompanionAble project (
          <year>2011</year>
          ) http://www.companionable.net
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>8. ksera.ieis.tue.nl</mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>M.</given-names>
            <surname>Pantic</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.J.M.</given-names>
            <surname>Rothkrantz</surname>
          </string-name>
          , “
          <article-title>Automatic analysis of facial expressions: The state of the art</article-title>
          ,
          <source>” IEEE Trans. on Pattern Analysis and Machine Intelligence</source>
          , vol.
          <volume>22</volume>
          , no.
          <issue>12</issue>
          , pp.
          <fpage>1424</fpage>
          -
          <lpage>1445</lpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>R.</given-names>
            <surname>Cowie</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Douglas-Cowie</surname>
          </string-name>
          , “
          <article-title>Automatic statistical analysis of the signal and prosodic signs of emotion in speech,”</article-title>
          <source>In Proc. International Conf. on Spoken Language Processing</source>
          , pp.
          <fpage>1989</fpage>
          -
          <lpage>1992</lpage>
          ,
          <year>1996</year>
          . [
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. N.
          <string-name>
            <surname>Bianchi-Berthouze</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kleinsmith</surname>
          </string-name>
          , “
          <article-title>A categorical approach to affective gesture recognition</article-title>
          ,
          <source>” Connection Science</source>
          , vol.
          <volume>15</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>259</fpage>
          -
          <lpage>269</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. G. Castellano,
          <string-name>
            <given-names>S.D.</given-names>
            <surname>Villalba</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Camurri</surname>
          </string-name>
          , “
          <article-title>Recognising Human Emotions from Body Movement and Gesture Dynamics,”</article-title>
          <source>In Proc. of 2nd International Conference on Affective Computing and Intelligent Interaction</source>
          , Berlin, Heidelberg,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>H. K. M. Meeren</surname>
          </string-name>
          , C. van Heijnsbergen and B. de Gelder, “
          <article-title>Rapid perceptual integration of facial expression and emotional body language</article-title>
          ,
          <source>” Proc. National Academy of Sciences of the USA</source>
          , vol.
          <volume>102</volume>
          , no.
          <issue>45</issue>
          , pp.
          <fpage>16518</fpage>
          -
          <lpage>16523</lpage>
          ,
          <year>2005</year>
          . [
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>A.</given-names>
            <surname>Metallinou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Katsamanis</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          , “
          <article-title>Tracking changes in continuous emotion states using body language and prosodic cues,”</article-title>
          <source>In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , pp.
          <fpage>2288</fpage>
          -
          <lpage>2291</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>C. Busso</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Deng</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Yildirim</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Bulut</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>A. Kazemzaeh. S.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Neumann</surname>
            and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Narayanan</surname>
          </string-name>
          , “
          <article-title>Analysis of Emotion Recognition using Facial Expressions, Speech and Multimodal information,”</article-title>
          <source>In Proc. of ACM 6th int'l Conf. on Multimodal Interfaces (ICMI2004)</source>
          , State College, PA, pp.
          <fpage>205</fpage>
          -
          <lpage>211</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>D.</given-names>
            <surname>Ververidis</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Kotropoulos</surname>
          </string-name>
          , “
          <article-title>Emotional speech recognition: Resources, features</article-title>
          , and methods,” Speech Communication, pp.
          <fpage>1162</fpage>
          -
          <lpage>1181</lpage>
          ,
          <year>2006</year>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>B.</given-names>
            <surname>Bogert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Healy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Tukey</surname>
          </string-name>
          , “
          <article-title>The quefrency alanysis of time series for echoes: cepstrum, pseudo-autocovariance, cross- cepstrum, and saphe-cracking,”</article-title>
          <source>Proceedings of the Symposium on Time Series Analysis</source>
          , Wiley,
          <year>1963</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <given-names>D.</given-names>
            <surname>Le</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Provost</surname>
          </string-name>
          , “
          <article-title>Emotion recognition from spontaneous speech using Hidden Markov models with deep belief networks,” in Automatic Speech Recognition and Understanding (ASRU</article-title>
          ),
          <source>2013 IEEE Workshop on</source>
          , pp.
          <fpage>216</fpage>
          -
          <lpage>221</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>J. Wagner</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Vogt</surname>
          </string-name>
          , and E. Andre, “
          <article-title>A systematic comparison of different ´ HMM designs for emotion recognition from acted and spontaneous speech</article-title>
          ,”
          <source>in Proceedings of the 2nd International Conference on Affective Computing and Intelligent Interaction (ACII)</source>
          , Lisbon, Portugal, pp.
          <fpage>114</fpage>
          -
          <lpage>125</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>S.A.</given-names>
            <surname>Firoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.A.</given-names>
            <surname>Raj</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.P.</given-names>
            <surname>Babu</surname>
          </string-name>
          , “
          <article-title>Automatic Emotion Recognition from Speech Using Artificial Neural Networks with Gender-Dependent Databases,” in Advances in Computing, Control and Telecommunication Technologies</article-title>
          ,
          <source>ACT '09</source>
          , pp.
          <fpage>162</fpage>
          -
          <lpage>164</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>C. Yu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Tian</surname>
          </string-name>
          , F. Cheng and S. Zhang, “
          <article-title>Speech Emotion Recognition Using Support Vector Machines</article-title>
          ,” in
          <source>Advanced Research on Computer Science and Information Engineering</source>
          . vol.
          <volume>152</volume>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Shen</surname>
          </string-name>
          and
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          , Eds., ed: Springer Berlin Heidelberg, pp.
          <fpage>215</fpage>
          -
          <lpage>220</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>M.</given-names>
            <surname>Feraru</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Zbancioc</surname>
          </string-name>
          , “
          <article-title>Speech emotion recognition for SROL database using weighted KNN algorithm</article-title>
          ,” in Electronics,
          <source>Computers and Artificial Intelligence (ECAI)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>C.-C. Lee</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Mower</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Busso</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Narayanan</surname>
          </string-name>
          , “
          <article-title>Emotion recognition using a hierarchical binary decision tree approach,” Speech Commun</article-title>
          , vol.
          <volume>53</volume>
          , pp.
          <fpage>1162</fpage>
          -
          <lpage>1171</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>B. De Carolis</surname>
            and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ferilli</surname>
          </string-name>
          .
          <article-title>A multiagent system providing situation-aware services in a smart environment</article-title>
          .
          <source>Workshop on Ambient Intelligence Infrastructures (WAmIi)</source>
          .
          <source>November 13</source>
          ,
          <year>2012</year>
          , Pisa, Italy.In conjunction with
          <source>International Joint Conference on Ambient Intelligence (AmI</source>
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Russell</surname>
          </string-name>
          ,
          <string-name>
            <surname>James</surname>
          </string-name>
          (
          <year>1980</year>
          ).
          <article-title>"A circumplex model of affect"</article-title>
          .
          <source>Journal of Personality and Social Psychology</source>
          <volume>39</volume>
          :
          <fpage>1161</fpage>
          -
          <lpage>1178</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>