<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Romeo2 Project: Humanoid Robot Assistant and Companion for Everyday Life: I. Situation Assessment for Social Intelligence 1</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Amit Kumar Pandey</string-name>
          <email>akpandey@aldebaran.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rodolphe Gelin</string-name>
          <email>rgelin@aldebaran.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rachid Alami</string-name>
          <email>rachid.alami@laas.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renaud Viry</string-name>
          <email>renaud.viry@laas.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Axel Buendia</string-name>
          <email>axel.buendia@cnam.fr</email>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roland Meertens</string-name>
          <email>rolandmeertens@gmail.com</email>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Chetouani</string-name>
          <email>mohamed.chetouani@upmc.fr</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laurence Devillers</string-name>
          <email>devil@limsi.fr</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marie Tahon</string-name>
          <email>marie.tahon@limsi.fr</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Filliat</string-name>
          <email>lliat@ensta-paristech.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yves Grenier</string-name>
          <email>yves.grenier@telecom-paristech.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mounira Maazaoui</string-name>
          <email>maazaoui@telecom-paristech.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Abderrahmane Kheddar</string-name>
          <email>kheddar@gmail.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fr´ed´eric Lerasle</string-name>
          <email>frederic.lerasle@laas.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laurent Fitte Duval</string-name>
          <email>ttedu@laas.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Aldebaran, A-Lab</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CNRS, LAAS</institution>
          ,
          <addr-line>7 avenue du colonel Roche, F-31400 Toulouse</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ENSTA ParisTech - INRIA FLOWERS</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>ISIR</institution>
          ,
          <addr-line>UPMC</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Inst. Mines-T ́el ́ecom; T ́el ́ecom ParisTech; CNRS LTCI</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>LIMSI-CNRS University Paris-Sorbonne</institution>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>Spirops/CNAM (CEDRIC)</institution>
          ,
          <addr-line>Paris</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>For a socially intelligent robot, dierent levels of situation assessment are required, ranging from basic processing of sensor input to high-level analysis of semantics and intention. However, the attempt to combine them all prompts new research challenges and the need of a coherent framework and architecture. This paper presents the situation assessment aspect of Romeo2, a unique project aiming to bring multi-modal and multi-layered perception on a single system and targeting for a unified theoretical and functional framework for a robot companion for everyday life. It also discusses some of the innovation potentials, which the combination of these various perception abilities adds into the robot's socio-cognitive capabilities.</p>
      </abstract>
      <kwd-group>
        <kwd>Situation Assessment</kwd>
        <kwd>Socially Intelligent Robot</kwd>
        <kwd>Human Robot Interaction</kwd>
        <kwd>Robot Companion</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>As robots started to co-exist in a human-centered environment, the human
awareness capabilities must be considered. With safety being a basic requirement, such
robots should be able to behave in a socially accepted and expected manner. This
requires robots to reason about the situation, not only from the perspective of
physical locations of objects, but also from that of ‘mental’ and ‘physical’ states
of the human partner. Further, such reasoning should build knowledge with the
human understandable attributes, to facilitate natural human-robot interaction.</p>
      <p>The Romeo2 project (website 1 ), the focus of this paper, is unique in that it
brings together dierent perception components in a unified framework for
reallife personal assistant and companion robot in an everyday scenario. This paper
outlines our perception architecture, the categorization of basic requirements, the
key elements to perceive, and the innovation advantages such a system provides.
Mr. Smith lives alone (with his Romeo robot
companion). He is elderly and visually
impaired. Romeo understands his speech, emotion
and gestures, assists him in his daily life. It
provides physical support by bringing the ‘desired’
items, and cognitive support by reminding about
medicine, items to add in to-buy list, playing
memory games, etc. It monitors Mr. Smith’s Fig. 1. Romeo robot and sensors.
activities and calls for assistance if abnormalities are detected in his behaviors. As
a social inhabitant, it plays with Mr. Smith’s grandchildren visiting him.</p>
      <p>
        This outlined partial
target scenario of Romeo2 project
(also illustrated in fig. 2),
depicts that being aware about
human, his/her activities, the
environment and the situation
are the key aspects towards
practical achievement of the
project’s objective.
Situation awareness is the ability to perceive and abstract information from
the environment [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. It is an important aspect of day-to-day interaction,
decisionmaking, and planning, so as important is the domain-based identification of the
elements and attributes, constituting the state of the environment. In this paper, we
will identify and present such elements from companion robot domain perspective,
sec. 2.2. Further, three levels of it have been identified (Endsley et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]): Level
1 situation awareness: To perceive the state of the elements composing the
surrounding environment. Level 2 situation awareness: To build a goal oriented
understanding of the situation. Experience and comprehension of the meaning are
important. Level 3 situation awareness: To project on the future. Sec. 2.1 will
present our sense-interact perception loop and map these levels.
      </p>
      <p>
        Further, there have been eorts to develop integrated architecture to utilize
multiple components of situation assessment. However, most of them are
specific for a particular task like navigating [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], intention detection [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], robot’s
self-perception [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], spatial and temporal situation assessment for robot passing
through a narrow passage [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], laser data based human-robot-location situation
assessment, e.g. human entering, coming closer, etc. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Therefore, they are either
limited by the variety of perception attributes, sensors or restricted to a particular
perception-action scenario loop. On the other hand, various projects on Human
Robot Interaction try to overcome perception limitations by dierent means and
focus on high-level semantic and decision-making. Such as, the detection of objects
is simplified by putting tags/markers on the objects, in the detection of people no
audio information is used, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], etc. In [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], dierent layers of perception have
been analyzed to build representations of the 3D space, but focused on eye-hand
coordination for active perception and not on high-level semantics and perception
of the human.
      </p>
      <p>
        In the Romeo2 project, we are making eort to bring a range of multi-sensor
perception components within a unified framework (Naoqi, [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]), at the same time
making the entire multi-modal perception system independent from a very specific
scenario or task, and explicitly incorporating reasoning about human, towards
realizing eective and more natural multi-modal human robot interaction. In this
regard, to the best of our knowledge, Romeo2 project is the first eort of its kind
for a real world companion robot. In this paper, we do not provide the details of
each component. Instead, we give an overview of the entire situation assessment
system in Romeo2 project (sec. 2.1). Interested readers can find the details in
documentation of the system [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and in dedicated publications for individual
components, such as [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], etc. (see the complete
list of publications 1). Further, the combined eort to bring dierent components
together helps us to identify some of the innovation potentials and to develop
them, as discussed in section 3.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Perceiving Situation in Romeo2 Project</title>
      <sec id="sec-2-1">
        <title>2.1 A Generalized Sense-Interact Perception Architecture for HRI</title>
        <p>We have adapted a simple yet meaningful, sensing-interaction oriented perception
architecture, by carefully identifying various requirements and their
interdependencies, as shown in fig. 3. The roles of the five identified layers are:
(i) Sense: To receive signals/data from various sensors. Depending upon the
sensors and their fusion. This layer can build 3D point cloud world; sense stimuli
like touch, sound; know about the robot’s internal states such as joint, heat; record
speech signals; etc. Therefore, it belongs to level 1 of situation assessment.</p>
        <p>(ii) Cognize: Corresponds to the ’meaningful’ (human-understandable level)
and relevant information extraction, e.g. learning shapes of objects; learning to
extract the semantics from 3D point cloud, the meaningful words from speech, the
meaningful parameters in demonstration, etc. In most of the perception-action
systems, this cognize part is provided a priori to the system. However, in Romeo2
projects we are taking steps to make cognize layer more visible by bringing together
dierent learning modules, such as to learn objects, learn faces, learn the meaning
of instructions, learn to categorize emotions, etc. This layer lies across level 1 and
level 2 of situation assessment, as it is building knowledge in terms of attributes
and their values and also extracting some meaning for future use and interaction.</p>
        <p>(iii) Recognize: Dedicated to recognizing what has been ’cognized’ earlier by
the system, e.g. a place, face, word, meaning, emotion, etc. This mostly belongs
to level 2 of situation assessment, as it is more on utilizing the knowledge either
learned or provided a priori, hence ’experience’ becomes the dominating factor.
(iv) Track: This layer corresponds to the requirement to track something
(sound, object, person, etc.) during the course of interaction. From this layer, level
3 of situation assessment begins, as tracking allows to update in time the state of
the beforehand entity (person, object, etc.), hence involves a kind of ’projection’.</p>
        <p>(v) Interact: This corresponds to the high-level perception requirements for
interaction with the human and the environment. E.g. activity, action and
intention prediction, perspective taking, social signal and gaze analyses, semantic and
aordance prediction (e.g. pushable objects, sitable objects, etc.). It mainly
belongs to level 3 of situation assessment, as involves ’predicting’ side of perception.</p>
        <p>Sometimes, practically there are some intermediate loops and bridges among
these layers, for example a kind of loop between tracking and recognition. Those
are not shown for the sake of making main idea of the architecture better visible.</p>
        <p>Note the closed loop aspect of the architecture from interaction to sense. As
shown in some preliminary examples in section 3, such as Ex1, we are able to
practically achieve this, which is important to facilitate natural human-robot interaction
process, which can be viewed as: Sense ae Build knowledge for interaction ae
Interact ae Decide what to sense ae Sense ae ...
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Basic Requirements, Key Attributes and Developments</title>
        <p>
          In Romeo2 project, we have identified the key attributes and elements of situation
assessment, to be perceived from companion robotics domain perspective, and
categorized along five basic requirements as summarized in table 1. In this section,
we describe some of those modules. See Naoqi [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] for details of all the modules.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>I. Perception of Human</title>
        <p>People presence: Perceives presence of people, assign unique ID to each detected
person. Face characteristics: To predict age, gender and degree of smile on a
detected face. Posture characterization (human): To find position and
orientation of dierent body parts of the human, shoulder, hand, etc. Perspective
taking: To perceive reachable and visible places and objects from the human’s
perspective, with the level of eort required to see and reach. Emotion
recognition: For basic emotions of anxiety, anger, sadness, joy, etc. based on multi-modal
audio-video signal analysis. Speaker localization: Localizes spatially the
speaking person. Speech rhythm analysis: Analyzing the characterization of speech
rhythm by using acoustic or prosodic anchoring, to extract social signals such as
engagement, etc. User profile: To generate emotional and interactional profile
of the interacting user. Used to dynamically interpret the emotional behavior as
well as to build behavioral model of the individual over a longer period of time.
Intention analysis: To interpret the intention and desire of the user through
conversation in order to provide context, and switch among dierent topics to talk.
The context also helps other perception components about what to perceive and
where to focus. Thus, facilitates closing the interaction-sense loop of fig. 3.</p>
      </sec>
      <sec id="sec-2-4">
        <title>II. Perception of Robot Itself</title>
        <p>Fall detection: To detect if the robot is falling and to take some human user and
self-protection measures with its arms before touching the ground.
Other modules in this category are self-descriptive. However it is worth to mention
that, such modules also provide symbolic level information, such as battery nearly
empty, getting charged, foot touching ground, symbolic posture sitting, standing,
standing in init pose, etc. All these help in achieving one of the aims of Romeo2
project: sensing for natural interaction with human.</p>
      </sec>
      <sec id="sec-2-5">
        <title>III. Perception of Object</title>
        <p>Object Tracker: It consists of dierent aspects of tracking, such as moving to
track, tracking a moving object and tracking while the robot is moving. Semantic
perception (object): Extracts high-level meaningful information, such as object
type (chair, table, etc.), categories and aordances (sitable, pushable, etc.)</p>
      </sec>
      <sec id="sec-2-6">
        <title>IV. Perception of Environment</title>
        <p>Darkness detection: Estimates based on the lighting conditions of the
environment around the robot. Semantic perception (place): Extracts meaningful
information from the environment about places and landmarks (a kitchen, corridor,
etc.), and builds topological maps.</p>
      </sec>
      <sec id="sec-2-7">
        <title>V. Perception of Stimuli</title>
        <p>Contact observer: To be aware of desired or non-desired contacts when they
occur, by interpreting information from various embedded sensors, such as
accelerometers, gyro, inclinometers, joints, IMU and motor torques’.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion on Innovation Potentials</title>
      <p>
        We will not go in detail of the individual modules and the results, as those can be
found online [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Instead, we will discuss some of the advantages and innovation
potentials, which such modules functioning on a unified platform could bring.
      </p>
      <p>Ex1: The
capability of multi-modal
perception, combining
input from the
interacting user, the events
triggered by other
perception components,
and the centralized
memorization
mechanism of robot, help
to achieve the goal of Fig. 4. Subset of interaction topics (right), and their dynamic
closing the interact- activation levels based on multi-modal perception and events.
sense loop and dynamically shaping the interaction.
(a) (b) (c)
Fig. 5. High-level situation assessment. (a) The semantics map of the environment. (b)
Eort and Perspective taking based situation assessment. (c) Combining (a) and (b), the
robot will be able to make the object accessible to the human.</p>
      <p>To demonstrate, we programmed an extensive dialogue with 26 topics that
shows the capabilities of the Romeo robot. During this dialogue the user often
interrupts Romeo to quickly ask a question, this leads to several ’conflicting’ topics
in the dialogue manager. The activation of dierent topics during an interaction
over a period is shown in fig. 4. The plot shows that around 136th second the user
has to take his medicine, but the situation assessment based memory indicates that
the user has ignored and not yet taken the medicine. Eventually, the system results
the robot urging the user to take his medication (pointed by blue arrow), making
it more important than the activity indicated by the user during the conversation
(to engage in reading a book, pointed by dotted arrow in dark green). Hence, a
close loop between the perception and interaction is getting achieved in a real
time, dynamic and interactive manner.</p>
      <p>Ex2: Fig. 5(a) shows situation assessment of the environment and objects at
the level of semantics and aordances, such as there is a ’table’ recognized at
position X, and this belongs to an aordance category on which something can
be put. Fig. 5(b) shows situation assessment by perspective taking, in terms of
abilities and eort of the human. This enables the robot to infer that the sitting
human (as shown in fig. 5(c)) will be required to stand up and lean forward to see
and take the object behind the box. Thanks to the combined reasoning of (a) and
(b), the robot will be able to make the object accessible to the human by placing it
on the table (knowing that something can be put on it), at a place reachable and
visible by the human with least eort (through the perspective taking mechanism),
as shown in fig. 5(c).</p>
      <p>
        In Romeo2 we also aim to use this combined reasoning about abilities and
eorts of agents, and aordances of the environment, for autonomous
humanlevel understanding of task semantics through interactive demonstration, for the
development of robot’s proactive behaviors, etc. as suggested the feasibility and
advantages in some of our complementary studies in those directions, [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        Ex3: Analyzing verbal and non-verbal
behaviors such as head direction (e.g. on-view or
oview detection) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], speech rhythm (e.g. on-talk
or self-talk) [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], laugh detection [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], emotion
detection [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], attention detection [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], and their
dynamics (e.g. synchrony [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]), combined with acoustic
analysis (e.g. spectrum) and prosodic analysis
altogether greatly allows to improve social engagement Fig. 6. Self-talk detection
characterization of the human during interaction.
      </p>
      <p>To demonstrate, we collected a database
of human-robot interaction during sessions
of cognitive stimulation. The preliminary
result with 14 users shows that on a 7 level
evaluation scheme, the average scores for
questions, ”Did robot show any empathy?”,
”Was it nice to you?” and ”Was it polite?”
were 6.3, 6.2 and 6.4 respectively. In
addition, the multi-modality combination of Fig. 7. Face, shoulder and face
orientathe rhythmic, energy and pitch character- tion detection of two interacting people.
istics seems to be elevating the detection of
self-talk (known to reflect the cognitive load of the user, especially for elderly) as
shown in table of fig. 6.</p>
      <p>Ex4: Inferring face gaze (as illustrated in fig.
7), combined with sound localization and object
detection, altogether provides enhanced
knowledge about who might be speaking in a
multipeople human-robot interaction, and further
facilitates analyzing the attention and intention.</p>
      <p>To demonstrate this, we conducted an
experiment with two speakers, initially speaking at the
dierent sides of the robot and then slowly
moving towards each other and eventually separate
away. Fig. 8 shows the preliminary result for the
sound source separation by the system based on oFnigly. 8a.uSdoiounbdassoedurc(eBsFe-pSaSr)atiaonnd,
beamforming. The left part (BF-SS) shows when audio-video based (AVBF-SS).
only the audio signal is used. When the system
uses the visual information combined with the audio signals, the performance is
better (AVBF-SS) in all the three types of analyses: signal-to-interference ratio
(SIR), signal-to-distortion ratio(SDR) and signal-to-artifact (SAR) ratio.</p>
      <p>
        Ex5: The fusion of rich information about visual clues, audio speech rhythm,
lexical content and the user profile is also opening doors for automated context
extraction, helping for better interaction and emotion grounding and making the
interaction interesting, like doing humor [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4 Conclusion and Future Work</title>
      <p>In this paper, we have provided an overview of the rich multi-modal perception
and situation assessment system within the scope of Romeo2 project. We have
presented our sensing-interaction perception architecture and identified the key
perception components requirements for companion robot. The main novelty lies
in the provision for rich reasoning about the human and practically closing the
sensing-interaction loop. We have pointed towards some of the work in progress
innovation potentials, achievable when dierent situation assessment components
are working on a unified theoretical and functional framework. It would be
interesting to see how it could serve as guideline in dierent context than companion
robot, such as robot co-worker.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Risager</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andersen</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravn</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Spacio-temporal situation assessment for mobile robots</article-title>
          .
          <source>In: Int. Conf. on Information Fusion (FUSION)</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bolstad</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          :
          <article-title>Situation awareness: Does it change with age</article-title>
          . vol.
          <volume>45</volume>
          , pp.
          <fpage>272</fpage>
          -
          <lpage>276</lpage>
          . Human Factors and Ergonomics
          <string-name>
            <surname>Society</surname>
          </string-name>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Buendia</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Devillers</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>From informative cooperative dialogues to long-term social relation with a robot</article-title>
          .
          <source>In: Natural Interaction with Robots</source>
          ,
          <source>Knowbots and Smartphones</source>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>151</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Caron</surname>
            ,
            <given-names>L.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Filliat</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gepperth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Neural network based 2d/3d fusion for robotic object recognition</article-title>
          .
          <source>In: Proc. European Symposium on Artificial Neural Networks, Computational Intelligence and Machine Learning (ESANN)</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A robot architecture based on higher order perception loop</article-title>
          .
          <source>In: Brain Inspired Cognitive Systems</source>
          <year>2008</year>
          , pp.
          <fpage>267</fpage>
          -
          <lpage>283</lpage>
          . Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>CHRIS-Project</surname>
          </string-name>
          :
          <article-title>Cooperative human robot interaction systems</article-title>
          . http://www.chrisfp7.eu/
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Delaherche</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chetouani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mahdhaoui</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saint-Georges</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viaux</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Interpersonal synchrony: A survey of evaluation methods across disciplines</article-title>
          .
          <source>Aective Computing, IEEE Transactions on 3(3)</source>
          ,
          <fpage>349</fpage>
          -
          <lpage>365</lpage>
          (
          <year>July 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Devillers</surname>
            ,
            <given-names>L.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soury</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A social interaction system for studying humor with the robot nao</article-title>
          .
          <source>In: ICMI</source>
          . pp.
          <fpage>313</fpage>
          -
          <lpage>314</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Endsley</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          :
          <article-title>Toward a theory of situation awareness in dynamic systems</article-title>
          .
          <source>Human Factors: Journal of the Human Factors and Ergonomics Society</source>
          <volume>37</volume>
          (
          <issue>1</issue>
          ),
          <fpage>32</fpage>
          -
          <lpage>64</lpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>EYESHOTS-Project</surname>
          </string-name>
          :
          <article-title>Heterogeneous 3-d perception across visual fragments</article-title>
          . http://www.eyeshots.it/
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Filliat</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Battesti</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bazeille</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duceux</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gepperth</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harrath</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jebari</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tapus</surname>
            , A., Meyer,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ieng</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benosman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cizeron</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mamanna</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pothier</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Rgbd object recognition and visual texture classification for indoor semantic mapping</article-title>
          .
          <source>In: Technologies for Practical Robot Applications</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Philippsen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siegwart</surname>
          </string-name>
          , R.:
          <article-title>Narrative situation assessment for humanrobot interaction</article-title>
          .
          <source>In: IEEE ICRA</source>
          . vol.
          <volume>1</volume>
          , pp.
          <fpage>1503</fpage>
          -
          <lpage>1508</lpage>
          vol.
          <volume>1</volume>
          (
          <issue>Sept</issue>
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>JOKER-Project</surname>
          </string-name>
          :
          <article-title>Joke and empathy of a robot/eca: Towards social and aective relations with a robot</article-title>
          . http://www.chistera.eu/projects/joker
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lallee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lemaignan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lenz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melhuish</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natale</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Skachek</surname>
            , S., van Der Zant,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Warneken</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dominey</surname>
            ,
            <given-names>P.F.</given-names>
          </string-name>
          :
          <article-title>Towards a platform-independent cooperative human-robot interaction system: I. perception</article-title>
          . In: IEEE/RSJ IROS. pp.
          <fpage>4444</fpage>
          -
          <lpage>4451</lpage>
          (
          <year>Oct 2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Le Maitre</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chetouani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Self-talk discrimination in human-robot interaction situations for supporting social awareness</article-title>
          .
          <source>J. of Social Robotics</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <fpage>277</fpage>
          -
          <lpage>289</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baek</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Cognitive robotic engine: Behavioral perception architecture for human-robot interaction</article-title>
          .
          <source>In: Human Robot Interaction</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Mekonnen</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lerasle</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herbulot</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Briand</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>People detection with heterogeneous features and explicit optimization on computation time</article-title>
          .
          <source>In: ICPR</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>18. NAOqi-Documentation: https://community.aldebaran-robotics.com/doc/2- 00/naoqi/index.html/</mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Pandey</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alami</surname>
          </string-name>
          , R.:
          <article-title>Towards human-level semantics understanding of humancentered object manipulation tasks for hri: Reasoning about eect, ability, eort and perspective taking</article-title>
          .
          <source>Int. J. of Social</source>
          Robotics pp.
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Pandey</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ali</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alami</surname>
          </string-name>
          , R.:
          <article-title>Towards a task-aware proactive sociable robot based on multi-state perspective-taking</article-title>
          .
          <source>J. of Social Robotics</source>
          <volume>5</volume>
          (
          <issue>2</issue>
          ),
          <fpage>215</fpage>
          -
          <lpage>236</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Pomerleau</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          :
          <article-title>Neural network perception for mobile robot guidance</article-title>
          .
          <source>Tech. rep., DTIC Document</source>
          (
          <year>1992</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Ringeval</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chetouani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuller</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Novel metrics of speech rhythm for the assessment of emotion</article-title>
          . Interspeech pp.
          <fpage>2763</fpage>
          -
          <lpage>2766</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Sehili</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Devillers</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Attention detection in elderly people-robot spoken interaction</article-title>
          .
          <source>In: ICMI WS on Multimodal Multiparty real-world HRI</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Tahon</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delaborde</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Devillers</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Real-life emotion detection from speech in human-robot interaction: Experiments across diverse corpora with child and adult voices</article-title>
          .
          <source>In: INTERSPEECH</source>
          . pp.
          <fpage>3121</fpage>
          -
          <lpage>3124</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>