<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Search and retrieval of audiovisual content by integrating non-verbal multimodal, affective, and social descriptors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Antonio Camurri</string-name>
          <email>antonio.camurri@unige.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Casa Paganini - InfoMus Intl Research Centre DIST- University of Genova Piazza Santa Maria in Passione 34</institution>
          ,
          <addr-line>16123 Genova</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>One of the research challenges for future search engines concerns the integration of multimodal and cross-modal, nonverbal, full-body, affective, social, and enactive interaction in the process of search and retrieval of audiovisual content. The paper gives a short presentation of the three-year EU project I-SEARCH (EU 7FP ICT STREP), aiming at creating a novel unified framework for multimodal and cross-modal content indexing, sharing, search and retrieval of audiovisual content. A couple of scenarios developing multimodal paradigms of search and retirieval of audiovisual content are introduced and briefly discussed to explain in concrete terms some of the main research challenges that are addressed in I-SEARCH. Finally, the paper presents preliminary results on a specific research challenge: analysis of nonverbal expressive and social behaviour to extract useful information from users for the retrieval of audiovisual content.</p>
      </abstract>
      <kwd-group>
        <kwd>non-verbal full-body multimodal interfaces</kwd>
        <kwd>descriptors</kwd>
        <kwd>emotion</kwd>
        <kwd>social signals</kwd>
        <kwd>sound and music computing</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Internet is quickly evolving towards providing richer and immersive experiences, in
which the user interact seamlessly and transparently with digital and physical
artefacts. Due to the widespread availability of digital recording devices, improved
modelling tools, advanced scanning mechanisms as well as display and rendering
devices, even on mobile environments, users are more and more empowered to have a
more immersive and interactive experience. Digital media are moving to “User
Centric Media” [
        <xref ref-type="bibr" rid="ref2 ref3">2,3</xref>
        ], enabling adaptive and active experiences of audiovisual content
(see for example the EU ICT SAME Project www.sameproject.org). Users become
“prosumers”, and are more and more participating in the updating process of the
information and in improving the resolution and richness of the media repositories.
The emergence of embodied and social interaction with content, enabled by the
dramatic advances of multimodal/intelligent/natural interfaces, enriches this scenario,
providing users with further degrees of freedom and channels to access the content, in
terms of full-body, non-verbal, expressive [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], social [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] interaction with content.
It is therefore now possible for users to rapidly move from a mainly textual-based to a
media-based “embodied” Internet, where rich audiovisual content (images, graphics,
sound, videos, 3D models, etc.), 3D representations (avatars), virtual and mirror
worlds, serious games, lifelogging applications, multimodal yet affective utterances
(gestures, facial expressions, eye movements,…) etc. become a reality. See for
example [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for an extensive survey on content-based multimedia information
retrieval, and the “white papers” from EU on Future internet and on User-Centric
Media [
        <xref ref-type="bibr" rid="ref2 ref3">2,3</xref>
        ] for an in depth analysis of future internet and emerging user-centric
media.
      </p>
      <p>Traditional search of music archives usually include descriptive methods (mainly
textual, e.g., author, title, etc.). Recent search engines also include alternative
querying modalities such as audio (Shazam, Google China Music) or image (Google
goggles, Google similar images). However, these search engines are suited for
queryby-content or query-by-example, where the research objective is clearly defined or
where a specific information is targeted.</p>
      <p>We aim at developing alternative, yet complementary, querying modalities, e.g.,
integrating expressive gesture and affective, emotional cues, that facilitate more
explorative and creative search.</p>
      <p>This paper presents some insights and preliminary research results on the problem of
search and retrieval in cases where textual information is either missing or it is not
sufficient or adequate, and therefore the need for the integration of non-verbal
multimodal, cross-modal, full-body, affective, social descriptors emerges.
Use case scenarios on music content search and retrieval are adopted in this paper to
discuss main research challenges and approaches.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The I-SEARCH EU Project</title>
      <p>The EU 7FP ICT STREP I-SEARCH project aims to create a novel unified
framework for multimedia and multimodal content indexing, sharing, search and
retrieval. I-SEARCH is coordinated by Dimitrios Tsovaras by ITI-CERTH (Centre for
Research and Technology Hellas – Informatics and Telematics Institute), and partners
include JCP-Consult, INRIA, Athens Technology Center, Engineering Ingegneria
Informatica S.p.A., Google, University of Genoa, Exalead, Erfurt University of
Applied Sciences, Accademia Nazionale di Santa Cecilia, EasternGraphics. The
project started in January 2010, with a duration of three years.
“I-SEARCH aims to create a novel unified framework for multimedia and multimodal
content indexing, sharing, search and retrieval. I-SEARCH aims to be the first search
engine able to handle specific types of multimedia (text, 2D image, sketch, video, 3D
objects, audio and combination of the above) and multimodal content (gestures, face
expressions, eye movements) along with real world information (GPS, temperature,
time, weather sensors, RFID objects,), which can be used as queries and retrieve any
available relevant content of any of the aforementioned types and from any end-user
access device. Towards this aim, I-SEARCH proposes the research and development
of an innovative Rich Unified Content Description (RUCoD). RUCoD will consist of
a multi-layered structure, which will integrate descriptors of all of the above types of
content, real-world information, even non-verbal yet implicit, emotional cues and
social descriptors, in order to better express what the user wants to retrieve. Another
objective of I-SEARCH is the development of intelligent content interaction
mechanisms, including personal, social-based and recommendation-based relevance
feedback and novel interoperable multimodal interfaces. This will result in a highly
user-centric search engine, able to deliver to the end-users only the content of interest,
satisfying their information needs and preferences, which is expected to dramatically
improve end-user experience and offer new market opportunities. Furthermore,
ISEARCH introduces the use of advanced visual analytic technologies for search
results presentation in order to facilitate their fast and easy interpretation and also to
support optimal results presentation under various contexts (i.e. user profile, end-user
terminal, available network bandwidth, interaction modality preference, etc.). Finally,
the search engine will be dynamically adapted to end-user's device, which will vary
from a simple mobile phone to a high-performance PC.” (from cordis.europe.eu
projects archive)</p>
    </sec>
    <sec id="sec-3">
      <title>3 Scenarios</title>
    </sec>
    <sec id="sec-4">
      <title>3.1 Music retrieval through expressive embodied queries</title>
      <p>Chiara is a music-lover, looking for music material that share common affective
features. Here, search aims at discovering unexpected filiations and similarities across
music artworks. It takes place in an environment equipped with devices enabling the
user to express herself through voice, hands and body gesture. Pre-recorded
multimedia content can be uploaded from an external device (e.g., a mobile phone).
For each digital content in the collection, descriptors related to low-level features,
real-world context data, and expressive/emotional/social cues that compose the
RUCoD (Rich Unified Content Description of the I-SEARCH project) standard are
stored.</p>
      <p>Inputs to the search module include text queries, audio capture of live user singing
incipit of music piece, audio file recorded on handheld device such as mobile phones.
Beat tracking can be captured by tapping on a microphone or through accelerometers
embedded in mobile devices. User gestures can be captured using either video camera
or accelerometers embedded in user‟s mobile devices.</p>
      <p>
        Chiara wants to explore music artworks that share affective features with the Ravel‟s
Bolero. She starts by using the I-SEARCH framework to retrieve audio information
that share similarities with this audio pattern. Using a tangible acoustic interface [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ],
she taps the beating of the rhythm, a constant 4/4 time with a prominent triplet on the
second beat of every bar, or smaller rhythmic cells (for example the beat and triplet).
The recorded audio fragment is used by the I-SEARCH engine as an audio query to
initialize the search. Specifically, the I-SEARCH framework extracts low-level
descriptors from the audio content (the tapping resulting audio) and create a query
based on the RuCoD format. Through template matching techniques [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], similar audio
results related to Bolero rhythmic pattern are retrieved. Related video files and music
scores are also retrieved through multimodal annotation propagation. Results are
displayed via visual analytics techniques on Chiara‟s terminal, using clusters
annotated with information like modality type, population size and others.
Chiara picks one of the results returned by the query, listens to it but decides that what
she needs is something more energetic, so she closes her fists and starts making sharp,
sudden vertical movements on the same Bolero rhythm. Through a video camera or
embedded accelerometers the environment captures such expressive features of the
gesture, and refine the search, resulting in changes in the displayed results to match it
(either by removing the items that don‟t convey that expression, or by moving the
suitable items closer to build a different cluster configuration). One of the results
captures Chiara‟s attention: a drum recording from Italian ethnomusicological
repertoire, the „Ritmo di tamburo‟, where various drum rhythms are played, sharing
indeed the same triplet of Bolero. A little further away she also finds a voice
recording.
3.2 Collective DJ - Social music retrieval through expressive
embodied queries
Four friends at a party wish to dance together, and to accomplish this they search
some music pieces resonating with their (collective) mood. They do not know in
advance the music pieces they want, and they use the I-SEARCH tool collaboratively
to find their music, and possible associated video. Alternatively, the search process
might not necessarily be „conscious‟ or intentional, but simply a part of a social game,
i.e., in a more fun/entertainment approach.
      </p>
      <p>The friends are at a party, but not necessarily they share the same physical
environment. One or more of them can be remotely connected via audiovisual links.
GPS and context aware information are available to the ISEARCH search engine.
Users have devices that enable to express themselves through voice, movement, face,
and full-body movements.</p>
      <p>For each digital music content of the archive under consideration, descriptors related
to low-level features, real-world data and expressive/emotional/social cues that
compose the RuCoD (Rich Unified Content Description) standard are available.
Users inputs include the following: (i) Rhythmic queries, using hands, clapping,
fullbody movement; (ii) Context data (GPS, compass, proximity of others etc); (iii)
Entrainment/synchronisation and dominance/leadership among users, measured by
on-body sensors (eg accelerometers or/and videocameras on their mobiles, or game
interfaces) and/or environment videocamera(s), to find a shared, collective
information to build the query based on non-verbal social signals; (iv) Gestures to
shape the query: again, the gesture are captured using either video camera and/or
accelerometers embedded in mobile devices (carried by the user or kept in hand) or in
environmental videocameras.</p>
      <p>The output of the experience include a shared enjoyment of performances of the
retrieved music pieces, possible video clips associated to such music pieces (music
videos).</p>
      <p>The experience may be described as follows:
1. The four users A, B,C, D start to dance;
2. Their movement acts as selector of music pieces coded in relation to their
motoric-affective-social behaviour (slow, fast, dionisiac, ...);</p>
      <p>3. If the movements of A,B,C,D are not sympathetic (i.e., low entrainment), an
overlapping of different music pieces will emerge, in a rather chaotic sound
environment (the different music pieces are heard simultaneously at different
changing levels according to users behaviour). As the joint experience goes on, the
music continuously changes, and may start to converge to a piece corresponding to
the user who results to dominate in the group (dominance/leadership features which
are automatically extracted by the system);</p>
      <p>4. When the movement of the group obtains a sufficient uniformity (contagion
from the one who results to act as a leader), the group will converge to dance on the
same music, chosen by such “collective gesture”: this is the first level of query result;
5. This entertaining task is integrated and is part of the search task: for example,
the search can take the priority as soon as one of the users, once obtained a shared
agreement and a single music, tries to trigger a change in the shared general emotion
of the group (with a consequent change of musical choices), for example a
perturbation to the current situation by a user, to explore music pieces similar to the
one obtained and experienced by the group. The effect can be a sort of “game on
leadership”: the user who is able to triggers a general perturbation of emotion in the
group causes a sort of “collective DJ-like” real-time interactive editing of
heterogeneous music fragments on the main music piece. The user who is able to act
as a leader (detected by the system), has the possibility to inject the consequences of
her gestural/movement/affective choices, which will take the power to determine the
change of the music context only in relation to the capability of her contagion on the
other users. If the others will be captured, they will follow her moving toward a new
music piece, by means of an audio and possibly video cross-fade or sudden change of
scene. Otherwise, if the user will not have enough power or dominance on the others,
her associated new music piece will fade off.</p>
      <p>
        This is a sort of a “Collective DJ” example, in which the collective behaviour
provides the source data for the search of the music. It is the opposite of the
traditional music experience: here the movement determines its own correct music
frame [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The music retrieval may also keep into account of GPS and spatial locations of users.
For example, GPS information may be kept into account to select music pieces
keeping into account their geographical region.
3.3</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>Social search can occur simultaneously or in different moments. In the first case,
which is considered in the scenario 2, users collaborate to a common shared objective.
In the latter, a user leaves a track of her activity, which can be used later by other
users to take inspiration for their search. Of course, mechanisms allowing the
intentional control of access to personal search schemas and data, similarly to shared
data in social networks, must be introduced.</p>
      <p>In situations of “affective non-verbal querying” the following three types of
difficulties emerge:
(i) The user is not familiar to non-verbal search by means of “affect”, “expressivity”,
“embodiment”: she is not sure if she is able to perform correctly the task, she
does not know if she is able to link an affect with a content. For example, if she
sings/hums a song but she has poor singing skills, this can be a case in which the
user is not able to express the affect she wants to convey (but associated gesture
may help). In a “game-like” scenario, the users may be more inclined to dancing
even if they are not professional dancers, for the sake of fun and social
interaction. In this case, the results of the scenario should be measured more in
terms of user experience than on the technical „efficiency‟ and performance of
the search task.
(ii) The user may be not capable to express what she is searching for;
(iii) The machine might not be able to execute the search according the user
intentions.</p>
      <p>Despite these difficulties, emotional/expressive social non-verbal queries, with their
inner “blurred” characterisation, can enhance “serendipitous discovery”, and can lead
to stimulating exploring-style querying, complementary to traditional querying
paradigms.
4. Multimodal queries based on non-verbal expressive and social
features
Several research challenges emerge from the scenarios sketched above, including the
understanding of users multimodal inputs, and in particular the non-verbal multimodal
cues conveying users’ intentions, and the mapping between user multimodal inputs to
features in the audiovisual content.</p>
      <p>In this section we focus on how to exploit users’ non-verbal expressive and social
signals useful to build multimodal queries.</p>
      <p>
        The proposed approach consists of two phases:
(i) Extraction of an array of expressive features describing each user behavior
[
        <xref ref-type="bibr" rid="ref11">12</xref>
        ];
(ii) Using such expressive features as the inputs to modules which extract social
features related users behavior, with particular focus on entrainment and
dominance [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">13, 14, 15</xref>
        ].
      </p>
      <p>The resulting array of individual and social features will be a subset of the user
component RuCoD descriptors.
4.1</p>
    </sec>
    <sec id="sec-6">
      <title>Expressive features</title>
      <p>
        As for the extraction of descriptors on non-verbal expressive behaviour, a particular
focus is on full-body movement and gesture, i.e., on recognizing how a gesture is
performed, including expressive and affective content [
        <xref ref-type="bibr" rid="ref11">12</xref>
        ]. Typically, a single gesture
can be performed in several different ways (e.g., fluid, hesitant, impulsive). A
collection of features characterizing the expressive qualities of a gesture has been
defined, starting from biomechanics, psychology, and humanistic theories [
        <xref ref-type="bibr" rid="ref11">12</xref>
        ].
Expressive features include the following:
      </p>
      <p>
        Quantity of Motion (QoM) is an index of motoric activation that provides an
estimation of the amount of overall movement (variation of pixels) the video-camera
detects. QoM computed on translational movements only (TQoM) provides an
estimation of how much the user is moving around the physical space. Using Laban’s
Effort [
        <xref ref-type="bibr" rid="ref16">17</xref>
        ] terminology, whereas Quantity of Motion measures the amount of
detected movement in both the Kinesphere and the General Space, its computation on
translational movements refers to the overall detected movement in the General Space
only. TQoM, together with speed of barycentre (BS) and variation of the Contraction
Index (dCI) are introduced to distinguish between the movement of the body in the
General Space and the movement of the limbs in the Kinesphere. Intuitively, if the
user moves her limbs but does not change her position in the space, TQoM and BS
will have low values, while QoM and dCI will have higher values.
      </p>
      <p>Impulsiveness (IM) is extracted using a model measuring it as a combination of
other features, mainly derived from QoM. The first one is the variance of QoM in a
sliding time window of 3s, i.e., a user is considered to move in an impulsive way if
the amount of movement the video-camera detects on her body changes considerably
in the time window. A second group of features is related to the analysis of the shape
of QoM along time. Such features include e.g., the ratio between the main peak of
QoM in the time window and the time duration of such a peak, the steepness of the
attack of a movement phase, the steepness of the main peak, the number of peaks
detected in the time window, the ratio between the main peak and the second biggest
one, the distribution of the peaks in the time window (i.e., whether they are uniformly
distributed along the time window or concentrated over a specific time range).
Another feature is related to the content of the QoM spectrum in the frequency band
over 5 Hz. Each feature is then weighted and combined in the model in order to
provide an overall index of Impulsiveness.</p>
      <p>
        Vertical and horizontal components of velocity of peripheral upper parts of the
body (VV, HV) are computed starting from the positions of the upper vertexes of the
body bounding rectangle. The vertical component, in particular, is used for detecting
upward movements that psychologists (e.g., [
        <xref ref-type="bibr" rid="ref15">16</xref>
        ]) identified as a significant indicator
of positive emotional expression.
      </p>
      <p>Space Occupation Area (SOA) is computed starting from the movement trajectory
integrated over time. In such a way a bitmap is obtained, summarizing the trajectory
followed along the considered time window (3s). An elliptical approximation of the
shape of the trajectory is then computed. The area of such ellipse is taken as the Space
Occupation Area. Intuitively, a trajectory spread over the whole space gets high SOA
values, whereas a trajectory confined in a small region gets low SOA values.</p>
      <p>Directness Index (DI) is computed as the ratio between the length of the straight
line connecting the first and last point of a trajectory (in this case the movement
trajectory in the selected 3s time window) and the sum of the lengths of each segment
composing the trajectory. It is inspired by the Space dimension of Laban’s Effort
Theory.</p>
      <p>Space Allure (SA) measures local deviations from the straight line trajectory. It is
inspired by composer Pierre Schaeffer’s Morphology. Whereas DI provides
information about whether the trajectory followed along the 3s time window is direct
or flexible, SA refers to waving movements around the straight trajectory in shorter
time windows. Currently, SA is approximated with the variance of DI in a time
window of 1s.</p>
      <p>
        The Amount of Periodic Movement (PM) provides a preliminary information
about the presence of rhythmic movements. Computation of PM starts from QoM.
Movement is segmented in motion and pause phases using an adaptive threshold on
QoM [
        <xref ref-type="bibr" rid="ref11">12</xref>
        ]; inter-onset intervals are then computed as the time elapsing from the
beginning of a motion phase and the beginning of the following motion phase. The
variance of such inter-onset intervals is taken as an approximate measure of PM.
      </p>
      <p>Symmetry Index (SI) is computed from the position of the barycenter and the left
and right edges of the body bounding rectangle. That is, it is the ratio between the
difference of the distances of the barycenter from the left and right edges and the
width of the bounding rectangle:
where xB is the x coordinate of the barycentre, xL is the x coordinate of the left edge of
the body bounding rectangle and xR is the x coordinate of the right edge.
The expressive features, can be analyzed using video input from videocameras, but
other sensor inputs can be considered, e.g. the 3D accelerometers embedded in mobile
systems.</p>
      <p>Current work aims at refining and extending the set of expressive features, to
contribute to the RuCoD standard.</p>
      <p>The expressive features are implemented as real-time software modules in the open
software platform EyexWeb XMI (www.eyesweb.org).
4.2</p>
    </sec>
    <sec id="sec-7">
      <title>Social features</title>
      <p>
        Research on the analysis of social descriptors include the development of models and
techniques for measuring entrainment, empathy, dominance, leadership, and salient
behaviour in small groups of users. We obtained preliminary results on entrainment
and dominance, based on theories of synchronization [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">13,14,15</xref>
        ].
      </p>
      <p>A basic assumption to approach nonverbal social behavior consists of the modeling of
a small group of users as a complex system consisting of single interacting
components able to auto-organize and to show global properties, which are not
obvious from the observation of their individual dynamics.</p>
      <p>A number of algorithms and related software modules for the automated analysis in
real-time of non-verbal cues related to expressive gesture in social interaction are
currently studied at our centre. Analysis of entrainment and dominance is based on
Phase Synchronisation and Recurrence Quantification Analysis. We start from the
hypothesis that phase synchronisation is one of the low-level social signals explaining
empathy and dominance in a small group of users. Another direction, based on Multi
Scale Entropy and other approaches is currently adopted to measure saliency and
rarity index in small group of users. Real time implementation of the algorithms
developed so far is available in the EyesWeb XMI Social Signal Processing Library
(www.eyesweb.org).
5</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <p>Some of the main research challenges faced in the EU 7FP ICT I-SEARCH project,
and preliminary results in the analysis of users behavior in terms of expressive and
social features have been presented, as a contribute to the RuCoD standard in
ISEARCH. Current work includes research on empathy, emotional entrainment,
leadership, co-creation, and attention.</p>
      <p>
        Other important directions of the research in I-SEARCH concern the study of
descriptors in audiovisual content, and the study of cross-modal descriptors [
        <xref ref-type="bibr" rid="ref10">11</xref>
        ].
      </p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>I am deeply grateful to my colleagues and friends Corrado Canepa, Paolo Coletta,
Nicola Ferrari, Alberto Massari, Gualtiero Volpe, Donald Glowinski, Maurizio
Mancini, Giovanna Varni.</p>
      <p>This research is partially supported by the 7FP EU-ICT three-year project I-SEARCH
no. 248296.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Lew</surname>
          </string-name>
          , Michael S.,
          <string-name>
            <surname>Sebe</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Djeraba</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Jain</surname>
          </string-name>
          , R..
          <article-title>Content-based multimedia information retrieval: State of the art and challenges</article-title>
          .
          <source>ACM Trans. Multimedia Comput. Commun. Appl.</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          , (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Laso</given-names>
            <surname>Ballestreros</surname>
          </string-name>
          , I. (Ed.) Research on Future Media Internet,
          <source>Future Media Internet Task Force, European Commission, 7FP ICT Networked Media Unit</source>
          ,
          <year>January 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Laso</given-names>
            <surname>Ballestreros</surname>
          </string-name>
          , I. (Ed.)
          <article-title>User Centric Media in the Future Internet, European Commission, 7FP ICT Networked Media Unit</article-title>
          ,
          <year>November 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Casey</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Veltkamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Leman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rhodes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Slaney</surname>
          </string-name>
          .
          <article-title>Content-based music information retrieval: current directions and future challenges</article-title>
          .
          <source>PROCEEDINGSIEEE</source>
          ,
          <volume>96</volume>
          (
          <issue>4</issue>
          ):
          <fpage>668</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Pentland</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,:
          <article-title>Socially aware, Computation and Communication</article-title>
          . Computer, IEEE CS Press (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Camurri</surname>
            , A.,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Canepa</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Ghisio</surname>
          </string-name>
          , G.Volpe (
          <year>2009</year>
          )
          <article-title>Automatic Classification of Expressive Hand Gestures on Tangible Acoustic Interfaces According to Laban‟s Theory of Effort</article-title>
          . In M.S.Dias,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gibet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.W.</given-names>
            <surname>Wanderley</surname>
          </string-name>
          , R.Bastos (Eds.),
          <source>Gesture-Based Human-Computer Interaction and Simulation</source>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>162</lpage>
          ,
          <issue>LNAI5085</issue>
          , Springer, ISSN
          <volume>0302</volume>
          -9743.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Poli</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volpe</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Toward Communicating Expressiveness and Affect in Multimodal Interactive Systems for Performing Art and Cultural Applications</article-title>
          , IEEE Multimedia, Vol.
          <volume>12</volume>
          , No.
          <issue>1</issue>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>53</lpage>
          , IEEE Computer Society Press (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <source>Interactive Dance/Music Systems, Proc. Intl. Computer Music Conference ICMC-95</source>
          , pp.
          <fpage>245</fpage>
          -
          <lpage>252</lpage>
          ,
          <article-title>The Banff Centre for the arts</article-title>
          ,
          <source>Sept.3-7</source>
          , Canada, ICMAIntl.Comp.Mus.
          <string-name>
            <surname>Association</surname>
          </string-name>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Canepa</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volpe</surname>
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Active listening to a virtual orchestra through an expressive gestural interface: The Orchestra Explorer</article-title>
          .
          <source>Proc. Intl Conf NIME-2007 New Interfaces for Music Expression</source>
          , New York University, (
          <year>2007</year>
          )
          <fpage>10</fpage>
          .
          <string-name>
            <surname>Varni</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coletta</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volpe</surname>
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Emotional Entrainment in Music Performance</article-title>
          .
          <source>Proc. 8th IEEE Intl Conf on Automatic Face and Gesture Recognition</source>
          ,
          <year>Sept</year>
          .
          <fpage>17</fpage>
          -
          <lpage>19</lpage>
          , Amsterdam (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          11.
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Coletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Drioli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Massari</surname>
          </string-name>
          , G.Volpe (
          <year>2005</year>
          ).
          <article-title>Audio processing in a multimodal framework</article-title>
          .
          <source>Proc. Intl. Conf. AES-05 Audio Engineering Society</source>
          , Barcelona, May
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          12.
          <string-name>
            <given-names>A.</given-names>
            <surname>Camurri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mazzarino</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Volpe (2004) Expressive Interfaces</article-title>
          . Cognition Technology &amp;
          <article-title>Work, special issue on "Presence: design and technology challenges for cooperative activities in virtual or remote environments”</article-title>
          , P.Marti (Ed.), Vol.
          <volume>6</volume>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>22</lpage>
          , SpringerVerlag.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          13.
          <string-name>
            <surname>Varni</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mancini</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volpe</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camurri</surname>
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2009</year>
          )
          <article-title>"Sync'n'Move: social interaction based on music and gesture"</article-title>
          .
          <source>In Proceedings of the 1st Intl. ICST Conference on User Centric Media (UCMedia</source>
          <year>2009</year>
          ), Venice, Italy,
          <year>December 2009</year>
          . LNICST vol.
          <volume>40</volume>
          , Springer,
          <year>2010</year>
          .
          <source>(ISBN 978-3-642-12629-1).</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          14.
          <string-name>
            <surname>Camurri</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varni</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volpe</surname>
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2009</year>
          )
          <article-title>Measuring Emotional Entrainment in Small Groups of Musicians</article-title>
          .
          <source>In Proceedings of International Conference on Affective Computing &amp; Intelligent Interaction (ACII</source>
          <year>2009</year>
          ), Lisbon.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          15.
          <string-name>
            <surname>Varni</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Camurri</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coletta</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volpe</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <article-title>(2009) Toward Real-time Automated Measure of Empathy and Dominance</article-title>
          .
          <source>In Proceedings of the 2009 IEEE International Conference on Social Computing SocialCom</source>
          , Vancouver, Canada,
          <year>August 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          16.
          <string-name>
            <surname>Boone</surname>
          </string-name>
          , R. T.,
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          ,
          <article-title>(1998) Children's decoding of emotion in expressive body movement: The development of cue attunement</article-title>
          ,
          <source>Developmental Psychology</source>
          ,
          <volume>34</volume>
          ,
          <fpage>1007</fpage>
          -
          <lpage>1016</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          17.
          <string-name>
            <surname>Laban</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawrence</surname>
            ,
            <given-names>F.C.</given-names>
          </string-name>
          ,
          <year>1947</year>
          . Effort. Macdonald &amp; Evans
          <string-name>
            <surname>Ltd</surname>
          </string-name>
          ., London.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>