<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Audiovisual Corpus of Guided Tours in Cultural Sites: Data Collection protocols in the CHROME Project</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Antonio Origlia</string-name>
          <email>antonio.origlia@unina.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renata Savy</string-name>
          <email>rsavy@unisa.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iolanda Alfano</string-name>
          <email>ialfano@unisa.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesca D'Errico</string-name>
          <email>francesca.derrico@ uniroma3.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Isabella Poggi</string-name>
          <email>isabella.poggi@uniroma3</email>
          <email>isabella.poggi@uniroma3. it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laura Vincze</string-name>
          <email>Laura.vincze@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Cutugno</string-name>
          <email>cutugno@unina.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Violetta Cataldo</string-name>
          <email>violetta.cataldo@live.itt</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical, Engineering and, Information Technology, University of Naples, "Federico II"</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Humanities, Studies, University of</institution>
          ,
          <addr-line>Salerno, Salerno</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Philosophy</institution>
          ,
          <addr-line>Communication and, Performing Arts, Roma Tre</addr-line>
          ,
          <institution>University</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>URBAN/ECO Research, Center, University of</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <institution>"Federico II"</institution>
          ,
          <addr-line>Naples</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>2091</volume>
      <abstract>
        <p>Creating interfaces for cultural heritage access is considered a fundamental research field because of the many beneficial efects it has on society. In this era of significant advances towards natural interaction with machines and deeper understanding of social communication nuances, it is important to investigate the communicative strategies human experts adopt when delivering contents to the visitors of cultural sites, as this allows the creation of a strong theoretical background for the development of eficient conversational agents. In this work, we present the data collection and annotation protocols adopted for the ongoing creation of the reference material to be used in the Cultural Heritage Resources Orienting Multimodal Experiences (CHROME) project to accomplish that goal.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Human-centered computing → User studies; HCI theory,
concepts and models; User models;
Corpus collection, guided tours, social signal processing</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        Developing Social Signal Processing [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] techniques for advanced,
natural interfaces requires a significant analysis efort on multiple
aspects of communication between individuals engaging in social
activity. Collecting meaningful corpora to document the multimodal
signals people exchange during these activities has been the subject
of a large amount of research. Among others, available corpora
document meetings [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], intercultural dynamics of first
acquaintance [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], phone calls between non acquainted subjects [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], and
two-person dialogues [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The Italian national project CHROME
aims at developing a data collection and annotation procedure to
support the development of new interactive technologies for
cultural heritage. The project concentrates on the three Campanian
Charterhouses: an integrated description of these from diferent
point of views (textual, behavioural, geometrical, etc. . . ) is being
developed.
      </p>
      <p>In this paper, we present the data collection and annotation
protocols adopted in the CHROME project to obtain reference material
of expert gatekeepers, intended as holders of knowledge for others
to refer to, accompanying visitors of cultural sites. This data will be
used to investigate the social communication strategies adopted by
the considered experts to deliver information to diferent groups of
visitors. By comparing diferent experts (inter-subject comparisons)
and diferent groups accompanied by the same expert (intra-subject
comparison) a Gatekeeper Computational Model will be obtained
and, on the basis of this model, a socially aware conversational
agent, in the form of a 3D avatar, will be developed. This is
expected to improve the capabilities of an interactive agent to involve
people in engaging presentations of cultural heritage. These will
make use of the 3D reconstructions of the three Campanian
Charterhouses, also collected in the framework of the CHROME project.</p>
      <p>Upon completion of the project, the dataset will be made freely
available for the scientific community.</p>
      <p>In the next sections, we will present the data collection protocol,
highlighting the chosen recording positions in the site of interest
and the recording setup. We will, then, present the multimodal
annotation protocol, designed to provide a formal description of
how the guide makes use of social signals exchange to adapt the
presentation and to efectively support the verbal transfer of
cultural contents. Next, we will describe the informative, syntactic
and prosodic annotations documenting the linguistic behaviour
that characterises the domain expert. The transcribed recordings,
together with the produced annotations, will be compared with a
corpus of textual resources describing the objects of interest. This
will support the development of a synthetic voice model for 3D
avatars designed to extract cultural contents from textual databases
and deliver them using social communication strategies. To improve
the quality of the model, the linguistic analysis will also include a
detailed annotation of disfluency phenomena, which are important
to produce a natural sounding voice.
2</p>
    </sec>
    <sec id="sec-3">
      <title>DATA COLLECTION</title>
      <p>The data collection plan foresees a campaign of audiovisual
recordings involving four art historians with strong experience in
accompanying groups of visitors. Given the limited number of
gatekeepers considered in the CHROME project, only female experts were
recruited to remove gender efects in multimodal and linguistic
analysis. Future extensions of the corpus will include male experts
as well.</p>
      <p>Recorded data include two Full-HD video recordings: the first
one is a fixed shot of the gatekeeper, taken from a position
immediately next to the attending group while the second one is a
ifxed shot of the visitors. A close range digital microphone with
background noise cancellation is used to record the gatekeeper’s
voice. Immediately after the visit, the recruited visitors compile
a questionnaire composed of 23 items including both Likert scale
evaluations and open answer questions. The items are designed to
collect anagraphic data, a self-evaluation of artistic competence, an
evaluation of personal satisfaction after the visit and an evaluation
of the gatekeeper’s performance. These data will be used to weight
objective measures of social behaviour.</p>
      <p>Each recruited expert accompanies four groups of four people in
an hour long guided tour at the San Martino Charterhouse in Naples.
Recruited members of the audience vary on a socio-demographic
basis and each group is gender balanced. The visit is divided into
six points of interest (POIs), selected as the most relevant parts of
the Charterhouse from an architectural and artistic point of view:
• Pronaos: outside the doorstep of the church. The
introductory part of the visit is recorded in this POI. Environmental
elements mainly consist of architectural details;
• Great cloister: a large external place, near the monks’
cemetery. Further details about the monks’ life are given.
Environmental elements consist of the natural setting of a large
garden and of the cemetery elements (e.g. memento mori);
• Parlor: the first internal setting. Specific details about the
Charthusians’ rules are given here. Environmental elements
mainly consist of frescoes;
• Chapter hall: next to the parlor. Specific details about the
Charthusians’ order are given here. Environmental elements
mainly consist of frescoes;
• Wooden choir: inside the church, behind the altar. The history
of the church decoration process is given here.
Environmental elements consist of both architectural details (e.g. the
choir and the harmonic chassis) and artistic elements
(frescoes and statues);
• Treasure hall: deeper inside the complex. Details about the
relationship between the monks and the diferent governing
parties in Naples are given. Environmental elements mainly
consist of architectural details.</p>
      <p>The selected POIs allow us to capture the social behaviour
visitors and gatekeepers exhibit to negotiate the approach to the visit
and to document postural and gestural behaviour of an art historian
presenting a complex environment.</p>
      <p>
        Videos and audio recordings are synchronised a posteriori using
a visual-acoustic marker. Linguistic and multimodal annotations,
performed on the synchronised versions of the collected material,
will be merged using the ELAN software [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. An ELAN project file
will be produced for each POI visit in order to allow cross-domain
research and closed vocabularies for the label sets belonging to
each annotation domain will be used to ensure consistency. An
example of the ELAN interface showing the two video shots and a
sample annotation tier is shown in Figure 1.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>MULTIMODAL ANNOTATION</title>
      <p>
        The video recording of the expert gatekeeper is annotated as to the
structure of verbal discourse and to body communicative behaviour.
The discourse structure point of view is based on a previous
analysis on videos of Art Commentators (ACs), that is, both museums
gatekeepers and art historians illustrating artworks in tv, where a
general script was extracted of what the AC can /should say in one’s
work. This allowed to outline the typical discourse structure of any
AC which, based on the analysis of discourse as a hierarchy of goals
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], distinguishes four main goals pursued by the gatekeeper: a
general goal of cultural elevation; encompassing favouring aesthetic
enjoyment, imagination and emotion triggering, and, subsumed to
it; the textual goals of providing information about the opera, its
history, function, cultural milieu, and the author; the corresponding
Audiovisual Corpus of Guided Tours in Cultural Sites
modal goals of attracting and sustaining attention, favouring
comprehension and inferential connections with the tourists’ previous
knowledge; interactional goals such as tuning, setting empathic
connection with tourists. Each particular performance of a gatekeeper
or other AC can be analysed in terms of this abstract script, and
this allows, among other things, to distinguish the idiosyncratic
styles of diferent ACs in terms of which nodes of the structure
they prefer to expand. Some mainly focus on the author and his
life, some on the deep symbolic meanings of the artwork, some on
the author’s style and the surrounding cultural milieu, and so on.
      </p>
      <p>
        The analysis of the gatekeeper’s multimodal communication takes
into account the following body communicative modalities:
gestures, postures, head movements, facial expression, gaze
communication. For each communicative item in each modality, the signal is
annotated in ELAN in terms of a detailed description of its
production: gestures are described according to their parameters of hand
configuration, location, orientation and movement; gaze in terms
of eye direction, eyebrows and eyelids movements; face in terms of
Ekman’s FACS; head movements in term of head nod, shake, toss,
canting; postures in terms of leg and trunk movements. Then, for
the signal described in this way, a verbal phrasing of its meaning is
provided (after [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]). Based on this meaning, the item is classified as
to its role and function within the gatekeeper’s discourse structure.
An example of multimodal annotation is shown in Table 1.
      </p>
    </sec>
    <sec id="sec-5">
      <title>LINGUISTIC ANNOTATION</title>
      <p>
        Using the close-mic recordings, speech produced by the expert
gatekeeper is analysed and annotated on diferent levels. From the
informative-syntactic point of view, an orthographic level is
produced on the basis of the indications provided by [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. This level
involves the transcription of a number of elements: lexical elements,
silent and filled pauses, noises, vocal (nonverbal) phenomena,
truncated words, interrupted words, false starts and lapsus linguae. A
phonetic level is included to store the phonetic transcription of the
utterances and markers of phonetic phenomena like coarticulation,
following the indications found in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. A syllabic level is produced
to allow speech fluency and speech rate analyses. A disfluency level,
involving the annotation of disfluency phenomena [
        <xref ref-type="bibr" rid="ref13 ref3">3, 13</xref>
        ], is also
included. This analysis level consists of four annotation tiers, detailed
in Table 2.
      </p>
      <p>
        To document the prosodic component of the experts’ linguistic
behaviour, a multilevel annotation, structured in diferent tiers,
has been produced. The considered aspects include: an intonative
level, using the INTSINT coding scheme [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ], providing a labels
sequence representing the f0 curve, obtained with the Prosomarker
tool [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]; a pragmatic - informative level, providing an analysis of
information structure considering topic (preposed or postposed)
and comment units [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; a macro-syntactic level, indicating the types
of clauses dividing independent clauses from dependent clauses and
specifying the type of subordination; a syntactic level, describing the
main syntactic functions; an intra-syntactic level, labelling the type
of phrase and its composition (between parenthesis); a measure of
syntactic weight, based on [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], which takes into account both the
structure and the length of constituents. It considers the following
features: ± presence of determiners, ± presence of modifiers, ±
presence of pronouns, ± verbal valency saturation. An annotation
example is shown in Figure 2.
      </p>
    </sec>
    <sec id="sec-6">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>We have presented the data collection and annotations protocols
for a work in progress on an audiovisual corpus documenting how
cultural heritage gatekeepers support people in accessing
architectural heritage and consists of both video and audio recordings
to capture the social interaction process taking place between the
)
z
H
(
h
c
ti
P
SUB</p>
      <p>NP4
NP(DET+N)</p>
      <sec id="sec-6-1">
        <title>PRED VP(V) VP4</title>
      </sec>
      <sec id="sec-6-2">
        <title>NP(DET+N+PP(PREP+DET+N))</title>
      </sec>
      <sec id="sec-6-3">
        <title>PP(PREP+NP(DET+N)) IO PP4 9.862</title>
        <p>group guide and the attending audience. Annotation levels cover
linguistic and multimodal aspects of communication to allow a
multi-faceted investigation of the ongoing communicative process.
The collected material will be used as reference to build a
computational model of a 3D virtual character presenting reconstructions
of architectural heritage sites.
6</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>ACKNOWLEDGMENTS</title>
      <p>Antonio Origlia’s work is funded by the Italian PRIN project
Cultural Heritage Resources Orienting Multimodal Experience (CHROME)
#B52F15000450001.
OBJ</p>
      <p>NP5</p>
      <sec id="sec-7-1">
        <title>Time (s) C</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jens</given-names>
            <surname>Allwood</surname>
          </string-name>
          , Nataliya Berbyuk Lindström, and
          <string-name>
            <given-names>Jia</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Intercultural dynamics of fist acquaintance: comparative study of swedish, chinese and swedishchinese first time encounters</article-title>
          .
          <source>In International Conference on Universal Access in Human-Computer Interaction</source>
          . Springer,
          <fpage>12</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Jeanette</surname>
            <given-names>K</given-names>
          </string-name>
          <string-name>
            <surname>Gundel</surname>
          </string-name>
          .
          <year>1988</year>
          .
          <article-title>Universals of topic-comment structure</article-title>
          .
          <source>Studies in syntactic typology 17</source>
          (
          <year>1988</year>
          ),
          <fpage>209</fpage>
          -
          <lpage>239</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Adolf</surname>
            <given-names>E</given-names>
          </string-name>
          <string-name>
            <surname>Hieke</surname>
          </string-name>
          .
          <year>1981</year>
          .
          <article-title>A content-processing view of hesitation phenomena</article-title>
          .
          <source>Language and Speech</source>
          <volume>24</volume>
          ,
          <issue>2</issue>
          (
          <year>1981</year>
          ),
          <fpage>147</fpage>
          -
          <lpage>160</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Hirst</surname>
          </string-name>
          and
          <string-name>
            <given-names>Albert Di</given-names>
            <surname>Cristo</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>A survey of intonation systems. Intonation systems: A survey of twenty languages (</article-title>
          <year>1998</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Hirst</surname>
          </string-name>
          ,
          <source>Albert Di Cristo, and Robert Espesser</source>
          .
          <year>2000</year>
          .
          <article-title>Levels of representation and levels of analysis for the description of intonation systems</article-title>
          .
          <source>In Prosody: Theory and experiment</source>
          . Springer,
          <fpage>51</fpage>
          -
          <lpage>87</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Iain</surname>
            <given-names>McCowan</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Jean</given-names>
            <surname>Carletta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W</given-names>
            <surname>Kraaij</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Ashby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Bourban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Flynn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Guillemot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Hain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Kadlec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V</given-names>
            <surname>Karaiskos</surname>
          </string-name>
          , and others.
          <year>2005</year>
          .
          <article-title>The AMI meeting corpus</article-title>
          .
          <source>In Proceedings of the 5th International Conference on Methods and Techniques in Behavioral Research</source>
          , Vol.
          <volume>88</volume>
          . 100.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Antonio</given-names>
            <surname>Origlia</surname>
          </string-name>
          and
          <string-name>
            <given-names>Iolanda</given-names>
            <surname>Alfano</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Prosomarker: a prosodic analysis tool based on optimal pitch stylization and automatic syllabi fication</article-title>
          ..
          <source>In Proc. of the International Conference on Language Resources and Evaluation (LREC)</source>
          .
          <volume>997</volume>
          -
          <fpage>1002</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Domenico</given-names>
            <surname>Parisi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Cristiano</given-names>
            <surname>Castelfranchi</surname>
          </string-name>
          .
          <year>1976</year>
          .
          <article-title>The discourse as a hierarchy of goals. Centro Internazionale di Semiotica e di Linguistica, Università di Urbino</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Isabella</given-names>
            <surname>Poggi</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Mind, hands, face and body: a goal and belief view of multimodal communication</article-title>
          .
          <source>Weidler.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hugues</surname>
            <given-names>Salamin</given-names>
          </string-name>
          , Anna Polychroniou, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Vinciarelli</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Automatic detection of laughter and fillers in spontaneous mobile phone conversations</article-title>
          .
          <source>In Systems, Man, and Cybernetics (SMC)</source>
          ,
          <source>2013 IEEE International Conference on. IEEE</source>
          ,
          <fpage>4282</fpage>
          -
          <lpage>4287</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Renata</given-names>
            <surname>Savy</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Specifiche per la trascrizione ortografica annotata dei testi</article-title>
          .
          <source>Italiano Parlato</source>
          ,
          <article-title>Analisi di un dialogo (</article-title>
          <year>2005</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Renata</given-names>
            <surname>Savy</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Specifiche per l'etichettatura dei livelli segmentali</article-title>
          .
          <source>Italiano Parlato</source>
          .
          <article-title>Analisi di un dialogo</article-title>
          .
          <source>Napoli: Liguori</source>
          (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Elizabeth</surname>
            <given-names>Ellen</given-names>
          </string-name>
          <string-name>
            <surname>Shriberg</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Preliminaries to a theory of speech disfluencies</article-title>
          .
          <source>Ph.D. Dissertation</source>
          . Citeseer.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Yasir</surname>
            <given-names>Tahir</given-names>
          </string-name>
          , Debsubhra Chakraborty, Tomasz Maszczyk, Shoko Dauwels, Justin Dauwels, Nadia Thalmann, and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Thalmann</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Real-time sociometrics from audio-visual features for two-person dialogs</article-title>
          .
          <source>In Digital Signal Processing (DSP)</source>
          ,
          <source>2015 IEEE International Conference on. IEEE</source>
          ,
          <fpage>823</fpage>
          -
          <lpage>827</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Alessandro</surname>
            <given-names>Vinciarelli</given-names>
          </string-name>
          , Maja Pantic, and
          <string-name>
            <given-names>Hervé</given-names>
            <surname>Bourlard</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Social signal processing: Survey of an emerging domain</article-title>
          .
          <source>Image and vision computing 27</source>
          ,
          <issue>12</issue>
          (
          <year>2009</year>
          ),
          <fpage>1743</fpage>
          -
          <lpage>1759</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Miriam</given-names>
            <surname>Voghera</surname>
          </string-name>
          and
          <string-name>
            <given-names>Giuseppina</given-names>
            <surname>Turco</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Il peso del parlare e dello scrivere</article-title>
          .
          <source>In Proc. of International Conf. Il Parlato Italiano</source>
          , Liguori, Napoli.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Wittenburg</surname>
          </string-name>
          , Hennie Brugman,
          <string-name>
            <surname>Albert Russel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Alex</given-names>
            <surname>Klassmann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Han</given-names>
            <surname>Sloetjes</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>ELAN: a professional framework for multimodality research</article-title>
          .
          <source>In Proc. of the International Conference on Language Resources and Evaluation (LREC)</source>
          .
          <volume>1556</volume>
          -
          <fpage>1559</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>