<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On the Role of Communicative Structure in Read Aloud Applications for the Elderly</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mónica Domínguez</string-name>
          <email>monica.dominguez@upf.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mireia Farrús</string-name>
          <email>mireia.farrus@upf.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alicia Burga</string-name>
          <email>alicia.burga@upf.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leo Wanner</string-name>
          <email>leo.wanner@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Catalan Institute for Research and Advanced Studies and, University Pompeu Fabra</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University Pompeu Fabra</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <fpage>40</fpage>
      <lpage>44</lpage>
      <abstract>
        <p>Conversational technologies that assist elderly people need to adapt to common disabilities in old age. Visual, hearing and even more so cognitive impairments pose serious dificulties for our seniors to handle a standard conversation with a human. Understanding a virtual agent may be ever harder. In this case, communicative strategies are key to adapt the virtual agent to the needs of elderly users. This paper addresses the role of the communicative structure for expressive speech prosody, which is known to be crucial for better speech comprehension. It reports on eforts to improve prosody within a text-to-speech system based on one aspect of the communicative structure, namely thematicity. The work has been implemented as an application in a social virtual agent, KRISTINA, which reads aloud news articles upon request for elderly users in German.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Social and professional topics → Seniors;</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>In the last decades, conversational interfaces involving
text-tospeech (TTS) applications have improved expressiveness and
overall naturalness to a reasonable extent. Conversational features, such
as speech acts, afective states and Information Structure have been
instrumental to derive more expressive prosodic contours. However,
synthetic speech is still perceived as monotonous, when a text that
lacks those conversational features is read aloud in the interface,
i.e., when it is fed directly to the TTS application. If users of the
conversational interface furthermore have some impairments, as it
is usually the case with elderly people using assisting technologies,
it is paramount to adapt the conversational agent’s speech to
guarantee the communication flow of the interaction, and thus improve
the acceptance of the agent by the user. This adaptation requires
advanced functionalities that usually involve several areas of
expertise. In this paper, we present how theoretical linguistics can be
used in computational approaches to achieve a more fine-grained
communicative interaction adapted to the elderly.</p>
      <p>
        Virtual agents with human interaction capabilities have a large
potential for the exploration of such user-oriented advanced
functionalities. We work with KRISTINA. KRISTINA is a
KnowledgeBased Information Agent with Social Competence and Human
Interaction Capabilities [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. KRISTINA interacts with the user
in diferent scenarios. One of these scenarios consists in reading
the newspaper to elderly people with eyesight impairments. This
target audience requires a varied range of expressiveness in the
synthetic voice, which state-of-the-art text-to-speech (TTS)
applications usually lack, especially when processing long monologue
discourse.
      </p>
      <p>This paper discusses the role of the Information (or
Communicative) Structure–prosody interface for reading aloud applications,
and scratches the surface of the theoretical framework behind this
interface. The discussion is based upon the authors’ implementation
of a thematicity-based prosody module that enriches raw texts
extracted from news with communicative information with the goal to
achieve a more expressive reading for targeted elderly users.1 The
aim is to analyze syntactic and Information Structure, and then use
high-level linguistic features derived from the analysis to generate
more expressive prosody in the synthesized speech. The proposed
methodology encompasses a modular pipeline consisting of (1) a
tokenizer, (2) a syntactic parser, (3) a theme/rheme parser, and (3) an
SSML prosody tag converter. The implementation has been tested
in an experimental setting for German, using web-retrieved news
articles.</p>
      <p>The rest of the paper is structured as follows. Section 2
introduces the motivation and background of this work. In Section 3,
we dive into the theoretical grounds that support the proposed
computational model from a linguistic perspective. Then, Section 4
sketches how this model has been implemented within the context
of KRISTINA. Finally, conclusions are drawn in Section 5.
2</p>
    </sec>
    <sec id="sec-3">
      <title>MOTIVATION AND BACKGROUND</title>
      <p>The way information is formally packaged in a sentence, known
as “Information Structure”, has been a fruitful field of research
in linguistic studies to better understand how communication is
1Such an application may also be handy for other users, not only elderly.
produced and perceived. Information Structure is a wide term and
its study usually involves various linguistic dimensions in
connection with how content is packaged, hence its interfaces at least
semantics, syntax and prosody.</p>
      <p>
        Diferent linguistic schools have long stated that Information
Structure, and, in particular, the dichotomy referred to as theme–
rheme [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], given–new [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], or topic–focus [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] is related to
intonation.2 Moreover, prosody structure on the grounds of thematicity
partitions plays a key role in the understanding of a message [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Empirical studies in diferent languages provide evidence that when
thematicity and prosody are appropriately put together,
comprehension of the message is positively afected (cf., e.g., [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] for German
and [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] for Catalan). Several works also show that a correlation
between thematicity and beat gestures, which are an important
non-verbal “prosodic” means to mark rythm and to “accentuate
speech” [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], improves discourse recall and comprehension [
        <xref ref-type="bibr" rid="ref18 ref20">18, 20</xref>
        ].
Therefore, there is reason to assume that a conversational
application considering the notions of content packaging by means of
the relation between thematicity and prosody will benefit from the
same advantages as in natural conversation environments. Most
of all, conversational avatars in applications for children in
educational settings [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], applications for those with special needs
[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] as well as for the elderly [
        <xref ref-type="bibr" rid="ref25 ref33">25, 33</xref>
        ] and, in particular, for those
with cognitive impairments [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ], would greatly benefit from such
a communicatively-oriented improvement.
      </p>
      <p>
        On the other hand, expressive speech that uses a varied range of
prosodic cues (variation in fundamental frequency, speech rate and
intensity) is often regarded as more understandable and
communicative. However, previous attempts to implement the concepts
of Information Structure in text-to-speech (TTS) applications are
rather scarce [
        <xref ref-type="bibr" rid="ref19 ref27">19, 27</xref>
        ]. Moreover, it is usually a simple binary theme–
rheme structure what is being tested in short sentences. A more
ifne-grained analysis of thematicity structure, as defined by Mel’čuk
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] has been proved to yield better results to predict a wider
variety of prosodic contours, which are furthermore perceived as more
natural when implemented in a TTS application; see, e.g. [
        <xref ref-type="bibr" rid="ref11 ref15">11, 15</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>COMMUNICATIVE STRUCTURE</title>
      <p>Despite the great eforts along the years for defining communicative
notions, studies on Information Structure have remained within
the field of theoretical linguistics. These studies sometimes explore
diferent linguistic phenomena in relation to Information Structure
(e.g., discourse, dialog, anaphora, and co-reference). The
Communicative Structure within the Meaning-Text Theory (MTT) comes
to cope with some of the limitations other theories on Information
Structure have, as this representation is devised in the context of a
theoretical production-oriented linguistic model, which is described
in what follows.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>A Theoretical Framework for</title>
    </sec>
    <sec id="sec-6">
      <title>Computational Linguistics</title>
      <p>
        The Meaning-Text Theory proposes a framework for language
analysis and generation suitable for Natural Language Processing (NLP)
applications [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In particular, the Meaning-Text Theory Model
2In our work, we use the first denotation, i.e., theme–rheme or thematicity. ‘Theme’
marks what a sentence is about, and ‘rheme’ what is said about the theme.
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] distinguishes diferent levels of representation. These levels
are sequentially mapped from an unordered semantic
representation (SemR) through a dependency tree structure of the Syntactic
Representation (SyntR) and linearized chain of lexemes onto the
Morphological Representation (MorphR) to get to the ordered string
of phonemes at the Phonetic Representation (PhonR). Starting from
SyntR and until PhonR, there is a subdivision into deep and surface
representations.
      </p>
      <p>The SemR includes four structures: (1) the Semantic Structure
(SemS), which is a predicate-argument (meaning) structure of the
message; (2) the Semantic Communicative Structure (SemCommS),
which consists of a representation of the communicative intention
of the speaker; (3) the Rhetorical Structure (RhetS), which encodes
the artistic intentions and stylistic decisions of the speaker (irony,
humorous, etc.); and (4) the Referential Structure (RefS), which
specifies real-world referents for semantic configurations. The
SemCommS superimposes on the SemS the communicative properties
of the meaning of the sentence to be synthesized rather than the
communicative properties of the sentence itself.3 Consequently,
the functions of SemCommS are:
• organizing initial meaning into a message;
• ensuring coherence of the text of which the sentence under
synthesis is supposed to be a part;
• reducing periphrastic potential of the initial SemS, specifying
more precisely the meaning.</p>
      <p>
        In other words, the same abstract Semantic Structure can be
shared by a given set of sentences, and it is by means of the
SemCommS that these sentences are distinguished at subsequent levels
(namely, SyntR, MorphR and PhonR). Figure 1 sketches the common
SemS of sentences from (1a) to (1d) taken from [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>(1a) John met the doctor at the airport.
(1b) The doctor was met at the airport by John.
(1c) The airport was where John met the doctor.</p>
      <p>(1d) It was John who met the doctor at the airport.</p>
      <p>
        The Deep Syntactic Structure (DSyntS), which may already
relfect some of the SemCommS features, is the central component of
the Deep-Syntactic Representation (DSyntR).4 Consider, for
illustration, the DSyntS’s of sentences (1a) (Figure 2) and (1d) (Figure
3). They show how SemCommS determines the diferent resulting
3In general linguistics, the term ‘communicative’ is usually linked to the idea of
‘communicative competence’ and refers to concepts related to the study of pragmatics;
see the definition of ‘linguistic competence’ and ‘performance’ by Chomsky [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
4Apart from DSyntS, DSyntR includes, in its turn, three further components:
DeepSyntactic Communicative Structure, Deep-Syntactic Anaphoric Structure and
DeepSyntactic Prosodic Structure (which represents semantically conditioned prosodies).
dependency trees. The communicative subject (Theme) may
coincide or not with the semantic subject (Actor) and syntactic subject
(Synt-Subject), as represented in Table 1. This underlines the idea
that CommS is a distinct dimension.
      </p>
      <p>
        In a nutshell, CommS is part of the SemR and DSyntR of
individual sentences. The communicative organization of text is not
covered by CommS, it rather accounts for the structure of the
socalled propositional content. Going back to example (1) taken from
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], the set of sentences may seem fully synonymous, but only (1a)
is an appropriate reply to D1, whereas (1d) better suits D2:
      </p>
      <sec id="sec-6-1">
        <title>D1 - Nobody saw the doctor last night?</title>
        <p>- John met him at the airport.</p>
      </sec>
      <sec id="sec-6-2">
        <title>D2 - Ask John. - Why John? - It was John who met the doctor at the airport.</title>
        <p>
          CommS is composed of eight distinct dimensions: ‘thematicity’,
‘givenness’, ‘focalization’, ‘perspective’, ‘emphasis’,
‘presupposedness’, ‘unitariness’ and ‘locutionality’. As CommS characterizes
the meaning of the sentence and the sentence itself, it is,
consequently, modeled at the semantic level, to be propagated then to the
deep-syntactic and surface-syntactic levels of the linguistic
description. Note that givenness, which is often treated as synonymous
to thematicity, is in Mel’čuk’s communicative structure theory a
distinct dimension from thematicity. According to Mel’čuk [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ],
the thematicity of the initial SemS has to do with psychologically
motivated choices of the speaker, who decides that he/she wants
to communicate some specific information (i.e., the rheme)
concerning some specific item (i.e., the theme), and thereby makes the
addressee follow him. In Mel’čuk’s words: “The Sem-Thematicity
is thus a SPEAKER-ORIENTED Comm-category.”
        </p>
        <p>
          In the following section, we sketch Mel’čuk’s definition of
thematicity, which is the dimension considered in previous work when
the correspondence of the Information Structure with prosody is
discussed.
In contrast to Information Structure models that propose a partition
of sentences into a theme and a rheme, Mel’čuk [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] argues in the
context of the Meaning–Text Theory for a tripartite hierarchical
division (‘theme’, ‘rheme’, and ‘specifier’ –the element which sets
the utterance’s context) within propositions that further permits
embeddedness of communicative spans; consider (1) for
illustration of hierarchical thematicity (annotated following the guidelines
established in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]) of the sentence Ever since, the remaining
members have been desperate for the United States to rejoin this dreadful
group. A total of five partitions are identified, including three spans
at level 1, a specifier (SP1), theme (T1) and rheme (R1), and two
embedded spans at level 2 in the rheme, a theme (T1(R1)) and a
rheme (R1(R1)).5
(1) [Ever since,]SP1 [the remaining members]T1 [have been
desperate [for the United States]T1(R1) [to rejoin this dreadful
group.]R1(R1)]R1
        </p>
        <p>
          A hierarchical thematicity structure of this kind has been shown
to correlate better with ToBI [
          <xref ref-type="bibr" rid="ref1 ref29">1, 29</xref>
          ] labels than binary flat
thematicity [
          <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
          ]. Such a correlation still does not solve the problem of a
one–to-one mapping between a specific intonation label (e.g., H*)
to a static acoustic parameter (e.g., an increase of 50% in
fundamental frequency). This is one of the reasons why we propose an
implementation using a more varied range of automatically derived
prosodic cues based on hierarchical thematicity spans, as described
in what follows.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>AUTOMATIC GENERATION OF</title>
    </sec>
    <sec id="sec-8">
      <title>THEMATICITY-BASED PROSODY IN</title>
    </sec>
    <sec id="sec-9">
      <title>KRISTINA</title>
      <p>
        In the use case of KRISTINA as social companion for the elderly, the
scenario of reading the newspaper involves a dialogue interaction
between the user (U) and KRISTINA (K). U requests K to read the
newspaper and K prompts U to pick up a piece of news. Upon
reading of the title, the system retrieves the selected text, which
is sent to the pipeline sketched in Figure 4. The pipeline tests the
formal representation of the Communicative Structure, in particular
of thematicity, proposed by Mel’čuk [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. In the context of the
conversational agent KRISTINA, text coming from a web-retrieved
service is processed in the pipeline before it arrives to the TTS
engine.
      </p>
      <p>
        The proposed pipeline in Figure 4 includes four modules:
5As more than one thematicity span may exist within the same proposition,
abbreviations include a number (e.g., ‘SP1’) that indicates the number of occurrences at each
level (e.g., ‘SP2’ would be the second specifier in a specific thematicity level).
(1) Tokenizer: Splits the text into sentences and words.
Punctuation marks are also tokenized as the syntactic parser
requires that.
(2) Syntactic parser: An of-the-shelf parser [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which is trained
on the TIGER Penn Treebank [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and which outputs a
fourteencolumned CoNLL file. 6
(3) Communicative parser: Derives using rules hierarchical
thematicity labels from syntactic structure. It outputs a CoNLL
ifle with an added column for communicative structure (i.e.,
the output CoNLL has fifteen columns).
(4) SSML prosody converter: Converts the thematicity spans
derived by the communicative parser to SSML spans and
assigns a variety of prosody tags to each span. This module
is based on the tool presented in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        The correspondence between hierarchical thematicity and prosody
is presented in terms of variations of referent SSML7 [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] prosody
tag values involving fundamental frequency (F0), speech rate (SR)
and insertion of breaks.
5
      </p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSIONS</title>
      <p>
        Theoretical studies on the Information Structure-prosody interface
have stated for some time that there is a correspondence between
how the linguistic content is structured communicatively and how
intonation is used in human speech to convey that content. In
previous work, this correspondence (in particular, the
relationship between hierarchical thematicity and prosodic variation) has
been brought to the foreground from an empirical perspective in
the context of expressive speech generation. Corpus-based
experiments and data-driven implementations [
        <xref ref-type="bibr" rid="ref12 ref13 ref14 ref15">12–15</xref>
        ] supported initial
expectations on the potential of the Information Structure–prosody
interface applied to speech technologies. The use of this potential is
an initial step ahead in communicative approaches for prosody
generation within TTS/CTS applications that is one of the key aspects
for a next generation of more expressive conversational virtual
agents.
      </p>
      <p>The implementation described above contributes in several
aspects to the state of the art: (i) a formal description of hierarchical
thematicity is used; (ii) a communicative parser that automatically
6Details about the CoNLL format are provided in
http://universaldependencies.org/docs/format.html
7SSML stands for Speech Synthesis Markup Language: details about this convention
can be found in https://www.w3.org/TR/speech-synthesis11/
derives thematicity labels is introduced; and (iii) a platform for
prosody testing in TTS applications is demonstrated. Evaluation
shows that the thematicity-based prosody enrichment is perceived
as more expressive than the default TTS output. Expressiveness
was assessed by means of a perception test using a Mean Opinion
Score (MOS) with a 5-point Likert scale (LS): 1-bad, 2-poor, 3-fair,
4-good, and 5-excellent. Average results for the tested sentences
proved that the automatic prosody modifications (LS = 3.30) achieve
statistical significance at p &lt;0.05 compared to the default score (LS
= 3.01). All in all, this study pivots the transition from theoretical
work on the IS–prosody interface to the integration of
thematicitybased prosody enrichment to achieve more expressive synthesized
speech. Future work is aimed at exploring other dimensions of
communicative structure like emphasis and foregroundedness within
the framework that has been discussed.</p>
      <p>
        Research carried out so far in this direction [
        <xref ref-type="bibr" rid="ref15 ref9">9, 15</xref>
        ] is a proof of
concept of the applicability of the Information Structure–prosody
interface in speech synthesis, but there are many issues that
remain unexplored. For now, only thematicity at the sentence level
has been tested. Other dimensions of the communicative
structure (like givenness and focus, as defined by Mel’čuk [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]) may
also have a strong correspondence with prosody. Corpora need to
be compiled in order to continue looking into this field from an
empirical perspective; see e.g. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. With respect to prosody, an
implementation with SSML tags does not sufice to address the
requirements for prosody modeling in a pre-processing stage for TTS
applications. Therefore, closer insights into how to model prosody
to reflect better the communicative structure of a text also need to
be investigated.
      </p>
      <p>Given the relevant role of the Information Structure–prosody
interface in human communication, it seems reasonable that next
generation conversational agents face new challenges in adopting
communicatively-oriented models. In this paper, we have
introduced some basic concepts on the theoretical framework behind
an implementation of a hierarchical thematicity model as well as
an overview of the research carried out so far in this area in its
correspondence to prosody.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Beckman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Hirschberg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Shattuck-Hufnagel</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The Original ToBI System and the Evolution of the ToBI Framework</article-title>
          . In Prosodic Models and Transcription: Towards Prosodic Typology,
          <string-name>
            <given-names>S.A.</given-names>
            <surname>Jun</surname>
          </string-name>
          (Ed.). Oxford University Press,
          <fpage>9</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Burga</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Towards the Annotation of Penn TreeBank with Information Structure</article-title>
          .
          <source>In Proceedings of the Sixth International Joint Conference on Natural Language Processing. Nagoya, Japan</source>
          ,
          <fpage>1250</fpage>
          -
          <lpage>1256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Nivre</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A Transition-Based System for Joint Part-of-Speech Tagging and Labeled Non-Projective Dependency Parsing</article-title>
          .
          <source>In Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL '12)</source>
          . Jeju Island, Korea,
          <fpage>1455</fpage>
          -
          <lpage>1465</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Open Source Graph Transducer Interpreter and Grammar Development Environment</article-title>
          .
          <source>In Proceedings of the Seventh Conference on International Language Resources and Evaluation (LREC)</source>
          .
          <source>European Language Resources Association (ELRA)</source>
          , Valletta, Malta.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bozkurt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yemez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Erzin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Multimodal analysis of speech and arm motion for prosody-driven synthesis of beat gestures</article-title>
          .
          <source>Speech Communication</source>
          <volume>85</volume>
          (12
          <year>2016</year>
          ),
          <fpage>29</fpage>
          -
          <lpage>42</lpage>
          . https://doi.org/10.1016/J.SPECOM.
          <year>2016</year>
          .
          <volume>10</volume>
          .004
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Brants</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dipper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Eisenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E</given-names>
            <surname>König</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lezius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rohrer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and H.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>TIGER: Linguistic Interpretation of a German Corpus</article-title>
          .
          <source>Journal of Language and Computation</source>
          <volume>2</volume>
          (
          <year>2004</year>
          ),
          <fpage>597</fpage>
          -
          <lpage>620</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Chomsky</surname>
          </string-name>
          .
          <year>1965</year>
          .
          <article-title>Aspects of the Theory of Syntax</article-title>
          . The MIT Press, Cambridge.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Clark</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Haviland</surname>
          </string-name>
          .
          <year>1977</year>
          .
          <article-title>Comprehension and the given-new contract. Discourse production and comprehension</article-title>
          .
          <source>Discourse processes: Advances in research and theory 1</source>
          (
          <year>1977</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The Information Structure-Prosody Interface: On the Role of Hierarchical Thematicity in an Empirically-grounded Model</article-title>
          .
          <source>Ph.D. Dissertation</source>
          . Universitat Pompeu Fabra.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Burga</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The Information Structure - Prosody Language Interface Revisited</article-title>
          .
          <source>In Proceedings of the 7th International Conference on Speech Prosody</source>
          . Dublin, Ireland,
          <fpage>539</fpage>
          -
          <lpage>543</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Burga</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Using hierarchical information structure for prosody prediction in content-to-speech applications</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Speech Prosody</source>
          . Boston, USA,
          <fpage>1019</fpage>
          -
          <lpage>1023</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Combining acoustic and linguistic features in phrase-oriented prosody prediction</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Speech Prosody</source>
          . Boston, USA,
          <fpage>796</fpage>
          -
          <lpage>800</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Thematicity-based Prosody Enrichment Tool for CTS</article-title>
          .
          <source>In Proceedings of the 18th Annual Conference of the International Speech Communication Association (INTERSPEECH</source>
          <year>2017</year>
          ). Stockholm, Sweden,
          <fpage>3421</fpage>
          -
          <lpage>2</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Compilation of Corpora to Study the Information StructureâĂŞProsody Interface</article-title>
          .
          <source>In 11th edition of the Language Resources and Evaluation Conference (LREC2018)</source>
          . Mijazaki, Japan.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Thematicity-based Prosody Enrichment for Text-to-Speech Applications</article-title>
          .
          <source>In 9th International Conference on Speech Prosody</source>
          <year>2018</year>
          (
          <article-title>SP2018)</article-title>
          . Poznan, Poland.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>E</given-names>
            <surname>Hajicˆova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B</given-names>
            <surname>Partee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P</given-names>
            <surname>Sgall</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <string-name>
            <surname>Topic-Focus</surname>
            <given-names>Articulation</given-names>
          </string-name>
          , Tripartite Structures, and
          <string-name>
            <given-names>Semantic</given-names>
            <surname>Content</surname>
          </string-name>
          . Kluwer Academic Publishers, Dordrecht.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.A.K.</given-names>
            <surname>Halliday</surname>
          </string-name>
          .
          <source>1967. Notes on Transitivity and Theme in English, Parts 1-3. Journal of Linguistics 3</source>
          ,
          <issue>1</issue>
          (
          <year>1967</year>
          ),
          <fpage>37</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Alfonso</surname>
            <given-names>Igualada</given-names>
          </string-name>
          , Núria Estebe-Gibert, and
          <string-name>
            <given-names>Pilar</given-names>
            <surname>Prieto</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Beat gestures improve word recall in 3- to 5-year-old children</article-title>
          .
          <source>Journal of Experimental Child Psychology</source>
          <volume>156</volume>
          (
          <year>2017</year>
          ),
          <fpage>99</fpage>
          -
          <lpage>112</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Frank</surname>
            <given-names>Kügler</given-names>
          </string-name>
          , Bernadett Smolibocki, and
          <string-name>
            <given-names>Manfred</given-names>
            <surname>Stede</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Evaluation of Information Structure in Speech Synthesis : The Case of Product Recommender Systems Perception</article-title>
          .
          <source>In ITG Conference on Speech Communication</source>
          , IEEE.
          <fpage>26</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J</given-names>
            <surname>Llanes-Coromina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I</given-names>
            <surname>Vilà-Giménez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>Kushch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Borràs-Comes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Prieto</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Beat gestures help preschoolers recall and comprehend discourse information</article-title>
          .
          <source>Journal of Experimental Child Psychology</source>
          <volume>172</volume>
          (
          <year>2018</year>
          ),
          <fpage>168</fpage>
          -
          <lpage>188</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>I. A</surname>
          </string-name>
          . Mel'čuk.
          <year>1988</year>
          .
          <article-title>Dependency Syntax: Theory and Practice</article-title>
          . SUNY Press, Albany, NY. 400 pages.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>I. A</surname>
          </string-name>
          . Mel'čuk.
          <year>2001</year>
          .
          <article-title>Communicative Organization in Natural Language: The semantic-communicative structure of sentences</article-title>
          . Benjamins, Amsterdam, Philadephia.
          <volume>393</volume>
          pages.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>B.</given-names>
            <surname>Mencía-López</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Trapote-Hernández</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Gómez-Hernández</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Embodied Conversational Agents in Interactive Applications for Children with Special Educational Needs</article-title>
          .
          <source>In Technologies for Inclusive Education: Beyond Traditional Integration Approaches</source>
          , David Griol Barres, Zoraida Callejas Carrión, and
          <string-name>
            <surname>Ramón</surname>
          </string-name>
          López-Cózar
          <string-name>
            <surname>Delgado</surname>
          </string-name>
          (Eds.).
          <source>IGI Global</source>
          , Hershey, USA,
          <fpage>59</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>D.</given-names>
            <surname>Meurers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ziai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ott</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kopp</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Evaluating Answers to Reading Comprehension Questions in Context: Results for German and the Role of Information Structure</article-title>
          .
          <source>In Proceedings of the TextInfer 2011 Workshop on Textual Entailment (TIWTE '11)</source>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Stroudsburg, PA, USA,
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ortiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. del Puy</given-names>
            <surname>Carretero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Oyarzun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Yanguas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Buiza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I.</given-names>
            <surname>Etxeberria</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Elderly Users in Ambient Intelligence: Does an Avatar Improve the Interaction</article-title>
          ? Springer Berlin Heidelberg, Berlin, Heidelberg,
          <fpage>99</fpage>
          -
          <lpage>114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>D.</given-names>
            <surname>Pérez-Marín</surname>
          </string-name>
          and
          <string-name>
            <given-names>I.</given-names>
            <surname>Pascual-Nieto</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>An exploratory study on how children interact with pedagogic conversational agents</article-title>
          .
          <source>Behaviour &amp; Information Technology 32</source>
          ,
          <issue>9</issue>
          (
          <year>2013</year>
          ),
          <fpage>955</fpage>
          -
          <lpage>964</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Schröder</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Trouvain</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>The German Text-to-Speech Synthesis System MARY: A Tool for Research, Development and Teaching</article-title>
          .
          <source>International Journal of Speech Technology</source>
          <volume>6</volume>
          ,
          <issue>4</issue>
          (
          <year>2003</year>
          ),
          <fpage>365</fpage>
          -
          <lpage>377</lpage>
          . https://doi.org/10.1023/A:1025708916924
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>R.</given-names>
            <surname>Schwarzschild</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>GIVENness, AvoidF and Other Constraints on the Placement of Accent</article-title>
          .
          <source>Natural Language Semantics</source>
          <volume>7</volume>
          ,
          <issue>1</issue>
          (
          <year>1999</year>
          ),
          <fpage>141</fpage>
          -
          <lpage>177</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>K.</given-names>
            <surname>Silverman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Beckman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pitrelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ostendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wightman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pierrehumbert</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Hirschberg</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>ToBI: A standard for labeling English prosody</article-title>
          .
          <source>In Proceedings of Interspeech. Makuhari, Japan</source>
          ,
          <fpage>146</fpage>
          -
          <lpage>149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>P.</given-names>
            <surname>Taylor</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Isard</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>SSML: A Speech Synthesis Markup Language</article-title>
          .
          <source>Speech Communication</source>
          <volume>21</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>February 1997</year>
          ),
          <fpage>123</fpage>
          -
          <lpage>133</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vanrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I</given-names>
            <surname>Mascaró</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Torres-Tamarit</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Prieto</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Intonation as an Encoder of Speaker Certainty: Information and Confirmation Yes-No Questions in Catalan</article-title>
          .
          <source>Language and Speech</source>
          <volume>56</volume>
          ,
          <issue>2</issue>
          (
          <year>2013</year>
          ),
          <fpage>163</fpage>
          -
          <lpage>190</lpage>
          . https://doi.org/10.1177/ 0023830912443942
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>André</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Blat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dasiopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Farrús</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fraga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kamateri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lingenfelser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Llorach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Martínez</surname>
          </string-name>
          , G. Meditskos,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Minker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pragst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schiller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Stam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Stellingwerf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sukno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vieru</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Vrochidis</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>KRISTINA: A Knowledge-Based Virtual Conversation Agent</article-title>
          .
          <source>In Proceedings of the 15th International Conference on Practical Applications of Agents and Multi-Agent Systems (PAAMS)</source>
          . Oporto, Portugal.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Blat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dasiopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Domínguez</surname>
          </string-name>
          , G. Llorach,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mille</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Sukno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kamateri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vrochidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Kompatsiaris</surname>
          </string-name>
          , et al.
          <year>2016</year>
          .
          <article-title>Towards a multimedia knowledge-based agent with social competence and human interaction capabilities</article-title>
          .
          <source>In Proceedings of the 1st International Workshop on Multimedia Analysis and Retrieval for Multimodal Interaction. ACM Digital Library</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>P.</given-names>
            <surname>Wargnier</surname>
          </string-name>
          , G. Carletti,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Laurent-Corniquet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Benveniste</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jouvelot</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Rigaud</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Field evaluation with cognitively-impaired older adults of attention management in the Embodied Conversational Agent Louise</article-title>
          .
          <source>In 2016 IEEE International Conference on Serious Games and Applications</source>
          for Health,
          <source>SeGAH</source>
          <year>2016</year>
          , Orlando, FL, USA, May
          <volume>11</volume>
          -13,
          <year>2016</year>
          . 1-
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>