<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Anticipation and its applications in human-machine interaction</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Stanislav Ondáš, Matúš Pleva Department of Electronics and Multimedia Communications, Faculty of Electrical Engineering and Informatics, Technical University of Košice</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Behinds the capability of a person to be an interlocutor of a conversation there lies many human capabilities, many of which are carried out unconsciously and very naturally in childhood. Human-like turn-taking in the human-machine interactions (HMI) can be seen as a critical issue to achieve natural conversational interaction. The production of the listener response starts before the speaker turn is finished, which means, that listener is often able to anticipate the remaining content of unfolded speaker turn. The ability to anticipate can be identified as a very important human capability, which supports rapid turn-taking. This anticipation process, which occurs during listening relates to sentence comprehension and turn-taking. The proposed paper study this phenomenon on the small corpus of Slovak interviews, where the attention is focused on overlapping segments and discussed possible applications, where anticipatory behavior on the machine side can bring benefits.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Nowadays spoken communication between human and
machine has become obvious, what relates to the new types
of devices, which start to be a part of our everyday life.
Good examples are Amazon Echo, Google Home, TV with
voice control, Voice search applications, social robots.</p>
      <p>We can classify that form of communication mostly as
simple dialogues, what means question-answer scenarios
and task-oriented domain-specific dialogue.</p>
      <p>Unlike mentioned scenarios, in case of human-human
spoken interaction, the situation is significantly more
complex. Here, information exchange is often performed
very quickly, a lot of information is omitted, but
supplemented by the listener based on the common
background and history (short-term, long-term,
topicrelated, speaker-related, situation-related). Behind a
capability of a person to be an interlocutor of a
conversation lies a lot of important human capabilities,
many of which are carried out unconsciously and very
naturally. People are easily able to track the content, to
detect a speech act (dialogue act) behind a speaker’s turn,
to perform effective and rapid turn-taking, to provide
feedback in a role of listener, or incrementally construct
their turns.</p>
      <p>In the proposed work, we focus on the three related
aspects of the spoken communication - turn-taking,
sentence comprehension and anticipation.</p>
      <p>
        “A turn is the time when a speaker is talking, and
turntaking is the skill of knowing when to start and finish a turn
in a conversation. It is an important organizational tool in
spoken discourse.” [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Turn-taking can be described as a
process in which one dialogue participant talks, then stops
and gives the floor to another participant. It is a human
skill, which we learn without any effort in childhood.
      </p>
      <p>
        Stivers et al. in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] observed, that human-human
conversations are characterized by rapid turn-taking often
with a minimal gap lower than 200 msec. Several other
studies (e.g. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) indicate that utterance production can
take more than 600msec (see [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). It indicates, that
utterance productions usually start during listening. This
finding is supported by measurements of EEG signals.
Magyari et al. in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] observed changes in EEG which
relates to the fact, that human brain is able to estimate turn
duration. This estimation is based on anticipating the way
the turn would be completed. They founded a neuronal
correlate of turn-end anticipation and a beta frequency
desynchronization as early as 1250 msec, before the end of
the turn. They suggest that anticipation of the speaker
utterance leads to accurately timed transitions in everyday
conversations. The ability to anticipate the content of the
speaker turn and turn end was researched and confirmed in
several papers (see e.g. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]).
      </p>
      <p>Figure 1 Turn-taking timing</p>
      <p>Anticipation relates to the sentence comprehension. It
can be inferred, that instead of the word sequence, the
meaning is anticipated. We can identify, that in the
moment, when a person is able to anticipate remaining part
of the speaker turn, there is partial or complete
comprehension, which can be signalized using one of the
following behavior: providing a feedback (backchannel
signals) or attempt to take the floor.</p>
      <p>
        The attempt to take the floor can result into the
overlapping speech. Overlapping speech is a segment of the
conversation, where both interlocutors speak
simultaneously. The speaker tries to finish his turn, but the
listener starts his own turn. Several studies quantified a
significant occurrence of overlaps in dyadic human-human
spoken interactions (e.g 18.9% in [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]).
      </p>
      <p>We supposed that overlaps are the best place to observe
anticipatory behavior of the listener.</p>
      <p>The reason, why we decided to study the relation
between rapid turn-taking, comprehension and anticipation
is the lack of fluentness in case of human-machine spoken
interaction. We can observe that machines are still not
enough skilled in turn-taking.</p>
      <p>Rapid turn-taking in HMI is especially important, when
we consider the fact, that in case of humanoid robots, social
and family robots people tend to expect more natural
dialogue interaction, often called “conversation” due to
their human-like embodiment.</p>
      <p>
        Human-like turn-taking in the human-machine
interactions can be seen as a critical issue to achieve natural
conversational interaction in HMI [
        <xref ref-type="bibr" rid="ref11 ref12">11 – 12</xref>
        ]. A great view
inside this area can be found in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>Together with observation of anticipatory behavior in
human-human interactions, several questions regarding
human-machine spoken interaction arise. E.g.:</p>
      <p>Would machines let their “predictable” turns unfinished,
when they observe a comprehension on the side of human
listener? Can unfinished turns increase a natural character
of human-machine interactions?</p>
      <p>Or:</p>
      <p>Would machines be able to interrupt human speaker turn
and if yes, in which situations?</p>
      <p>How can be anticipation integrated on the side of
machine?</p>
      <p>How can machines catch the moment when the listener
has enough information to comprehend?</p>
      <p>And finally:</p>
      <p>Could be helpful if the machine would be able to
interrupt the human interlocutor?</p>
      <p>To answer proposed questions, we decided to collect
human-human interactions to analyze turn-taking, overlaps
and anticipatory behavior.</p>
      <p>The paper is organized as follows: The second section
deals with anticipation in human-human spoken interaction,
description of the prepared corpus and results of its
analysis. Section 3 provides discussion of applications,
where anticipation can bring benefits.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Anticipation</title>
      <p>interactions
in
human-human
spoken</p>
      <p>
        Anticipation play a very important role in the
humanhuman as well as human-machine spoken interaction,
because it influences the speed of response in conversation
(see [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) and it is a critical ability to enable “rapid”
turntaking. Anticipation allows the speaker to interpret partial
utterances. Sagae et al highlights the importance of
anticipation, when they conclude in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] „To achieve more
flexible turn-taking with human users, for whom
turntaking and feedback at the sub-utterance level is natural,
the system needs the ability to start interpretation of user
utterances before they are completed.“ and „it also includes
an utterance completion capability, where a virtual human
can make a strategic decision to display its understanding
of an unfinished user utterance by completing the
utterances itself”.
      </p>
      <p>Anticipation is routinely used by human interlocutors in
dyadic interactions. Expect the exact results that indicate
anticipatory behavior, according our opinion, there can be
identified also other indicators of interlocuters anticipation:
 Unfinished turns: We believe that in case
of unfinished turns speaker considers that the
listener can anticipate remaining part of his
turn.</p>
      <p> Overlaps and interruptions: We believe
that the listener can try to take floor, when he
has come to the comprehension before the end
of the speaker turn.</p>
      <p>
        Interruptions differ from overlaps in the timing. Listener
usually use a small pause in the speaker turn to interrupt the
speaker and to take a floor without causing an overlap.
Such pauses are usually marked as “Transition relevant
places” (TRP) as defined by Sacks et al. in [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. They can
also indicate the place, where the anticipation core [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] or
“moment of the maximum understanding” [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is located.
We cannot claim that each TRP is the place, where a
listener is able to anticipate. It could be interesting to
research, whether places, where anticipation core is located
can be marked as TRP.
      </p>
      <p>Unfinished turns and occurrence of overlapping segments
can have several other reasons expect comprehension based
on anticipation. Therefore, the careful analysis needs to be
done.</p>
      <p>
        The mutual understanding or comprehension on the
listener side can be indicated by backchannel signals,
which are usually produced by human listeners.
Backchannel signals was defined by Yngve [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] as an
acoustic and visual signals provided during the speaker’s
turn. Allwood et al. and Poggi in [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], described
meaning of acoustic and visual feedback that they provide
information about the basic communicative functions, as
perception, attention, interest, understanding, attitude (e.g.,
belief, liking) and acceptance towards what the speaker is
saying. Bevacqua et al. in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] defined some associations
between the listener’s communicative functions and a set of
backchannel signals. They performed an experiment with
the 3D Embodied Agent Greta [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], which confirm defined
associations. In the described experiment it has been
shown, that there exist the association between
understanding and following multimodal backchannel
signals: raise eyebrows+“ooh”, head nod+“ooh”, head
nod+“really”, head nod+“yeah” and head nod.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Corpus</title>
        <p>Corpus of the investigative interviews of the TV program
“Na rovinu” and episodes of TV discussions “Pod lampou”
were selected for the analysis of turn-taking mechanisms.</p>
        <p>The overall length of the analyzed corpus was approx. 8
hours. There were 9 speakers (8 males, 1 female), two of
them play a role of the moderator (male).</p>
        <p>
          Recordings processing have two main levels. In the first
level, recordings were transcribed. The first transcriptions
were generated by our automatic transcription system [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ].
Then, transcriptions were corrected manually.
Transcriptions have a form of .trs files, which are generated
by Transcriber tool. At the second level, turns and overlaps
of both interlocutors were marked in Anvil annotation tool,
what enables to analyze turn-taking.
        </p>
        <p>In the first step of the analysis we focused mainly on
overlapping segments. To classify different types of
overlaps, we designed 13 categories according intent
behind the overlaps (see Fig. 2.)</p>
        <p>Two basic categories – concurrent and cooperative
overlaps are further divided according their function.</p>
        <p>Concurrent/competitive overlaps represent overlaps,
where we can identify an attempt to grab the floor. On the
other side cooperative/non-competitive category mark
overplaps, where listener’s goal is to assist the speaker to
continue his turn.
before the end of the speaker turn, what causes overlapping
speech.</p>
        <p>More than 84% of overlaps were caused by the
moderator, who jumped into respondent speech.
Respondent jumped into moderator turns approx. in 16% of
all cases. The explanation may be that the moderator has a
responsibility for interaction, and he need to maintain topic,
topic shifts and timing (duration) of the interaction.</p>
        <p>Table in Fig.3 shows the distribution of intentions behind
the overlaps, which we detected manually.</p>
        <p>According obtained results, we can conclude significant
differences between moderator and respondent. On the side
of the moderator, the most numerous categories were:
supporting the speaker (approx. 32%), request for
completing information or providing additional information
(around 20%) and asking the additional question (18%).</p>
        <p>On the side of the respondent, the most numerous
categories imagine overlaps, which occur when a listener
expressed involvement and agreement (more than 57%).</p>
        <p>Obtained results show that roles, which play
interlocutors affects the number of overlaps and also
intentions behind attempts to take the floor.</p>
        <p>Next issue is that designed categories contain several
items, which can be identified as backchannel signals
(express involvement/agreement/disagreement, express
noninterest). As was concluded above, backchannel signals
indicate rather partial understanding, than anticipatory
behavior. However, information about this dialogue parts is
important too, to analyze backchannel mechanisms in
different scenarios and roles, which interlocutor plays.</p>
        <p>
          Designed categories of overlaps correspond wih
speech/dialogue acts, which are conveyed through
utterances that interrupt the speaker turn. Confirmation
request, Clarification request, Agreement, Disagreement
are typical dialogue acts (see e.g. [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]). On the other hand,
categories like “Stop the speaker”, “Support speaker”,
“Move to another topic”, “Try to take floor” can be seen as
a turn-management commands, which help to manage
speaker changing.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Anticipation and its practical application</title>
      <p>2.2</p>
      <sec id="sec-3-1">
        <title>Results</title>
        <p>Total number of both speakers performed turns was
1646. In 34.8% of cases, turns were taken by the listener
The main reason, why anticipation needs to be
considered is to support rapid turn-taking, which enables
smooth and fluent human-machine spoken interaction. But
anticipation can enable also other human-like capabilities
in HMI.</p>
        <p>
          In case of application anticipation in human-machine
dialogue interactions, there exists only few works that deals
with this phenomenon. We can mention the work of
Dominey et al., described in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], which focuses on next turn
anticipation, based on dialogue history, but in this case the
anticipation is not focused inside the turn. Sagae et al. focus
their work on interpretation of partial utterances, where
they are using words prediction – anticipation. They realize
an idea of the incremental user utterance processing, which
enables to increase speed of turn-taking (switching the
speaker).
        </p>
        <p>The anticipatory behavior on the machine side means the
ability to find an enough reliable hypothesis of unfold
utterance just uttered by human interlocutor. Machines that
can communicate with user through spoken language
implements a human machine communication chain.
Modules of this chain have anticipatory potential, because
they use resources with related data, as are recognition
network, language models or dialogue model.
3.1</p>
      </sec>
      <sec id="sec-3-2">
        <title>Applications</title>
        <p>While, the anticipatory behavior can be considered as
non-important for simple task-oriented dialogue systems, it
can bring more human-like character of the interaction and
advantages in case of systems for multi-party
conversations, for human-machine collaboration scenarios
or for crisis scenarios, where can be useful or necessary for
the machine to be able to rudely interrupt a human
speaker(s), to be able to take a floor and to propose ideas or
solutions.
3.1.1</p>
        <sec id="sec-3-2-1">
          <title>Machine in a role of a moderator</title>
          <p>Machine in the role of discussion moderator can be the
next example of application, where anticipatory behavior
can play important role. Emotionally colored multi-party
spoken interactions as are e.g. political discussions often
require a lot of effort on the side of moderator to lead the
interaction and to manage turn-taking for all speakers.
Moreover, he often needs to enforce good behavior or
compliance with specified time.
3.1.2</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Machine in a task of simultaneous interpretation</title>
          <p>
            Anticipating in simultaneous interpreting simply means
that interpreters say a word or a group of words before the
speaker actually says them [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ]. It means that interpreters,
familiar with the domain and content, can interpret
beforehand thank to the anticipation. Anticipation is a key
competence that interpreters need to learn before they can
become professionals [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ], [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ].
          </p>
          <p>Nowadays, machine translation is a common application.
There also exist applications, which perform interpretation,
but they interpret whole sentences after their pronunciation.
But, if we imagine machine as the simultaneous interpreter,
it must be able to anticipate, to produce fluent simultaneous
interpretations without meaningless gaps. In that case,
anticipation can enhance perceived quality and clarity of
the translation.</p>
          <p>Anticipatory function can help also human interpreters in
the way, that it can suggest them hypothesis about next
words in the speaker utterance. Architecture of such
machine-supported simultaneous interpretation system is
sketched on the Fig. 4.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>Machine-supported simultaneous Figure 4 interpretation</title>
        <p>One of the scenarios of machine-supported simultaneous
interpretation is that the system will listen to the speaker
and try to provide early predictions of the next words
according just recognized partial utterance. Speaker speech
will be continually recognized by Automatic Speech
Recognition (ASR) module, where recognition hypothesis
will be provided as often as possible. Each recognition
hypothesis will serve as the input into Hypothesis
construction module, in which, next words will be
predicted according general, speaker-dependent and topic
models. Statistical n-gram language models/DNN networks
that are used in ASR module can also serve for final
hypothesis generation.</p>
        <p>The same scenario can be used also on the side of the
interpreter, where the system can suggest next words of his
interpretation.</p>
        <p>
          Especially interesting can be an automatic transfer of
emotions, which can help interpreter to choose the most
appropriate words, which consider emotional coloring of
interpreted speech. Work proposed in [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] can be used for
desired emotion transfer.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>The aim of the paper is to stimulate discussion on the use
of anticipation and anticipatory behavior in the HMI and in
the practical applications. Certainly, anticipation lies
behind the smooth human-human interactions and rapid
turn-taking and we believe that it can significantly
accelerate human-machine spoken interaction and to
support human-like character of the HMI.</p>
      <p>We believe that machine-supported anticipation can
decrease cognitive load in several applications, e.g. in
simultaneous interpretation, where such a system can
suggest hypothesis for the interpreter in advance or to
support preparation of the interpretation result.</p>
      <p>We realize that obtained results are relevant only to the
interview scenario and the distribution of analyzed overlaps
category will change according roles, relationship,
emotions of interlocutors and according discussed topics.
The first challenge of our future work will be collecting and
analysing of dialogue interactions in other scenarios.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgment</title>
      <p>The research presented in this paper was supported by
the Slovak Research and Development Agency projects
APVV SK-TW-2017-0005, APVV-15-0731, Ministry of
Education, Science, Research and Sport of the Slovak
Republic under the research project VEGA 1/0511/17 and
by Cultural and Educational Grant Agency of the Slovak
Republic, grant No. KEGA 009TUKE-4/2019.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>[1] https://www.teachingenglish.org.uk/article/turn-taking</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Magyari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.C.</given-names>
            <surname>Bastiaansen</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.P. de Ruiter</surname>
            , and
            <given-names>S.C.</given-names>
          </string-name>
          <string-name>
            <surname>Levinson</surname>
          </string-name>
          , “
          <article-title>Early Anticipation Lies behind the Speed of Response in Conversation”</article-title>
          ,
          <source>Journal of Cognitive Neuroscience</source>
          , pp.
          <fpage>2530</fpage>
          -
          <lpage>2539</lpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.S.</given-names>
            <surname>Gisladottir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bögels</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.C.</given-names>
            <surname>Levinson</surname>
          </string-name>
          , “
          <article-title>Oscillatory Brain Responses Reflect Anticipation during Comprehension of Speech Acts in Spoken Dialog”</article-title>
          ,
          <source>Front. Hum. Neurosci</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Liddicoat</surname>
          </string-name>
          , “
          <article-title>The projectability of turn constructional units and the role of prediction in listening”</article-title>
          ,
          <source>Discourse Studies</source>
          , Vol.
          <volume>6</volume>
          , No.
          <issue>4</issue>
          , pp.
          <fpage>449</fpage>
          -
          <lpage>469</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.F.</given-names>
            <surname>Dominey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Metta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Nori</surname>
          </string-name>
          , and L. Natale, “
          <article-title>Anticipation and initiative in human-humanoid interaction“</article-title>
          ,
          <source>Humanoids 2008 - 8th IEEE-RAS International Conference on Humanoid Robots, Daejeon</source>
          , pp.
          <fpage>693</fpage>
          -
          <lpage>699</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Paľová</surname>
          </string-name>
          , E. Kiktová “
          <article-title>Prosodic anticipatory clues and reference activation in simultaneous interpretation”</article-title>
          ,
          <source>in XLinguae 12(1XL)</source>
          , pp.
          <fpage>13</fpage>
          -
          <lpage>22</lpage>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Stivers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Enfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Brown</surname>
          </string-name>
          , C. Englert,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heinemann</surname>
          </string-name>
          , et al. (
          <year>2009</year>
          ).
          <article-title>Universals and cultural variation in turn-taking in conversation</article-title>
          .
          <source>Proceedings of the National Academy of Sciences, U.S.A.</source>
          ,
          <volume>106</volume>
          ,
          <fpage>10587</fpage>
          -
          <lpage>10592</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Indefrey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. J. M.</given-names>
            <surname>Levelt</surname>
          </string-name>
          , “
          <article-title>The spatial and temporal signatures of word production components”</article-title>
          .
          <source>Cognition</source>
          ,
          <volume>92</volume>
          ,
          <fpage>101</fpage>
          -
          <lpage>144</lpage>
          ,
          <year>2004</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T. T.</given-names>
            <surname>Schnurr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Costa</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Caramazza, “
          <article-title>Planning at the phonological level during sentence production</article-title>
          .”,
          <source>Journal of Psycholinguistics Research</source>
          ,
          <volume>35</volume>
          ,
          <fpage>189</fpage>
          -
          <lpage>213</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Holler</surname>
          </string-name>
          ,
          <string-name>
            <surname>K</surname>
          </string-name>
          , H. Kendrick,
          <string-name>
            <given-names>M.</given-names>
            <surname>Casillas</surname>
          </string-name>
          , S. C. Levinson, eds. (
          <year>2016</year>
          ).
          <article-title>Turn-Taking in Human Communicative Interaction</article-title>
          .
          <source>Lausanne: Frontiers Media. doi: 10.3389/978-2-88919-825-2</source>
          ,
          <fpage>2016</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Thórisson</surname>
          </string-name>
          , “
          <article-title>Natural turn-raking needs no manual: computational theory and model, from perception to action,” in Multimodality in Language and Speech Systems</article-title>
          , eds B.
          <string-name>
            <surname>Granström</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>House</surname>
            ,
            <given-names>and I. Karlsson</given-names>
          </string-name>
          (Netherlands: Springer),
          <volume>173</volume>
          -
          <fpage>207</fpage>
          ,
          <year>2002</year>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Heldner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Edlund</surname>
          </string-name>
          , “
          <article-title>Pauses, gaps and overlaps in conversations</article-title>
          .
          <source>” J.Phon</source>
          .
          <volume>38</volume>
          ,
          <fpage>555</fpage>
          -
          <lpage>568</lpage>
          ,
          <year>2010</year>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <article-title>Anticipation in simultaneous interpreting</article-title>
          , https://www.languageconnections.com/,
          <year>March 2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>V.</given-names>
            <surname>Yngve</surname>
          </string-name>
          , “
          <article-title>On getting a word in edgewise”</article-title>
          ,
          <source>in Papers from the Sixth Regional Meeting of the Chicago Linguistic Society</source>
          , pp.
          <fpage>567</fpage>
          -
          <lpage>577</lpage>
          ,
          <year>1970</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Allwood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nivre</surname>
          </string-name>
          , and E. Ahlsn, “
          <article-title>On the semantics and pragmatics of linguistic feedback”</article-title>
          ,
          <source>Semantics</source>
          Vol.
          <volume>9</volume>
          , No.
          <volume>1</volume>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>I. Poggi</surname>
          </string-name>
          , “
          <article-title>Mind, hands, face and body. A goal and belief view of multimodal communication”</article-title>
          , Weidler, Berlin,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Bevacqua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pammi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.J.</given-names>
            <surname>Hyniewska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schröder</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Pelachaud</surname>
          </string-name>
          , “
          <article-title>Multimodal Backchannels for Embodied Conversational Agents”</article-title>
          , in: Intelligent Virtual Agents,
          <source>IVA 2010. Lecture Notes in Computer Science</source>
          , Vol.
          <volume>6356</volume>
          . Springer, Berlin, Heidelberg,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>R.</given-names>
            <surname>Niewiadomski</surname>
          </string-name>
          , E. Bevacqua,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mancini</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Pelachaud</surname>
          </string-name>
          , “
          <article-title>Greta: an interactive expressive eca system”</article-title>
          ,
          <source>in: AAMAS 2009 - Autonomous Agents and MultiAgent Systems</source>
          , Budapest, Hungary,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>K.</given-names>
            <surname>Sagae</surname>
          </string-name>
          , D. DeVault, and
          <string-name>
            <given-names>D.R.</given-names>
            <surname>Traum</surname>
          </string-name>
          , “
          <article-title>Interpretation of partial utterances in virtual human dialogue systems”</article-title>
          ,
          <source>Proc. of the NAACL HLT</source>
          <year>2010</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Stroudsburg, PA, USA, pp.
          <fpage>33</fpage>
          -
          <lpage>36</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kiktova</surname>
          </string-name>
          , J. Zimmermann, “
          <article-title>Detection of Anticipation Nucleus using HMM and Fuzzy based Approaches”, (in review process</article-title>
          )
          <source>DISA</source>
          <year>2018</year>
          ,
          <article-title>August</article-title>
          , Košice, Slovakia
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>K. G.</given-names>
            <surname>Seeber</surname>
          </string-name>
          , “Intonation and Anticipation in Simultaneous Interpreting”, Cahiers de Linguistique Française, vol.
          <volume>23</volume>
          ,
          <year>2001</year>
          , pp.
          <fpage>61</fpage>
          -
          <lpage>97</lpage>
          , ISSN 1661-
          <issue>3171</issue>
          ,
          <year>2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ondáš</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Juhár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pleva</surname>
          </string-name>
          , et. al.:
          <article-title>“Speech technologies for advanced applications in service robotics”</article-title>
          ,
          <source>in Acta Polytechnica Hungarica</source>
          . Vol.
          <volume>10</volume>
          , no.
          <issue>5</issue>
          (
          <issue>2013</issue>
          ), p.
          <fpage>45</fpage>
          -
          <lpage>61</lpage>
          ., ISSN
          <volume>1785</volume>
          -
          <fpage>8860</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mikula</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Machová</surname>
          </string-name>
          , “
          <article-title>Combined approach for sentiment analysis in Slovak using a dictionary annotated by particle swarm optimization</article-title>
          .”,
          <source>in: Acta Electrotechnica et Informatica</source>
          , Vol.
          <volume>18</volume>
          , No.
          <volume>2</volume>
          ,
          <year>2018</year>
          ,
          <fpage>27</fpage>
          -
          <lpage>34</lpage>
          , DOI: 10.15546/aeei-2018
          <source>-0013</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ondáš</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Juhár</surname>
          </string-name>
          ,
          <article-title>"Distance-based dialog acts labeling,"</article-title>
          <source>2015 6th IEEE International Conference on Cognitive Infocommunications (CogInfoCom)</source>
          , Gyor,
          <year>2015</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Lojka</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viszlay</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hladek</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juhar</surname>
          </string-name>
          , J.:
          <article-title>Slo-vak Broadcast News Speech Recognition and TranscriptionSystem</article-title>
          . International Conference on
          <article-title>Network-Based Infor-mation Systems</article-title>
          . In: Barolli L.,
          <string-name>
            <surname>Kryvinska</surname>
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Enokido</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takizawa</surname>
            <given-names>M</given-names>
          </string-name>
          .
          <article-title>(eds) Advances in Network-Based Informa-tion Systems</article-title>
          .
          <source>NBiS 2018. Lecture Notes on Data Engineer-ing and Communications Technologies - LNDECT</source>
          , vol
          <volume>22</volume>
          .Springer, Cham, pp.
          <fpage>385</fpage>
          -
          <lpage>394</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>I. Siegert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wendemuth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vlasenko</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ohnemus</surname>
          </string-name>
          ,
          <article-title>"Overlapping speech, utterance duration and affective content in HHI and HCI - An comparison,"</article-title>
          <source>2015 6th IEEE International Conference on Cognitive Infocommunications (CogInfoCom)</source>
          , Gyor,
          <year>2015</year>
          , pp.
          <fpage>83</fpage>
          -
          <lpage>88</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>H.</given-names>
            <surname>Sacks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Schegloff</surname>
          </string-name>
          , and G. Jefferson, “
          <article-title>A simplest systematics for the organization of turn-taking for conversation</article-title>
          ,
          <source>” Language</source>
          , vol.
          <volume>50</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>696</fpage>
          -
          <lpage>735</lpage>
          , Dec.
          <year>1974</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>