<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluating Cognitive Bias in Two-P arty and Multi-P arty Spoken Interactions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christina Alexandris</string-name>
          <email>calexandris@gs.uoa.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National and Kapodistrian University of Athens</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Targeting to by -pass Cognitive Bias in two-party discussions and interviews containing longer speech segments, a proposed semi-automatic procedure involves “takingthe temperature” of a transcribed dialog by measuring the number of detected points of possible tension and/or conflict between speakers-participants.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Human Computer Interaction (HCI) systems may assist in
the evaluation of complex Human-Human interaction, as in
the case of designed applications for journalists
        <xref ref-type="bibr" rid="ref2">(Alexandris ,
Nottas and Cambourakis, 2015)</xref>
        .
      </p>
      <p>In spoken dialogs concerning complex interactions
between speakers-participants -as in the case of spoken
journalistic texts, there are aspects that can be evaluated by
semi-automatic or interactive procedures, targeting to by
pass Cognitive Bias and there are aspects that can be
evaluated by interactive procedures, targeting to register
Cognitive Bias.</p>
      <p>
        Unlike task-specific dialogs
        <xref ref-type="bibr" rid="ref31">(Tung et al., 2013)</xref>
        and
typical collaborative dialogs
        <xref ref-type="bibr" rid="ref33 ref38">(Yang, Levow and Meng, 2012,
Wang et al., 2013)</xref>
        , the Speech Acts performed by one or
multiple speakers -participants, often may involve comple x
Illocutionary Acts beyond the defined framework of the
interaction. Specifically, the Illocutionary Act
        <xref ref-type="bibr" rid="ref25 ref5">(Searle,1969,
Austin, 1962)</xref>
        performed by the Speaker concerned may not
be restricted to “Obtaining Information Asked” or
“Providing Information Asked” in a discussion or interview:
Speakers-participants may have other or additional intentions
regarding their presence and their role in the discussion or
interview concerned. In the spoken journalistic texts
concerned, Illocutionary Acts not restricted to “Obtaining
Information Asked” or “Providing Information Asked ” are
related to other or additional Speaker intentions. For example ,
a Speaker may focus in emphasizing opinion (or the policy
of the network concerned) or in (purposefully) creating
tension in the interview or discussion. Furthermore, a
consistent avoidance of the topics addressed may indicate that
the Speaker is more interested in showing a mere presence
in the discussion or interview, rather than sharing any
information.
      </p>
      <p>
        The existence of additional, “hidden” Illocutionary Acts
can be identified, by procedures evaluating the behavior of
speakers-participants in relation to specific values and
benchmarks. The presentation and calculation of these
values allows the possibility of by-passing or registering
Cognitive Bias. The Cognitive Bias by-passed or registered
concerns primarily the evaluation of Cognitive Bias of (i) the
speakers-participants concerned but may also serve for the
evaluation of the Confidence Bias of (ii) the user-evaluator
of the recorded and transcribed discussion or interview.
In smaller speech segments with constant and quick change
of speaker turns and with discourse structure compatible to
models where each participant selects self
        <xref ref-type="bibr" rid="ref19 ref23 ref37">(Wilson, 2005,
Sacks, Schegloff, and Jefferson 1974)</xref>
        , topic tracking (and
topic change) allows the evaluation of speaker behavior and
enables the identification of speaker’s intentions and
Illocutionary Speech Acts performed
        <xref ref-type="bibr" rid="ref25 ref5">(Searle,1969, Austin, 1962)</xref>
        .
Topic tracking can be applied especially in short speech
segments with two or multiple speakers -participants (Alexan
dris, 2018). The content of relatively short utterances can be
summarized with the use of keywords chosen from each
utterance by the user-evaluator
        <xref ref-type="bibr" rid="ref1">(Alexandris, 2018)</xref>
        , with the
assistance of the Stanford POS Tagger for the automatic
signalization of nouns in each turn taken by the Speakers in the
respective segment in the dialog structure. The registered
and tracked keywords, treated as local variables, signalize
each topic and the relations between topics, since automatic
Rhetorical Structure Theory (RST) analysis proce
        <xref ref-type="bibr" rid="ref10">dures
(Stede, Taboada and Das, 2017</xref>
        ,
        <xref ref-type="bibr" rid="ref41">Zeldes, 2016</xref>
        ) usually
involves larger (written) texts and may not produce the
required results.
      </p>
      <p>
        The System generates a visual representation from the
user’s interaction, tracking the corresponding selected
topickeywords in the dialog flow, as well as the chosen types of
relations between them. The interactive generation of
registered paths is similar to the paths with generated sequences
of recognized keywords in spoken dialog systems, in the
domains of consumer complaints and mobile phone services
call centers
        <xref ref-type="bibr" rid="ref11">(Nottas et al., 2007, Floros and Mourouzidis,
2016)</xref>
        . This function is similar to user-independent
evaluations of spoken dialog systems
        <xref ref-type="bibr" rid="ref36">(Williams, Asadi and Zweig,
2017)</xref>
        for by-passing User bias
        <xref ref-type="bibr" rid="ref19 ref8">(Nass and Brave, 2005,
Cohen, 1997)</xref>
        . Keywords (topics) may be repeated or related to
a more general concept (or global variable)
        <xref ref-type="bibr" rid="ref17">(Lewis, 2009)</xref>
        or
related to keywords (topics) concerning similar functions
(corresponding to the Repetition, Generalization and
Association relations respectively and the visual representations
of Distances 1 (value “1”),2 (value “2”) and 3 (value “3”)
respectively)
        <xref ref-type="bibr" rid="ref1">(Alexandris, 2018)</xref>
        . A keyword involving a
new command or function is registered as a new topic (New
Topic, visual representation of Distance 4, corresponding to
value: “0”). The sequence of topics chosen by the user and
the perceived relations between them generates a “path” of
interaction, forming distinctive visual representations stored
in a database currently under development: Topics and
words generating diverse reactions and choices from users
result to the generation of different forms of generated
visual representations for the same conversation and
interaction
        <xref ref-type="bibr" rid="ref1">(Alexandris, 2018)</xref>
        .
      </p>
      <p>
        The generated visual representations depict topics
avoided, introduced or repeatedly referred to by each
speaker-participant, and in specific types of cases may
indicate the existence of additional, “hidden” Illocutionary Acts
other than “Obtaining Information Asked” or “Providing
Information Asked” in a discussion or interview. Thus, the
evaluation of speaker-participant behavior targets to by-pass
Cognitive Bias, specifically, Confidence Bias
        <xref ref-type="bibr" rid="ref15">(Hilbert ,
2012)</xref>
        of the user-evaluator, especially if multiple users
evaluators may produce different forms of generated visual
representations for the same conversation and interaction
and compared to each other in the database. In this case,
chosen relations between topics may describe Lexical Bias,
        <xref ref-type="bibr" rid="ref29">(Trofimova, 2014)</xref>
        and may differ according to political,
socio-cultural and linguistic characteristics of the
user-evaluator, especially if international users are concerned
        <xref ref-type="bibr" rid="ref22 ref39 ref4">(Yu et al.,
2010, Alexandris, 2010, Ma, 2010, Pan, 2000)</xref>
        due to to lack
of world knowledge of the language community involved
        <xref ref-type="bibr" rid="ref14 ref21 ref35">(Paltridge, 2012, Hatim, 1997, Wardhaugh, 1992)</xref>
        . The
envisioned further development of generated visual
representations is their modeling in a form of graphs, similar to
discourse trees
        <xref ref-type="bibr" rid="ref7">(Marcu, 1999, Carlson, Marcu and Okurowski,
2001)</xref>
        .
      </p>
    </sec>
    <sec id="sec-2">
      <title>Evaluation and Be nchmarks</title>
      <p>The types of relations-distances between word-topics
chosen by the user-evaluator are registered and counted. If the
number of (a) the “Repetitions” label or (b) the number of
the “Generalizations” or (c) the number of the “Topic
Switches” exceeds well over 50% of the registered
relationsdistances between word-topics, the interaction is signalized
for further evaluation, containing Illocutionary Acts not
restricted to “Obtaining Information Asked” or “Providing
Information Asked”. The following benchmarks indicate
interactions with Illocutionary Acts beyond the predefined
framework of the dialog for multiple Speaker discussions
and/or short speech segments, where Ds = Number of
Distances and Sp = Number of Speaker turns:
• X= Ds ≤ Sp (calculating over 50% of “Repetitions” (Dis
tance = 1, value “1”) ) or “Topic Switches” (Distance =
4, value “0”).
• X= Ds &gt; Sp × Gen (Gen = Sp × 3 ÷ 2) (calculating over
50% of “Generalizations” (Distance = 3, value “3”).
These benchmarks for dialogs with short speech segments
can be referred to as “(Topic) Relevance” benchmarks with
a value of “X” or “Relevance (X)”.</p>
    </sec>
    <sec id="sec-3">
      <title>By-Passing and Registering Cognitive Bias in</title>
    </sec>
    <sec id="sec-4">
      <title>Two-Party Discussions and Interviews or in</title>
    </sec>
    <sec id="sec-5">
      <title>Long Speech Segments</title>
      <p>
        The further development of the database containing
registered spoken interaction for determining and evaluating
Cognitive Bias in spoken journalistic texts
        <xref ref-type="bibr" rid="ref1">(Alexandris ,
2018)</xref>
        involved the processing of discussions and interviews
containing larger speech segments. Similarly to the
abovedescribed multiple speaker discussions and in short speech
segments, the Illocutionary Act performed by the Speaker
concerned may not be restricted to “Obtaining Information
Asked” or “Providing Information Asked” in a discussion or
interview.
      </p>
      <p>
        In two-party discussions and interviews containing longer
speech segments, the discourse structure is more compatible
to turn-taking in “push-to-talk conversations”, with a strict
protocol in managing the interview or discussion and turn
taking
        <xref ref-type="bibr" rid="ref26">(Taboada, 2006)</xref>
        . In this case, speakers-participants
usually not have the liberty of modifying or changing the
topic, resulting to the strategy of topic tracking being
insufficient for the identification of speaker’s intentions. In larger
speech segments mostly occurring in interviews with a strict
protocol and a set of predefined topics, automatic Rhetorical
Structure Theory (RST) analysis proce
        <xref ref-type="bibr" rid="ref10">dures (Stede,
Taboada and Das , 2017</xref>
        ,
        <xref ref-type="bibr" rid="ref41">Zeldes, 2016</xref>
        ) can be performed in
the transcribed text, with the condition that the speaker is
allowed sufficient time to elaborate on the topic in question.
The extent to which automatic RST analysis procedures can
be executed in the transcribed text indicates the degree of
collaborative interaction between the speakers -participants,
especially from the journalist-interviewer (referred to as
Speaker 1), since the speaker-participant is allocated enough
time to elaborate and/or argument on the topic co ncerned.
      </p>
      <p>In the case of discussions and interviews containing larger
speech segments, the identification of speaker’s intentions
and “hidden” Illocutionary Act detection follows a process
locating points of possible tension and/or conflict between
speakers-participants. In points of possible tension and/or
conflict between speakers -participants, Cognitive Bias can
both be by-passed or registered. Cognitive Bias is by-passed
by signalizing and counting the points of possible tension
and/or conflict between speakers-participants henceforth
referred to as “hotspots”. The signalization of “hotspots” is
based on the violation of the Quantity, Quality and Manner
Maxims of the Gricean Cooperativity Principle (Gric e ,
1975). Cognitive Bias is registered by comparing content of
the Speaker turns in the signalized “hotspots” and assigning
a respective value.</p>
    </sec>
    <sec id="sec-6">
      <title>By-Pas s ing Cognitive Bias: Automatic Signaliza</title>
      <p>tion of “Hots pots” and the Grice an Cooperativity</p>
    </sec>
    <sec id="sec-7">
      <title>Principle</title>
      <p>Targeting to by-pass Cognitive Bias in two-party
discussions and interviews containing longer speech segments , a
proposed semi-automatic procedure involves “taking the
temperature” of a transcribed dialog by measuring the
number of detected points of possible tension and/or conflict
between speakers-participants. These points are henceforth,
referred to as “hot spots” and concern in speech segments
where there is a recognition of speaker turns, namely a
switch between Speaker 1 and Speaker 2 by the Speech
recognition module of the transcription tool. The signaliza
tion of multiple “hot spots” indicates a more argumentative
than a collaborative interaction, even if speakers
-participants display a calm and composed behavior. In particular,
the Illocutionary Act performed by the Speaker concerned
may not be restricted to “Obtaining Information Asked” or
“Providing Information Asked” in a discussion or interview.</p>
      <p>
        A “hot spot” consists of the pair of utterances of both
speakers, namely a question-answer pair or a
statement-response pair or any other type of relation between speaker
turns. In longer utterances, the first 60 words of the second
speaker’s (Speaker 2) utterance are processed
(approximately 1 -3 sentences, depending on length, with the
average sentence length of 15-20 words,
        <xref ref-type="bibr" rid="ref9">(Cutts 2013)</xref>
        and the
last 60 words of the first speaker’s (Speaker 1) utterance are
processed (approximately 1 -3 sentences, depending on
length). The automatically signalized “hot spots” are
extracted to a separate template for further processing. The
extraction contains not only the detected segments but also the
complete utterances consisting of both speaker turns of
Speaker 1 and Speaker 2.
      </p>
      <p>
        For a segment of speaker turns to be automatically
identified as a “hot spot”, at least two of the following three
conditions (1), (2) and (3) must apply to one or to both of the
speaker’s utterances, of which conditions (1), (2) are
directly or indirectly related to flouting of Maxims of the
Gricean Cooperative Principle
        <xref ref-type="bibr" rid="ref12">(Grice, 1975)</xref>
        . These conditions
are the following:
• (1) Additional, modifying features: In one or in both
speakers’ utterances in the segment of speaker turns there
is at least one phrase containing a sequence of two
adjectives (ADJ ADJ) (a) or an adverb and an adjective (or
more adjectives) (b) (ADV ADJ) or two adverbs (ADV
ADV) (c). These forms of adjectival or adverbial phrases
are detectable with a POS Tagger (for example, the
Stanford POS Tagger.
• (2) Reference to the interaction itself and to its
participants with negation. In one or in both speakers’
utterances, the subject of the sentence containing the negation
is “I” or “you” ((I/You) “don’t”, “do not”,“cannot”) (a)
and in the verb phrase (VP) there is at least one speech
related or behavior verb-stem referring to the dialog itself
(b) (for example, “speak”, “listen”, “guess”,
“understand”). This applies to parts of speech other than verbs
(i.e. “guessing”, “listener”) as well as to words
constituting parts of expressions related to speech or behavior
(“conclusions”, “words”, “mouth”, “polite”, “nonsense”,
“manners”). The different forms of negation are
detectable with a POS Tagger. The respective words and word
categories may constitute a small set of entries in a
specially created lexicon or may be retrieved from existing
databases or WordNets .
• (3) Prosodic emphasis and/or Exclamations. (a) Excla ma
tions include expressions such as such as “Look”, “Wait”
and “Stop”. As in the above-described case (2), the
respective words and word categories may constitute a
small set of entries in a specially created lexicon or may
be retrieved from existing databases or WordNets. (b)
Prosodic emphasis, detected in the speech processing
module, may occur in one or more of the above-described
words of categories (1a, 1b, 1c, 2a and 2b) or in the noun
or verb following (modified by) 1a, 1b and 1c.
      </p>
      <p>
        In the case of 1a, 1b and 1c, there is extra information added
to the basic content of the utterance consisting the necessary
information required to fulfil the Gricean Cooperative Prin
ciple in respect to the Maxim of Quantity. (“Do not make
your contribution more informative than is required “).
Here, the Speaker violates the Maxim of Quantity in the
Gricean Cooperative Principle. In the case of 2a and 2b, the
Speaker perceives a violation of the Gricean Cooperative
Principle by the previous Speaker. In particular, the content
of the speaker’s utterance is not limited to the current topic
in question but refers to the dialog itself, mostly functioning
as a comment. Specifically, 2a and 2b imply a violation of
the Gricean Cooperative Principle in respect to the Maxim
of Quality (“1. Do not say what you believe to be false”, “2.
Do not say that for which you lack adequate evidence”)
        <xref ref-type="bibr" rid="ref12">(Grice, 1975)</xref>
        and/or in respect to the Maxim of Manner
(Submaxim 2. “Avoid ambiguity”)
        <xref ref-type="bibr" rid="ref12">(Grice, 1975)</xref>
        in the
utterance of the previous Speaker. In other words, in 2a and
2b, the Speaker considers the content of the previous
Speaker’s utterance to be unacceptable, ambiguous, false or
controversial.
      </p>
      <p>The number of automatically signalized ”hot spots”
indicates the degree in which discussions and interviews
containing larger speech segments constitute dialog with many
points of tension and/or conflict. The average time of
discussions and interviews containing larger speech segments
in the Media is 30 to 45 minutes (30-45 mins). A typical
example of a dialog with many detected points of possible
tension and/or conflict between speakers -participants is an
approximately 32 minute long interview with seven (7)
registered “hot spots” (BBC (British Broadcasting
Corporation): HARDtalk interview by journalist Stephen Sackur on
16th of April 2018). One or both speakers’ utterances may
display two or more of features (1), (2) and (3).</p>
    </sec>
    <sec id="sec-8">
      <title>Evaluation and Be nchmarks</title>
      <p>The benchmark for evaluating a remarkable degree of
tension in a discussion is signalized by multiple “hotspots”
detected and not sporadic occurrences of “hotspots”. Thus, the
number of 1-2 “hotspot” occurrences in longer speech
segments in question (30-45 mins) signalizes a low degree of
tension. A remarkable degree of tension in a 30-45 minute
discussion or interview is related to a number of at least 4
detected “hotspots” (where the number of 3 hotspots
constitutes a marginal value). Considering the above, th e
benchmark for evaluating a remarkable degree of tension concerns
the calculation of the time of discussion / interview in the
Media (for example, 35 mins) and the number of signalized
”hot spots” (SPEECH SEGMENT-count) in Speaker turns.
The defined benchmark for evaluating Speaker behavior is
the number of minutes divided by the number of identified
speech segments signalized as “hot spots” which should
contain a single digit number, if the above-described
minimal number of at least 4 detected “hotspots” is calculated.
For example, the acceptable values are “8.75”, “7” or,
ideally, “5” (for a file of 35 minutes ) versus “17.5” or “11.6”
(for a file of 35 minutes). Interactions with Illocutionary
Acts beyond the predefined framework of the dialog
-discussion (with two speakers –participants and long speech
segments) are based on the detected points of possible tension
and/or conflict indicated by the following benchmark, where
Y = wav file length in minutes divided by (÷) the number of
“hot spot” signalized speech segments :
• Y &lt; 10.
• Example: File length = 35 mins, SPEECH SEGM ENT
count: 5, Evaluation: 7.</p>
      <p>These benchmarks for dialogs with long speech segments
can be referred to as “Tension” benchmarks with a value of
“Y” or “Tension (Y)”.</p>
    </sec>
    <sec id="sec-9">
      <title>Registering Cognitive Bias: Interactive Comparison of Speaker Turns</title>
      <p>
        The registration of Cognitive Bias concerns the comparison
of the actual content of the pair of utterances of Speaker 1
and Speaker 2 in the signalized “hot spots”. As stated above,
the automatically signalized ”hot spots” are extracted to a
separate template for interactive processing, where the “hot
spot” utterances of both speakers are compared. If the last
60 words (approximately 1 -3 sentences, with average
sentence length of 15-20 words,
        <xref ref-type="bibr" rid="ref9">(Cutts 2013)</xref>
        of the first
speaker’s utterance contain at least two of the
above-described features (1), (2) and (3), the Quantity and Manner
Maxims of the Gricean Cooperative Principle
        <xref ref-type="bibr" rid="ref12">(Grice, 1975)</xref>
        are violated. Specifically, Cognitive Bias is registered by
comparing content of the Speaker turns in the signalized
“hotspots” and assigning the following respective values :
• (a) Each “hot spot” is marked with a (1,1) if both
speakers’ utterances are considered equally non-collaborative.
• (b) If this is the case for one of the two speakers, in
particular, Speaker 1, the “hot spot” is marked with a (1,0)
for Speaker 1 (in this case, the journalist-reporter). In this
case, the style of question or statement uttered is not
considered acceptable- contains features violating the
Gricean Cooperative Principle - in respect to the Maxim of
Manner or the Maxim of Quantity (Irony) or in respect to
the Maxim of Quality (content is considered false (“F”).
• (c) If this is the case for Speaker 2, the “hot spot” is
marked with a (0,1), if the interviewee’s (Speaker 2)
reaction is not justified in respect to the style and content of
the utterance of Speaker 1.
• (d) If a “hot spot” speech segment is evaluated by the User
not as a point of possible tension and/or conflict between
speakers-participants, the false “hot spot” is marked with
a (0,0) for both Speakers.
      </p>
    </sec>
    <sec id="sec-10">
      <title>Evaluation and Be nchmarks</title>
      <p>Both Speakers may have an equal number of a grading of
“1” in all extracted “hot spots” detected or one of the
Speakers may have a slightly higher/lower or a considerably
higher/lower grading of “1”. A grading of “1” in 50% or
more of the “hot spots” signalizes that the Illocutionary Act
performed by the Speaker concerned is not restricted to
“Obtaining Information Asked” or “Providing Informatio n
Asked”. Speaker behavior indicating that Illocutionary Acts
performed are not restricted to the predefined interaction
framework is evaluated by the following benchmarks, where
Z = the number of “hot spot” signalized speech segments
divided by (÷): 2 (50%):
• Sum of Speaker grades ≥ Z.
• Example: Evaluation of Speaker Behavior (Speaker 1 is
less collaborative than Speaker 2).
• SPEAKER1: (1), (1), (1), (0), (1).
• SPEAKER2: (0), (0), (1), (1), (0).
• File length: 35 mins: SPEECH-SEGM ENT-count “hot
spots”: 5 (sum of grades =6, 6 ≥ Z where Z = 2.5).
These benchmarks for dialogs with long speech segments
can be referred to as “Collaboration” benchmarks with a
value of “Z” or “Collaboration (Z)”.</p>
    </sec>
    <sec id="sec-11">
      <title>Conclusions and Further Research</title>
      <p>By-passing and registering Cognitive Bias in HCI systems
assisting in the evaluation of Human-Human interaction
involves both automatic and interactive procedures.
Interactive topic tracking in the dialog structure and automatic “hot
spot” generation involving points of tension and/or c onflict
contribute to an evaluation of speakers -participants behavior
and intentions during the interaction.</p>
      <p>The behavior and Cognitive Bias of (i) speakers
-participants is evaluated in relation to the values of the “Relevance
(X)”, “Tension (Y)” and “Collaboration (Z)” benchmarks.
However, the same benchmarks may be used for evaluating
the Cognitive Bias - Confidence Bias of (ii) the
user-evaluator of the recorded and transcribed discussion or interview.</p>
      <p>Spoken dialogs concerning complex interactions between
speakers -participants are not limited to spoken journalistic
texts. As the variety and complexity of spoken HCI
applications increases, Speech Acts performed by one or multiple
users -participants, even by the System itself, often may
involve Illocutionary Acts beyond the predefined framewo rk
of a task-oriented dialog, especially in systems with emotion
recognition, virtual negotiation, psychological support or
decision-making.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Alexandris</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>M easuring Cognitive Bias in Spoken Interaction and Conversation: Generating Visual Representations</article-title>
          . In:
          <string-name>
            <surname>Beyond M achine Intelligence</surname>
          </string-name>
          <article-title>: Understanding Cognitive Bias and Humanity for Well-Being AI</article-title>
          .
          <source>In Proceedings from the AAAI Spring Symposium</source>
          , Stanford University,
          <fpage>204</fpage>
          -206
          <source>Technical Report</source>
          ,
          <fpage>SS18</fpage>
          -03, Palo Alto, CA: AAAI Press.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Alexandris</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Nottas</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          ; and Cambourakis,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Interactive Evaluation of Pragmatic Features in Spoken Journalistic Texts</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>In</given-names>
            <surname>Kurosu</surname>
          </string-name>
          , M . ed.,
          <string-name>
            <surname>Human-Computer</surname>
            <given-names>Interaction</given-names>
          </string-name>
          , Users and Contexts,
          <source>LNCS Lecture Notes in Computer Science</source>
          , Vol.
          <volume>9171</volume>
          :
          <fpage>259</fpage>
          -
          <lpage>268</lpage>
          , Heidelberg, Germany: Springer.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Alexandris</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>English, German and the International “Semiprofessional” Translator: A M orphological Approach to Implied Connotative Features</article-title>
          .
          <source>Journal of Language and Translation</source>
          ,
          <year>September 2010</year>
          , Vol.
          <volume>11</volume>
          ,
          <issue>2</issue>
          , Sejong University, Korea:
          <fpage>7</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Austin J. L.</surname>
          </string-name>
          <year>1962</year>
          .
          <article-title>How to Do Things with Words</article-title>
          .
          <source>Urmson. J.O.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>and Sbisà</surname>
          </string-name>
          , M . eds.,
          <source>2nd edition.</source>
          ,
          <year>1976</year>
          , Oxford, UK: Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Carlson</surname>
            ,
            <given-names>L.;</given-names>
          </string-name>
          <article-title>M arcu, D. ;</article-title>
          and Okurowski,
          <string-name>
            <surname>M . E.</surname>
          </string-name>
          <year>2001</year>
          .
          <article-title>Building a Discourse-Tagged Corpus in the Framework of Rhetorical Structure Theory</article-title>
          .
          <source>In Proceedings of the 2nd SIGDIAL Workshop on Discourse and Dialogue, Eurospeech</source>
          <year>2001</year>
          , Denmark,
          <year>September 2001</year>
          ,
          <string-name>
            <given-names>ACL</given-names>
            <surname>Anthology</surname>
          </string-name>
          ,
          <fpage>W01</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Johnston</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          ; M cGee,
          <string-name>
            <given-names>D.</given-names>
            ;
            <surname>Oviatt</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; Pittman,
          <string-name>
            <given-names>J.</given-names>
            ;
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ;
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ; and
            <surname>Clow</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>1997</year>
          .
          <article-title>Quickset: M ultimodal Interaction for Distributed Applications</article-title>
          .
          <source>In Proceedings of the 5th ACM International Multimedia Conference</source>
          ,
          <volume>31</volume>
          -
          <fpage>40</fpage>
          , New York, NY: ACM Digital Library .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Cutts</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          <year>2013</year>
          . Oxford Guide to Plain
          <source>English. 4th edition.</source>
          , Oxford, UK: Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Du</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Alexandris,
          <string-name>
            <given-names>C.</given-names>
            ; M ourouzidis, D. ;
            <surname>Floros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            ; and
            <surname>Iliakis</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Controlling Interaction in M ultilingual Conversation Revisited: A Persp ective for Services and Interviews in M andarin Chinese</article-title>
          . In Kurosu, M . ed.,
          <source>Lecture Notes in Computer Science LNCS</source>
          <volume>10271</volume>
          :
          <fpage>573</fpage>
          -
          <lpage>583</lpage>
          , Heidelberg, Germany: Springer.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Floros</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>M ourouzidis</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>M ultiple Task M anagement in a Dialog System for Call Centers</article-title>
          .
          <source>M aster's Thesis</source>
          , Department of Informatics and Telecommunications, National University of Athens, Greece.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Grice</surname>
            ,
            <given-names>H.P.</given-names>
          </string-name>
          <year>1975</year>
          .
          <article-title>Logic and conversation</article-title>
          . In: Cole,
          <string-name>
            <surname>P.</surname>
          </string-name>
          , M organ, J.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          eds.,
          <source>Syntax and Semantics</source>
          , Vol.
          <volume>3</volume>
          . Academic Press, New York.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Hatim</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>1997</year>
          .
          <article-title>Communication Across Cultures: Translation Theory and Contrastive Text Linguistics</article-title>
          . Exeter, UK: University of Exeter Press.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Hilbert</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Toward a Synthesis of Cognitive Biases: How Noisy Information Processing Can Bias Human Decision M aking</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Psychological</given-names>
            <surname>Bulletin</surname>
          </string-name>
          , Vol
          <volume>138</volume>
          (
          <issue>2</issue>
          )
          <string-name>
            <surname>,</surname>
            <given-names>M ar</given-names>
          </string-name>
          <year>2012</year>
          :
          <fpage>211</fpage>
          -
          <lpage>237</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          <year>2009</year>
          .
          <article-title>Introduction to Practical Speech User Interface Design for Interactive Voice Response Applications</article-title>
          , IBM Software Group, USA, Tutorial T09 presented at HCI 2009 San Diego, CA, USA M a,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>A comparative analysis of the ambiguity resolution of two English-Chinese M T approaches</article-title>
          : RBM T and SM T. Dalian University of Technology Journal,
          <volume>31</volume>
          (
          <issue>3</issue>
          ):
          <fpage>114</fpage>
          -
          <lpage>119</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>M arcu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>1999</year>
          .
          <article-title>Discourse trees are good indicators of importance in text</article-title>
          . In M ani, I. and M aybury , M . (eds),
          <source>Advances in Automatic Text Summarization</source>
          , Cambridge M A,
          <string-name>
            <surname>The</surname>
            <given-names>M</given-names>
          </string-name>
          IT Press:
          <fpage>123</fpage>
          -
          <lpage>136</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Nass</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Brave</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Wired for Speech: How Voice Activates and Advances the Human-Computer Relationship</article-title>
          . Cambridge M A:
          <string-name>
            <surname>The M IT</surname>
          </string-name>
          <article-title>Press</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          2007.
          <article-title>A Hybrid Approach to Dialog Input in the CitzenShield Dialog System for Consumer Complaints</article-title>
          .
          <source>In Proceedings of HCII</source>
          <year>2007</year>
          , Beijing, Peoples Republic of China.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Paltridge</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <year>2012</year>
          .
          <article-title>Discourse Analysis: An Introduction</article-title>
          . London, UK: Bloomsbury Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2000</year>
          .
          <article-title>Politeness in Chinese Face-to-Face Interaction</article-title>
          .
          <source>Advances in Discourse Processes Series</source>
          Vol.
          <volume>67</volume>
          ,
          <string-name>
            <surname>Stamford</surname>
            ,
            <given-names>CT</given-names>
          </string-name>
          : Ablex Publishing Corporation.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Sacks</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Schegloff</surname>
            ,
            <given-names>E. A.</given-names>
          </string-name>
          ; and Jefferson,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>1974</year>
          .
          <article-title>A simplest systematics for the organization of turn-taking for conversation.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Language</surname>
          </string-name>
          , Vol.
          <volume>50</volume>
          :
          <fpage>696</fpage>
          -
          <lpage>735</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Searle J. R.</surname>
          </string-name>
          <year>1969</year>
          .
          <article-title>Speech Acts: An Essay in the Philosophy of Language</article-title>
          . Cambridge,
          <string-name>
            <surname>M A</surname>
          </string-name>
          : Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Spontaneous and non-spontaneous turn-taking.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Pragmatics</surname>
          </string-name>
          , Vol.
          <volume>16</volume>
          (
          <issue>2-3</issue>
          ):
          <fpage>329</fpage>
          -
          <lpage>360</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Stede</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Taboada</surname>
            ,
            <given-names>M .</given-names>
          </string-name>
          , ; and
          <string-name>
            <surname>Das</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Annotation Guidelines for Rhetorical Structure</article-title>
          . M anuscript. University of Potsdam and Simon Fraser University. M arch
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Trofimova</surname>
            <given-names>I.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>Observer Bias: An Interaction of Temperament Traits with Biases in the Semantic Perception of Lexical M aterial</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>PLoS ONE 9</source>
          (
          <issue>1</issue>
          ):
          <fpage>e85677</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Tung</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Gomez</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Kawahara,
          <string-name>
            <surname>T.</surname>
          </string-name>
          ; and M atsuyama,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>M ulti-party Human-M achine Interaction Using a Smart M ultimodal Digital Signage</article-title>
          . In Kurosu, M . ed.,
          <string-name>
            <surname>Human-Computer Interaction</surname>
          </string-name>
          .
          <source>Interaction Modalities and Techniques, Lecture Notes in Computer Science</source>
          , Vol.
          <volume>8007</volume>
          ,
          <year>2013</year>
          ,
          <fpage>408</fpage>
          -
          <lpage>415</lpage>
          , Heidelberg, Germany: Springer.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Gailliot</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hyden</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and Lietzenmayer,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>A Knowledge</given-names>
            <surname>Elicitation</surname>
          </string-name>
          <article-title>Study for Collaborative Dialogue Strategies Used to Handle Uncertainties in Speech Communication While Using GIS</article-title>
          . In Kurosu, M . ed.,
          <source>Human-Computer Interaction, Lecture Notes in Computer Science</source>
          <volume>8007</volume>
          ,
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          , Heidelberg, Germany: Springer.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Wardhaugh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>1992</year>
          .
          <article-title>An Introduction to Sociolinguistics. 2nd edition</article-title>
          . Oxford, UK: Blackwell.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Asadi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; and Zweig,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Hybrid Code Networks: practical and efficient end-to-end dialog control with supervised and reinforcement learning</article-title>
          .
          <source>In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics</source>
          , Vancouver, Canada,
          <source>July 30 - August 4</source>
          ,
          <year>2017</year>
          ,
          <fpage>665</fpage>
          -
          <lpage>677</lpage>
          , Association for Computational Linguistics -ACL.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>Wilson</surname>
            ,
            <given-names>K. E.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>An oscillator model of the timing of turntaking</article-title>
          .
          <source>Psychonomic Bulletin and Review</source>
          <year>2005</year>
          :
          <volume>12</volume>
          (
          <issue>6</issue>
          ):
          <fpage>957</fpage>
          -
          <lpage>968</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Levow</surname>
            <given-names>G.A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>M eng F. H.</surname>
          </string-name>
          <year>2012</year>
          .
          <article-title>Predicting User Satisfaction in Spoken Dialog System Evaluation With Collaborative Filtering</article-title>
          .
          <source>IEEE Journal of Selected Topics in Signal Processing</source>
          , Vol.
          <volume>6</volume>
          ,
          <issue>Issue</issue>
          : 8,
          <string-name>
            <surname>Dec</surname>
          </string-name>
          .
          <year>2012</year>
          :
          <fpage>971</fpage>
          -
          <lpage>981</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Aoyama</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ; Ozeki,
          <string-name>
            <surname>M .</surname>
          </string-name>
          ; and Nakamura,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <surname>Capture</surname>
          </string-name>
          , Recognition, and
          <article-title>Visualization of Human Semantic Interactions in M eetings</article-title>
          .
          <source>In Proceedings of PerCom</source>
          , Mannheim, Germany,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <surname>Zeldes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>"rstWeb - A Browser-based Annotation Interface for Rhetorical Structure Theory and Discourse Relations"</article-title>
          .
          <source>In Proceedings of NAACL-HLT 2016 System Demonstrations</source>
          , San Diego, CA,
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          ,
          <string-name>
            <given-names>ACL</given-names>
            <surname>Anthology</surname>
          </string-name>
          , NAACL-HLT
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>