<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Concerns: User-Specific Pitfalls of Favoring Voice over Text in Conversational Recom mender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alain D. Starke</string-name>
          <email>alain.starke@wur.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minha Lee</string-name>
          <email>m.lee@tue.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Conversational User Interfaces</institution>
          ,
          <addr-line>Recommender Systems, Accessibility, Inclusion, Voice-based Systems</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Eindhoven University of Technology</institution>
          ,
          <addr-line>Groene Loper 3, 5612 AE Eindhoven</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Environments (ComplexRec) Joint Workshop @ RecSys 2021</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Bergen</institution>
          ,
          <addr-line>P.O. Box 7800, 5020 Bergen</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Wageningen University &amp; Research</institution>
          ,
          <addr-line>Droevendaalsesteeg 4, 6708 PB Wageningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <abstract>
        <p>In the context of Conversational Recommender Systems (CRSs) and Conversational User Interfaces (CUIs; e.g., Digital Assistants, such as Siri), an increasing number of voice-based applications are emerging, often at the expense of text-based applications. In this position paper, we argue that the possible first-mover advantage of adopting voice-based technologies may put specific groups of users at a profound disadvantage, as they are likely to run into accessibility issues. For example, users that stammer or whom are not fluent in the English language have a hard time using voice-based conversational recommender systems. Along this line, we describe a number of challenges and issues for current and future systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Voice-based technologies are difusing through society at
a fast pace. Reportedly 4.2 billion digital voice assistants
were in use in 2020 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], including well-known
technologies such as Amazon Alexa and Siri. Their role in the
‘Internet of Things’ system is becoming increasingly
important [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], in the sense incumbent technologies, such as
recommender systems, are often made compatible with
voice-based applications [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
ever, most conversational recommender systems (CRSs)
to date are text-based [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. They focus on mining textual
user input, such as through fixed messages in clickable
menus or by open-ended text queries [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. In
comparison, the number of voice-based conversational
recommender systems are still limited, but is likely to expand
in the coming years [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        The user-system dynamic between text-based and
voice-based interactions difers greatly. Whereas
textbased CRSs can rely on either open-ended queries or fixed
input (e.g., the user selects an answer option), voice-based
queries tend to be impromptu and are more complex to
process. Nonetheless, given the current share and
expected growth of digital assistant use [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the emergence
of digital assistants, such as Amazon Alexa and Google
      </p>
      <sec id="sec-1-1">
        <title>Home, suggests that designing for voice-based interac</title>
        <p>Systems (KaRS) &amp; 5th Edition of Recommendation in Complex</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Conversational Systems</title>
      <sec id="sec-2-1">
        <title>2.1. Text-based systems</title>
        <sec id="sec-2-1-1">
          <title>Text-based conversational systems, which are also known</title>
          <p>
            as chatbots, have been around for decades, such as
Weizenbaum’s ELIZA in the 1960s [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ]. Chatbots now
exist on many business-to-consumer websites, for
example as an automated customer service agent [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. In
terms of technical implementation, two approaches are
taken to build chatbots, which typically also applies to
text-based conversational recommenders [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. They are understand multimodal cues like gestures or gaze that
aceither built as command-based systems that respond to company users’ speech [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ], which is likely to afect the
user queries as if they are commands or as bots that use interpretability of the voice-based query. Finally, when
natural language understanding. For example, on Slack, there is more than one person talking, the system has to
one can issue commands (e.g., unsubscribe) that chatbots distinguish whose voice to zone in on [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ], which might
can easily understand rather than using natural language lead to conflicts of agency [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]. Each of these problems
(e.g., “please get me out of this channel”) that can be are challenging. Yet, even if these technical issues are
vague for systems to understand [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ]. Also, people may resolved, they will not create a headway for a seamless
easily misspell or give incomplete input that chatbots experience for individuals; inclusion is not about
onecannot accurately interpret. size-fits-all, but about how these technical issues do not
          </p>
          <p>
            An important challenge of conversational systems is disproportionally afect specific userse.
to mitigate misunderstandings or a conversational
breakdown [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ]. There are inadequate responses to user
requests, false positives [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ], either because of unclear 3. Critique of Voice-based systems
query or a missing database category [
            <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
            ]. In such
cases, conversational repair strategies becomes important
like how the system should correct for misunderstood
phrases or unclear user intentions [
            <xref ref-type="bibr" rid="ref12">12</xref>
            ]. An example of a
repair strategy is a system giving potential options that
people can choose from, such as “I did not understand
that. Did you mean X or Y?”, to keep the conversation
going.
          </p>
          <p>In sum, the two strategies then are to design 1)
command-based systems that allow for minimum
flexibility on user input for eficiency or 2) natural
languagebased systems that allow for greater input flexibility, but
also then, with an increased number of possible repair
strategies that are not always successful.</p>
          <p>
            Voice-based queries tend to be ‘messier’ than text-based
inputs [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ]. The current state-of-the-art in Natural
Language Processing methods opens up possibilities for more
open-ended conversational strategies in
recommendation. However, even NLP-based systems are often still
limited to familiar input, i.e., requiring an explicit
understanding of users’ messages, which often gets
misunderstood. Problems like of environmental noise distortion
of user input are common [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. Yet when it comes to
user adoption, voice-based technologies have difused
at a large speed in terms of innovation adoption [
            <xref ref-type="bibr" rid="ref17 ref2">17, 2</xref>
            ].
          </p>
          <p>These have both positive and negative side efects.</p>
          <p>
            On the one hand, it seems that voice-based
applications be integrated in a modular way, as they can work
with recommendation libraries without designing an
appropriate user interface. On the other hands, it seems
that innovation in technology may only benefit those
that can work it. Even for people who are considered to
be “regular users”, there is a lot of trial and error when
it comes to learning how to interact with voice-based
agents [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] and recommender systems. But there are
people who have additional dificulties due to various
diferences in abilities; technical solutions often are built
around the normative assumption that users are fully
able-bodied, i.e., with sight, hearing, and other abilities
intact [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ].
          </p>
          <p>
            In recommender domains that use more traditional
interfaces, marginalized people are often ‘served’ by a
simple fix. For example, a tourism recommender system
for people with physical disabilities would apply
postifltering to an appropriate set of recommendations [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ].
          </p>
          <p>
            The problem for conversational recommenders, however,
not only applies to the appropriateness of the suggested
context, but also to the usability of the technology in
the first place. For example, people who stammer, being
a small subset of the population, face dificulties at the
start of their interaction: a voice-based system often
cannot understand what they say due to the lack of training
data and design choices [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ]. Their speech does not fit
the normative template of how people should “normally”
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Voice-based systems</title>
        <p>
          Voice-based conversational user interfaces (VUIs) are
becoming more popular in everyday use. In particular,
smart home assistants such as Google Home and
Amazon Alexa are used more frequently to help its users find
content that they are looking whether [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ], regardless of
whether the query is mundane and factual (e.g., ‘What
weather is it today?’), or more exploratory (e.g., ‘Play me
a song for my dinner party’). The latter is more
commonly explored in recommender systems research, for it
seeks to retrieve an item that a user does not explicitly
know about.
        </p>
        <p>
          Current voice-based systems face a number of
technical issues that are often situational. For example,
considering the use of voice-based assistant in a car [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], there
may be environmental noise, people not formulating
clearly due to multi-tasking driving and voice
interaction, among other causes. Technical issues that come
with VUIs are many, and will also impact the design of
voice-based conversational recommender systems. To
list three, they at times lack noise robustness, multimodal
understanding, and addressee detection [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. In most
contexts, users’ environmental conditions will feature a
certain degree of background noise, such as when people are
away from home. Moreover, voice-based system cannot
talk. Furthermore, many smart home assistants are only text-based chatbots with voice-based agents. In terms
compatible with a few languages (e.g., they are ‘biased’ of users being “better understood”, the decades old
texttowards English [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]), and speech recognition may be dis- based interactions may be better suited. Perhaps
countertorted because of fluctuations in human emotions [
          <xref ref-type="bibr" rid="ref22 ref4">4, 22</xref>
          ]. intuitively, due to the limits of query and text-based
conWe realize that these issues on inclusion in the use of versational recommendation, the odds are smaller that
voice-based assistants is nuanced [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. This means that it ‘gets it wrong’. Or, more technically, that it generates
notions on who can easily use voice-based agents de- a negative adversarial response [25], or has a
conversapends on multiple factors, such as accents, speech pat- tional breakdown [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Although the usability is arguably
terns, and the access to commercial agents like Alexa, lower in the sense that one needs to “touch” an
interwhich introduces many ways that eforts to include can face, this poses a huge advantage to those who have to
also be exclusive, e.g., by prioritizing one accent over concentrate to interact with such a voice-based
applicaothers. tion. However, the consumer trend is shaping up to favor
voice-based applications; IBM, Google, and Amazon have
product lines that promote voice-first interactions.
4. Suggestions for Conversational We ofer two suggestions moving forward. To optimize
Recommender Systems accessibility for all users, a move towards ‘voice-enabled’
rather than ‘voice-first’ or voice-based recommender
sysWe have highlighted diferent challenges for both text- tems would be desirable, akin to technology in which
based and voice-based interactions. What stands out is ‘voice’ is a feature rather a key characteristic (e.g., Siri on
that some challenges are easier to resolve with user train- an Apple iPhone). Although this requires the deployment
ing or adaptation (e.g., lacking suficient technical knowl- of two diferent retrieval and recommendation pipelines,
edge to use a text-based interface), than other challenges it maximizes accessibility by combining ‘the best of two
(e.g., non-native users lack vocal skils, such as because of worlds’. To note, we did not consider multimodality, e.g.,
stammering) [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. What these challenges have in com- combination of voice, gaze, body movements, and more,
mon is how people’s assumptions about conversational which will become more important in the coming years
agents, be they chatbots or Alexas, shape their interac- [26].
tions. People may have expectations that conversational We also suggest that diversity of data for retrieval and
agents cannot meet, as the systems cannot yet to com- recommendation is essential to design inclusive
converplex tasks such as email management by voice [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. Even sational recommender systems, or systems that cater to
text-based chatbots often do not meet people’s needs, as specific users. Eforts are ongoing when it comes to
makusers expect a higher level of understanding from bots ing voice-based interactions more accessible; Google’s
that they were not designed for [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Hence, for most of Project Euphonia1 aims to collect more data on atypical
us, going beyond simple interactions and towards more speech, e.g., from people with cerebral palsy. Similarly,
complex exchanges is a problem that we all share due to more time should be spend on collecting dificult data
the state of the technology. when it comes to voice in research, in terms of responding
        </p>
        <p>
          Some studies describe that conversational recom- to “unconventional voices”.
mender systems are distinct from the more traditional
chatbots and dialogue-based systems [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. However, we
argue that the retrieval of conversational elements in 5. Conclusion
conjunction with ‘task-related items’ are two sides of the
same coin. A task-based conversation can be dialogue- This paper has reflected on current practices in
converbased, by supporting a task at hand. Instead of focusing sational recommender systems. In particular, we have
on a false dichotomy between task-based or dialogue- pitted text-based systems against voice-based systems,
based systems, a better way forward is being attentive to observing that while voice-based recommender systems
how diferent users’ capacities get highlighted or ignored are becoming more common because of their integration
by systems. The problem to focus on is inclusion vs. ex- with digital assistant [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], it may put specific users at a
clusion of user groups based on systems’ assumptions of disadvantage. We have identified a number of challenges
diferent abilities that people may or may not have. to make CRSs more inclusive, particularly for the
emerg
        </p>
        <p>
          How should we move forward with conversational ing domain of voice-based user interfaces. We emphasize
recommender systems? Recommender systems are tradi- lastly that inclusion for some may mean exclusion for
tionally applied in domains where one-shot recommen- others. In order to recommend to all users, we need to
dations are efective [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], such as movies, e-commerce, understand all users. Specifically, understanding users
and books. The use of conversations, however, makes not only in terms of preferences, but also in terms of the
for more complex interactions which introduces greater
technical challenges. We above diferentiated between 1https://sites.research.google/euphonia/about/
fundamental conversational elements, such as speech,
should be a priority.
generation of recommender systems: A survey of
the state-of-the-art and possible extensions, IEEE
transactions on knowledge and data engineering
17 (2005) 734–749.
[25] G. Penha, C. Hauf, What does bert know about
books, movies and music? probing bert for
conversational recommendation, in: Fourteenth ACM
Conference on Recommender Systems, 2020, pp.
        </p>
        <p>388–397.
[26] Y. Deldjoo, J. R. Trippas, H. Zamani, Towards
multimodal conversational information seeking, in:
Proceedings of the ACM Conference on Research and
Development in Information Retrieval, SIGIR,
volume 21, 2021.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Vailshery</surname>
          </string-name>
          ,
          <article-title>Number of digital voice assistants in use worldwide from 2019 to 2024 (in billions</article-title>
          ),
          <year>2021</year>
          . URL: https://www.statista.com/statistics/973815/ worldwide-digital
          <article-title>-voice-assistant-in-use/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Arpnikanondt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Funilkul</surname>
          </string-name>
          , W. Chutimaskul,
          <article-title>The adoption analysis of voice-based smart iot products</article-title>
          ,
          <source>IEEE Internet of Things Journal</source>
          <volume>7</volume>
          (
          <year>2020</year>
          )
          <fpage>10852</fpage>
          -
          <lpage>10867</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Iovine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Narducci</surname>
          </string-name>
          , G. Semeraro,
          <article-title>Conversational recommender systems and natural language:: A study through the converse framework</article-title>
          ,
          <source>Decision Support Systems</source>
          <volume>131</volume>
          (
          <year>2020</year>
          )
          <fpage>113250</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. de Rijke</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Advances and challenges in conversational recommender systems: A survey</article-title>
          ,
          <source>arXiv preprint arXiv:2101.09459</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Hernandez-Bocanegra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <article-title>Conversational review-based explanations for recommender systems: Exploring users' query behavior</article-title>
          ,
          <source>in: CUI 2021-3rd Conference on Conversational User Interfaces</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>A survey on conversational recommender systems</article-title>
          ,
          <source>ACM Computing Surveys (CSUR) 54</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Angkititrakul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. E.</given-names>
            <surname>Shriberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Heck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Peters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <article-title>Conversational in-vehicle dialog systems: The past, present, and future</article-title>
          ,
          <source>IEEE Signal Processing Magazine</source>
          <volume>33</volume>
          (
          <year>2016</year>
          )
          <fpage>49</fpage>
          -
          <lpage>60</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Weizenbaum</surname>
          </string-name>
          ,
          <article-title>Eliza-a computer program for the study of natural language communication between man and machine</article-title>
          ,
          <source>Communications of the ACM</source>
          <volume>9</volume>
          (
          <year>1966</year>
          )
          <fpage>36</fpage>
          -
          <lpage>45</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dale</surname>
          </string-name>
          ,
          <article-title>The return of the chatbots</article-title>
          ,
          <source>Natural Language Engineering</source>
          <volume>22</volume>
          (
          <year>2016</year>
          )
          <fpage>811</fpage>
          -
          <lpage>817</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Frank</surname>
          </string-name>
          , W. IJsselsteijn,
          <article-title>Brokerbot: A cryptocurrency chatbot in the social-technical gap of trust, Computer Supported Cooperative Work (CSCW) 30 (</article-title>
          <year>2021</year>
          )
          <fpage>79</fpage>
          -
          <lpage>117</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Følstad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Taylor</surname>
          </string-name>
          ,
          <article-title>Conversational repair in chatbots for customer service: the efect of expressing uncertainty and suggesting alternatives</article-title>
          , in: International Workshop on Chatbot Research and Design, Springer,
          <year>2019</year>
          , pp.
          <fpage>201</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ashktorab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Weisz</surname>
          </string-name>
          ,
          <article-title>Resilient chatbots: Repair strategy preferences for conversational breakdowns</article-title>
          ,
          <source>in: Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>Corti</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Gillespie,</surname>
          </string-name>
          <article-title>Co-constructing intersubjectivity with artificial conversational agents: people are more likely to initiate repairs of misunderstandings with agents represented as human</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>58</volume>
          (
          <year>2016</year>
          )
          <fpage>431</fpage>
          -
          <lpage>442</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Cowan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Doyle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Edwards</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garaialde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hayes-Brady</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. P.</given-names>
            <surname>Branigan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cabral</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <article-title>What's in an accent? the impact of accented synthetic speech on lexical choice in human-machine dialogue</article-title>
          ,
          <source>in: Proceedings of the 1st International Conference on Conversational User Interfaces</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Heylen</surname>
          </string-name>
          ,
          <article-title>Head gestures, gaze and the principles of conversational structure</article-title>
          ,
          <source>International Journal of Humanoid Robotics</source>
          <volume>3</volume>
          (
          <year>2006</year>
          )
          <fpage>241</fpage>
          -
          <lpage>267</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Noortman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zaga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Starke</surname>
          </string-name>
          , G. Huisman,
          <string-name>
            <given-names>K.</given-names>
            <surname>Andersen</surname>
          </string-name>
          ,
          <article-title>Conversational futures: Emancipating conversational interactions for futures worth wanting</article-title>
          ,
          <source>in: Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname>McLean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Osei-Frimpong</surname>
          </string-name>
          ,
          <article-title>Hey alexa… examine the variables influencing the use of artificial intelligent in-home voice assistants</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>99</volume>
          (
          <year>2019</year>
          )
          <fpage>28</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>C. M. Myers</surname>
            ,
            <given-names>L. F.</given-names>
          </string-name>
          <string-name>
            <surname>Laris Pardo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Acosta-Ruiz</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Canossa</surname>
          </string-name>
          , J. Zhu, “try, try, try again:
          <article-title>” sequence analysis of user interaction data with a voice user interface</article-title>
          ,
          <source>in: CUI 2021-3rd Conference on Conversational User Interfaces</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Costanza-Chock</surname>
          </string-name>
          ,
          <article-title>Design justice: Towards an intersectional feminist framework for design theory and practice</article-title>
          ,
          <source>Proceedings of the Design Research Society</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Mahmoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>El-Bendary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Mokhtar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. E.</given-names>
            <surname>Hassanien</surname>
          </string-name>
          ,
          <article-title>Similarity measures based recommender system for rehabilitation of people with disabilities</article-title>
          ,
          <source>in: The 1st International Conference on Advanced Intelligent System and Informatics (AISI2015)</source>
          ,
          <source>November 28-30</source>
          ,
          <year>2015</year>
          ,
          <string-name>
            <given-names>Beni</given-names>
            <surname>Suef</surname>
          </string-name>
          , Egypt, Springer,
          <year>2016</year>
          , pp.
          <fpage>523</fpage>
          -
          <lpage>533</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>L.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Cowan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lindsay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Sheers</surname>
          </string-name>
          ,
          <article-title>Speech diversity and speech interfaces: Considering an inclusive future through stammering</article-title>
          ,
          <source>in: Proceedings of the 2nd Conference on Conversational User Interfaces</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pittermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pittermann</surname>
          </string-name>
          , W. Minker,
          <article-title>Emotion recognition and adaptation in spoken dialogue systems</article-title>
          ,
          <source>International Journal of Speech Technology</source>
          <volume>13</volume>
          (
          <year>2010</year>
          )
          <fpage>49</fpage>
          -
          <lpage>60</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>E.</given-names>
            <surname>Luger</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Sellen, ”
          <article-title>like having a really bad pa” the gulf between user expectation and experience of conversational agents</article-title>
          ,
          <source>in: Proceedings of the 2016 CHI conference on human factors in computing systems</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>5286</fpage>
          -
          <lpage>5297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>G.</given-names>
            <surname>Adomavicius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tuzhilin</surname>
          </string-name>
          , Toward the next
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>