<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>S. Li, W. Lei, Q. Wu, X. He, P. Jiang, T.-S. Chua, Seamlessly unifying attributes and items:
Conversational recommendation for cold-start users, ACM Trans. Inf. Syst.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1145/3397481.3450657</article-id>
      <title-group>
        <article-title>Mixed-Modality Interaction in Conversational Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuan Ma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Timm Kleemann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jürgen Ziegler</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Duisburg-Essen</institution>
          ,
          <addr-line>Duisburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>39</volume>
      <issue>2021</issue>
      <fpage>480</fpage>
      <lpage>490</lpage>
      <abstract>
        <p>Recent advances in natural language processing have made modern chatbots and Conversational Recommender Systems (CRS) increasingly intelligent, enabling them to handle more complex user inputs. Still, the interaction with a CRS is often tedious and error-prone. Especially when using written text as the form of conversation, the interaction is often less eficient in comparison to conventional GUIstyle interaction. To keep the flexibility and mixed-initiative style of language-based conversation while leveraging the eficiency and simplicity of interacting through graphical widgets, we investigate the design space of integrating GUI elements into text-based conversations. While simple response buttons have already been used in chatbots, the full range of such mixed-modality interactions has not yet been investigated in existing research. We propose two design dimensions along which integrations can be defined and analyze their applicability for preference elicitation and for critiquing the CRS's responses at diferent levels. We report a user study in which we investigated user preferences and perceived usability of diferent techniques based on video prototypes.</p>
      </abstract>
      <kwd-group>
        <kwd>conversational recommender systems</kwd>
        <kwd>user interface</kwd>
        <kwd>preference elicitation</kwd>
        <kwd>critique-based recommendations</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In recent years, conversational styles of interaction have increasingly been applied in the field
of recommender systems. Conversational Recommender Systems (CRS) aim to provide a more
human-like and more comprehensible form of eliciting users’ preferences and recommending
suitable items [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. While many e-commerce sites present recommendations in a static fashion
allowing little or no user interaction, the need for more flexible and personalized ways of
providing recommendations is increasingly recognized. For this purpose, CRS are utilized,
whereby a virtual agent interacts with the user [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A variety of techniques has been explored
for providing conversational interaction with a recommender, some systems, for example, follow
a strict rule-based process requiring the user to answer questions with predefined answers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
while other approaches are designed to mimic a natural conversation in which users can freely
formulate their questions and answers [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Rather than providing one-shot recommendations with only limited user intervention, CRS
enable users to respond to the recommendations they receive, to criticize them, or to provide
more precise indications of their own preferences in interactive conversations with the virtual
agent [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The mode of interaction in current CRS is predominantly text-based where textual
input by the user is followed by textual responses. In some cases, however, CRS also ofer the
user a set of potential answers in a GUI style, presenting the options as buttons. Augmenting
the text-only interaction by GUI elements serves two purposes: (1) it provides users with a more
eficient input technique and (2) it reduces the number of misrecognitions which may occur
when just relying on typed (or spoken) input. While combining textual and GUI interaction in a
CRS can increase its efectiveness and usability, the design space of possible multimodal CRS
interactions has not yet been suficiently explored yet and there is a number of options that
have not been investigated or even considered in CRS research.
      </p>
      <p>In this paper, we aim at providing a first, more complete investigation of the diferent ways
how textual interaction can be combined with GUI-like interactions. Keeping the flexibility of
free text interaction, we introduce and investigate a variety of additional interactions by which
the textual interaction may be augmented. The options investigated include directly changing
(critiquing) features of an item shown as recommendation in the dialog as well as interaction
with the textual responses given by the system.</p>
      <p>We developed video prototypes for these interactions and investigated them in an online
user study. The results provide initial insights into interaction methods that users may prefer.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        CRS exhibit a number of potential advantages over conventional GUI-based recommenders.
They can provide a more natural and less obtrusive way of obtaining information about the user’s
preferences which is essential for generating personalized recommendations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In contrast
to upfront elicitation steps that are detached from the actual recommending, such as rating a
number of sample items e.g. MovieLens,1 or completing initial interviews [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] or personality
questionnaires [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ], the elicitation of needs and preferences can be smoothly integrated into the
dialog flow, thus also mitigating the cold-start problem. Provided a suficient level of language
understanding on the part of the system, the expression of user intentions [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], preferences or
even dislikes is more flexible in comparison to system-initiated GUI interactions and
closedform questions. This flexibility may come, however, at the cost of lower eficiency, especially
when users need to type their questions and responses instead of just clicking one of several
pre-defined options. Therefore, CRS should aim to achieve an acceptable flexibility-eficiency
trade-of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        A further, and so far less considered aspect of CRS is their capability to provide means for
critiquing the system’s recommendations, thus increasing user control over the
recommendations. According to Chen and Pu [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] three critiquing approaches can be distinguished: natural
language dialog-based critiquing (NLC), system-suggested critiquing (SC) and user-initiated
critiquing (UC). NLC as a specific form of conversation, either text-based or voice-based, is
well compatible with the general CRS approach. NLC can be performed in a human-like style,
simulating, for instance, the conversation with a salesperson (e.g. ExpertClerk [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]). SC on
the other hand provides system-initiated critiquing options and asks users for one or more
responses, such as, for example, in multi-attribute utility theory (MAUT)-based Compound
Critiques [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The benefits of SC lie mainly in their ability of guiding users concerning relevant
and acceptable feedback criteria, so that the system can better understand users and enhance
its recommendation efectiveness. UC ofers users a more user-initiated form of critiquing, such
as in Example Critiquing [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The main advantage of UC is that it allows for a higher level of
user control. Also, hybrid critiquing techniques have been proposed [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and compared [15].
By means of dialogs in CRS, both UC and SC may be combined and used flexibly. SC and UC are
often realized using graphical user interface elements, such as buttons, sliders and checkboxes
to either respond to system questions (SC) or to change properties of a recommended item (UC).
      </p>
      <p>
        There exist a large and increasing variety of techniques for realizing CRS [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Recent
approaches for improving CRS performance include, among others, knowledge graph-based
methods [16], contextual bandits [17], bandit approaches unifying items and features [18], or
topic guided methods [19]. A recent survey [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] provides a good overview of the current status of
CRS. Most of the recent works are focused on the underlying methods and algorithms, providing
users with a text-only form of interaction. However, multi-turn conversations can be built on
any form of interaction or mixed-modality interactions instead of merely textual form [20].
      </p>
      <p>Only limited research has thus far focused on the question how conversational interactions
in recommender systems using diferent modalities impact CRS performance and, in particular,
users’ perception of a CRS. Several works have investigated diferent interaction methods,
although in other domains. For example, Ciechanowski et al. [21] evaluate diferent interactions
to avoid the uncanny valley efect in chatbots, their interactions including text, voice with
human-like avatar animations. In their study, text-based interaction was considered to be more
pleasant by the participants, compared to voice interaction with a human-like avatar.</p>
      <p>The combination of digital assistants and CRS has been investigated recently, here, the results
indicate that a combination of buttons and natural language is particularly beneficial [ 22]. Jin
et al. [23] conducted an experiment to explore the correlation between the user’s interaction and
personal characteristics. What is interesting for our study is that they deploy several interactive
methods in MusicBot: text, voice, button, radio button (ratings). From their results, one can see
that participants used buttons most frequently, then radio buttons, followed by text and voice.
This indicates that text-only interaction in a CRS might not be the most useful and preferred
technique, providing a motivation for our research presented here.</p>
      <p>Valério et al. [24] performed a comparison of diferent chatbot interaction paradigms. Chatbot
Kino used only text to communicate with users, while the alternative chatbot Cinemito used text
in combination with buttons and images for providing quick feedback. Their analysis revealed
that there is not a clearly preferred way of interaction. Their work was a qualitative study (n = 10)
and mainly focused on user’ perception, aiming provide design guidance for chatbots. The study
presented in this paper extends existing work by focusing on conversational recommenders,
by introducing mixed-modality interaction in CRS and by providing empirical evidence of the
benefits of combining interaction modalities for preference capture and critiquing.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Mixed-Modality Interaction in CRS</title>
      <p>Combining diferent modalities in human-system interaction can generally bring about various
benefits such as increased eficiency of the interaction or better disambiguation in probabilistic
input recognition. Text-based or speech-based dialogs provide the user with a natural and flexible
interaction style which, however, is also error-prone and often not very transparent since the
user needs to anticipate the comprehension capabilities of the system to avoid misinterpretation
or rejection of the input. When the input options are limited, selecting from the available
options is also mostly quicker, resulting in higher eficiency and often reduced user frustration.
While multimodal interfaces may employ a wide range of diferent modalities [ 25], we focus
in this paper on the visual channel and on the prevalent combination of input techniques in
CRS which is textual, language-based dialogue combined with graphical interaction widgets.
To distinguish this type of interaction from more general multimodal interfaces, we use the
term mixed modality for this combination. Even though restricted in the number of modalities,
the design space for interaction based on textual conversation with integrated GUI elements is
an under-researched area. To more systematically explore the design options we propose two
dimensions for characterizing the interactions.</p>
      <p>1
2
3
4</p>
      <p>The first dimension refers to the location where the GUI element is integrated in the
conversation flow. We call this the anchor dimension. An anchor can be located in or near the user input
area where typically buttons are used as shortcuts for otherwise textual user responses. Widgets
can also be attached to a presented recommendation itself, be that inline in the textual flow or
in a separate recommendation area. Responding to such prompts is essentially equivalent to
critiquing the recommendation since the widget will allow the user to change feature values,
and thus user preferences, directly on the displayed item. As a third anchor, we propose to
embed interactive elements directly in the textual output of the system. This way, the user can
respond directly to terms that appear in the output, such as features mentioned or the intended
usage of a product. Various options exist for making such feedback available, for instance,
through links or embedded drop-down lists. We assume that the user can in all cases also
respond by typing a textual question or response thus providing a flexible style of interaction.
Depending on where the widget is located it can either serve for specifying preferences in the
dialog or for changing, i.e. criticizing, general or item-specific values.</p>
      <p>The second dimension refers to the type of interactive element (widget dimension) integrated
in the conversation. Here, we consider the standard widgets which can be selected depending
on the purpose and constraints of the input. Buttons, checkboxes, drop-down lists or sliders
can be ofered for responding to system questions, or be attached to a recommended item to
show and modify its features. As a novel option in CRS, we propose to also make parts of the
system output interactive by embedding links, drop-down lists, or buttons directly in the textual
stream to let users react directly on questions, assumptions or suggestions made by the system.
In the present study, we investigated six diferent interaction techniques with respect to their
usefulness in CRS (Fig. 1 and 2). In the following section, we will discuss the diferent forms of
providing feedback and critique in more detail.</p>
      <sec id="sec-3-1">
        <title>3.1. Critiquing in CRS</title>
        <p>
          Item-based critiquing (Fig. 2 (2)) is a technique that has been gaining considerable interest
in recommender systems research. Since recommendations may not meet user preferences,
critiquing the features of a recommended item allows users to modify or incrementally refine
their preference in an interactive fashion, thus also increasing their control over the
recommendations provided [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. In a CRS, the critiquing approach can be extended to also providing
feedback on other concepts that appear in the conversation, such as the system’s assumptions
about the intended usage of an item, or any other aspect of the user model that is explicitly
mentioned in the system output,
        </p>
        <p>Integrating item-based critiquing in CRS could help users conveniently supplement or modify
preference when the first recommendation occurs. Besides the limitation of flexibility, another
drawback is the learning cost of interaction, users need time to adapt to this interaction.</p>
        <sec id="sec-3-1-1">
          <title>3.1.1. Inline critiquing</title>
          <p>We propose a novel inline critiquing interaction by which users can conveniently state and
modify their preferences. The basic idea is that once the system presents a recommendation
to the user, it also simultaneously generates a response, summarizing relevant item properties
as well as user preferences collected so far. In this summary, some keywords are marked that
can be directly modified by the user in the text. This avoids the problem that users would have
to refer verbally to previous system responses to criticize their content in a purely text-based
interface. This method can, in principle, be applied both to previous system outputs or user
inputs. We assume that this form of feedback provides advantages with respect to eficiency as
well as error-avoidance. In the following, we describe three diferent styles used in our study of
emphasizing interactive keywords and the widgets used for providing feedback.
1
2</p>
          <p>Underlined Text Fig. 3 (1) shows the first style of critiquing textual elements. The keyword
is underlined and shown in a diferent color, as is common for visualizing web links in HTML.
Once the user clicks on this link, a list pops up showing selectable—possibly popular—options.
Other feedback can be freely entered in an editable text field. The technique is well-known to
most users, although some might misinterpret it as a Web link.</p>
          <p>Highlighted Text The second style of indicating interactive keywords uses highlighting,
showing the text with a colored background (Fig. 3 (2)). Highlighting can easily attract user’s
attention and diferent from underlined text, it does not have hyperlink misunderstanding
problem. However, users may interpret it as indicating importance, not interaction.
Drop-down Button Fig. 3 (3) shows the third style of inline critiquing we investigated. This
form inserts a button in the text indicating a drop-down list which has equivalent functionality
as the other two styles. Buttons are easily recognizable as interactive objects and avoid the
problem of potential misinterpretation as is the case with the other two styles, but may look
awkward inside a running text.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>In this section, we describe an empirical comparison of the interaction patterns described in
the previous section and their meaningful combination in a mixed-modality CRS against a
conventional CRS. We conducted a user study to determine whether there exist preferences for
the diferent techniques and critiquing styles when engaging with conversational agents. In
particular, we intended to investigate whether users prefer a purely text-based CRS (TBCRS) or
a mixed-modality interaction CRS (MICRS). Besides, we were interested in identifying which
interaction modes are favored for diferent communicative tasks.</p>
      <sec id="sec-4-1">
        <title>4.1. Method</title>
        <p>In order to investigate these questions, we performed a study using video prototypes of the
diferent techniques described. To obtain a deeper understanding of users’ perception of the
techniques, our study was designed to capture quantitative data as well as qualitative feedback
from the participants. We split our study into three parts to investigate these research questions.</p>
        <sec id="sec-4-1-1">
          <title>4.1.1. Comparison between TBCRS and MICRS</title>
          <p>Since the focus of this study was centered on obtaining initial insights about users’ preferences
and the perceived usability of diferent interaction patterns, we did not yet implement a working
CRS with the interaction integrated. Instead, the evaluation was done by means of videos
showing the interactions based on a conversational recommender scenario in a fictitious bicycle
shop. We created two videos exemplifying diferent levels of interaction with a fictitious online
bicycle CRS. During the first part of the experiment, we presented these two videos to the
participants, showing conversations with the CRS in the form of a chatbot. In the videos, a
ifctitious user tried to find a suitable bicycle for himself by means of the chatbot. Both videos
were identical in their content, only the user’s interaction possibilities with the chatbot varied
as follows:
1. Text-based CRS (TBCRS): The conversation with the chatbot is solely text-based. Besides
the text-based method, there is no alternative option to respond to the chatbot’s questions.2
2. Mixed-modality Interaction CRS (MICRS): The conversation with the chatbot is both
textbased and via direct feedback using i.e. buttons, drop-down lists. Each of the interaction
methods described in Section 3 are exhibited. For all actions that the fictive user has to
perform in the conventional system through text input, there is an alternative interaction
possibility in this version. However, at all times, the user is able to enter simple textual
input instead.3</p>
          <p>Both videos were presented to all participants. Participants were allowed to pause, resume
and restart the videos at any time. There was no time limit for watching the videos. We
counterbalanced the order of the videos, resulting in a within-subject design.</p>
          <p>After each video, participants were asked to fill in a questionnaire. If not indicated otherwise,
all questionnaire items had to be answered on positive 1–5 Likert response scales. For this
purpose, we asked them to imagine themselves interacting with the chatbot shown and to
evaluate the interaction possibilities. To assess user interface satisfaction, we applied the factors
of “overall reaction to the software” from the QUIS questionnaire [26], consisting of six items.
These items were assessed by means of a polarity profile. Besides, we constructed nine items that
2Video of text-based CRS interaction: https://intsys.info/tbcrs
3Video of mixed-modality CRS interaction: https://intsys.info/micrs
were specifically intended to evaluate the interaction methods shown. Furthermore, we assessed
domain knowledge of participants with self-constructed items and collected demographic data.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Interaction Methods in Detail</title>
          <p>During the second part of the study, we sought to obtain more detailed feedback on the six
diferent interaction methods described in Section 3: Free text, buttons, checkboxes, sliders,
item-based critiquing and inline critiquing. Here, we successively showed participants the
individual interaction methods as screenshots. All interaction methods were already shown in
the videos during the first part of the study, thus participants have already seen the interaction
process with the respective method. We asked them to rate each interaction opportunity by
means of the self constructed questions regarding enjoyability, supportiveness, eficiency and
precision. Besides, we asked a specific question regarding critiquing eficiency for the interaction
methods free-text, inline critiquing as well as for the item-based critiquing method. Additionally,
we asked what they particularly liked or disliked about each interaction method. This optional
questions were open-ended.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.1.3. Inline critiquing styles</title>
          <p>Finally, we aimed to identify the preferred presentation style for the inline critiquing method.
Therefore, we asked participants to choose one of three diferent designs described in Section
4.1.3, as the most appropriate one for directly modifying features in the text. Additionally, they
were asked to briefly describe why they selected a particular style. These two question were
optional.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Participants</title>
        <p>We recruited 70 participants using Prolific, 4 a tool commonly used for academic surveys [27],
of whom 63 finished the study. We pre-selected Prolific users based on the following criteria
to maximize quality: (1) participants should be fluent in English; (2) their success rate should
be greater than 95 %; and (3) the survey should not be conducted on smartphones or tablets to
ensure that the interaction methods shown in the videos and screenshots can be recognized
easily. The average duration of the survey was 13.18 minutes (SD = 3.28) and each participant
received a compensation of 1.25£ if they successfully completed the survey. In our analysis, we
only considered participants who watched the videos completely, leaving us with 54 participants.
Demography Out of 54 participants, 32 were female. Their age ranged from 18 to 81 (M =
35.2, SD = 14.29). The majority had a university degree (46.3 %), 27.8 % had a higher education
entrance qualification and 11.1 % had a general certificate of secondary education. The majority
originated from the United Kingdom (85.2 %). All other participants originated from South
Africa (3.7 %) and further countries (11.4 %). The domain knowledge of participants was rather
low (M = 2.16, SD = 1.12).</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Results</title>
        <p>We present the quantitative and qualitative results of the comparison between the two diferent
video prototypes: TBCRS and MICRS (Section 4.3.1), followed by addressing specific quantitative
and qualitative evaluations of each interaction method separately (Section 4.3.2). Finally, we
detail the user comments regarding the proposed inline critiquing styles described in Section
4.3.4. Therefore, we will quote exemplary statements made by participants.</p>
        <sec id="sec-4-3-1">
          <title>4.3.1. Comparison between TBCRS and MICRS</title>
          <p>Except for the “rigid/flexible” item, the MICRS version shown in the video prototype received
better average ratings in all elicited items from the QUIS questionnaire. For the item
“frustrating/satisfying”, we identified a significant diference between the two tested versions. For all
other factors tested, we did not observe any significant diferences between the two conditions.
Furthermore, we could not identify any significant correlations between the values reported
here and those reported in the remainder of this work and domain knowledge.</p>
          <p>Comparing the comments on each of the video prototypes, it seems that the participants
appreciate that it feels “natural to interact” with the virtual agent in the TBCRS version: “I liked
this advisor [TBCRS] better as it made me feel like I’m talking to a real person”. The virtual
agent appears “more human” and “it felt more personal” than the MICRS variant showcased.
Conversely, other participants worry that it might be more dificult to express requirements
because “[. . . ] people don’t always know exactly what they want and it would be dificult to
articulate properties eficiently” and “with free text input, it is dificult to know what answer
the chatbot is probing for, which can lead to frustration.” In addition to the “potential for higher
error margins for misunderstanding the customer,” some participants also mentioned dificulties
in “[. . . ] reasoning for recommendations because it is less clear what information is considered.”</p>
          <p>The positive comments on the video prototype of the MICRS mainly refer to the increased
eficiency (“It was very eficient and time-saving”; “I like how easy it is to fine-tune my preferences”)
and possibility of easier specification of personal preferences: “I liked the given options, which
saved time and gave ideas you might not have necessarily thought of.” Also, the participants
perceived the interaction options used as “straightforward and self explanatory.”</p>
          <p>The reasons why participants disliked this prototype were primarily that its options were
rather “specific and seemed less flexible.” Also, some participants were not aware that besides
the suggested interaction methods, they could continue to provide open text input. Few of them
stated that it “did not feel authentic.”</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>4.3.2. Free Text vs. GUI-Responses</title>
          <p>To compare the diferent interaction methods, we combined the assessed values for the
GUIbased interaction methods (checkboxes, buttons and sliders) into one score. We did this because
the text-only approach may be used universally, but not every GUI-based interaction method is
equally suitable for all response types.</p>
          <p>We conducted a paired t-test to compare text-based input with GUI-based methods. As shown
in Tab. 2 the GUI-based interaction method for responding to the virtual agent were rated
consistently higher. However, we only observed a significant diference for enjoyability.</p>
          <p>In addition to the quantitative data, we also analyzed the participants’ comments for each of
the interaction styles presented. Tab. 4 shows a meta analysis of the comments separated by
each interaction style. On average, there were 35.6 positive comments followed by 13.0 neutral
comments for the GUI-responses. On average, there were 4.7 negative comments for these
interaction styles. All interaction methods received fewer negative comments than the solely
text-based interaction. In fact, more than half of the comments on the text-based interaction
were neutral or negative (Tab. 4).</p>
          <p>Participants commented negatively on the text-only input that “with open-ended responses
there are so many ways to respond that I probably would be unsure that my answer would be
interpreted correctly.” Others commented positively that “it felt more like writing with a real
person” and that they “preferred this kind of input because it is more descriptive.”</p>
          <p>Positive comments concerning the GUI-responses often referred to the simplicity (“very clear
and helpful”, “straightforward”) and precision of the input: “It is easy to select what you want
instead of typing and potentially making a typo which could impact results.” However, other
participants noted that these interactions were “potentially limiting” and “narrow”, and that
“they may not cover all possible responses.” For other participants, these forms of interaction
were “[. . . ] too similar to conventional filtering systems in online stores.”
aItem description: This interaction gives me the possibility to criticize the displayed features of the recommendations in an easy and eficient way.</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>4.3.3. Free Text vs. Critiquing Methods</title>
          <p>Additionally, we performed multiple Repeated Measures ANOVAs to compare the text-based
interaction with the inline and item-based critiquing variants. The item-based variant was
assessed consistently better than the text-only and inline critiquing variants (Tab. 3).</p>
          <p>Regarding enjoyability, we found a significant diference between the tested conditions (Tab.
3). Post-hoc analysis revealed a significant diference (  = .003) between the interactions
freetext and item-based critiquing (− 0.667, 95 %–CI[− 1.14, − 0.19]). Post-hoc tests performed here
and in subsequent results were Bonferroni-adjusted.</p>
          <p>In terms of supportiveness, results indicate a significant diference between the conditions
(Tab. 3). Again, post-hoc tests revealed a significant diference (  = .030) between the free-text
interaction and item-based critiquing (− 0.389, 95 %–CI[− 0.75, − 0.03]).</p>
          <p>Furthermore, the eficiency of the interaction variants is evaluated significantly diferently
(Tab. 3). Compared to the free text and the inline critiquing variant, the item-based critiquing
is rated better. However, the performed post-hoc tests only revealed a significant diference
( = .030) between item-based and inline critiquing (0.389, 95 %–CI[0.028, 0.75]).</p>
          <p>Although item-based had higher means for the last two items (precision and critiquing
eficiency) shown in Tab. 3 than inline critiquing and text-based interaction, no significant
diferences were detected.</p>
          <p>The received comments regarding the diferent critiquing methods were also rather positive.
Item-based critiquing received more positive comments than inline critiquing. Overall, inline
critiquing received the most neutral comments (Tab. 4).</p>
          <p>Regarding the item-based critiquing option, participants liked the “[. . . ] ability to directly
select diferent options” and to be able “[. . . ] to criticize options directly based on the given
items.” Others appreciated that “previous specifications were already taken into account” as
well as “[. . . ] being able to specify further features.” However, others were critical and noted
that “less technological afine users could possibly be overwhelmed.”</p>
          <p>Some participants felt that inline critiquing was “not as seamless as the other GUI options”. In
addition, some participants were critical that this interaction option might not be understood by
everyone: “If people are not familiar with the Internet (e.g. the older generation) they may not
understand how to use this.” Others experienced a similar problem in terms of the representation:
“I would not have realized they were drop-down lists and assumed they were links [. . . ].”</p>
        </sec>
        <sec id="sec-4-3-4">
          <title>4.3.4. Inline Critiquing Styles</title>
          <p>Next, we present the assessment results of the proposed inline critiquing styles. Of the 54
participants, 50 participants specified a favored style for inline critiquing. From this group, the
majority (60.0 %) preferred the drop-down button (Fig. 3 (2)). Less often, the other two variants
highlighted text (24.0 %; Fig. 3 (2)) and underlined text (16.0 %; Fig. 2 (1)) were chosen. The other
4 participants did not indicate a preference, but still provided comments on the styles shown.</p>
          <p>A  2-Goodness-of-Fit-Test shows that there are significant diferences between the observed
frequencies ( 2(2,  = 50) = 16.485,  &lt; .001). Post-hoc analysis revealed significant
differences between drop-down button and underlined text styles ( = .003), and between the
drop-down button and highlighted text ( = .024). However, we could not detect any significant
diference (  &gt; .999) between the underlined text and highlighted text styles.</p>
          <p>Comments from those favoring the drop-down button stated that it is “[. . . ] obvious that one
can click and alter something [. . . ]” in this variation. Others noted that it is “[. . . ] particularly
obvious that there are additional options” and “[. . . ] it cannot be confused with a hyperlink.”
Other participants stated, that “it appears as an option whereas the other styles might be missed.”</p>
          <p>Participants who favored the highlighted text style argue that it “[. . . ] stands out most” and
it is “easiest” and “clearest to see without no thought, because of the words highlighted with
the background color.”</p>
          <p>Those participants who favored the variant with the underlined text justified their decision
by stating that “[. . . ] it looks like a hyperlink and thus makes it clearer that one can click on it.”
Others noted that this inline critiquing style “is simple and less cluttered.”</p>
          <p>Participants who did not nominate a favorite justified this by stating that all three styles
shown are “useful when there are a lot of options to choose from” or that all variations “are
clearly illustrating options.”</p>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Discussion</title>
        <p>In this section, we discuss the results of our empirical study. Therefore, we first elaborate on
the comparison between the two conditions and subsequently discuss the findings regarding
the various interaction methods.</p>
        <sec id="sec-4-4-1">
          <title>4.4.1. Comparison between TBCRS and MICRS</title>
          <p>First, the comparison of the two video prototypes showed slightly better scores for the MICRS
condition on most factors, however, using paired t-test, we were only able to observe a significant
diference for one factor (frustrating/satisfying). These results are in line with the participants’
comments. Here, they rated the MICRS as less error-prone. They also noted that it is simpler to
recognize possible options and to specify their preferences. Similarly, the ability to respond
quickly to questions is also rated slightly better. As expected, entering feedback through buttons
or drop-down lists is faster and easier in comparison to formulating text—presuming that
adequate response mechanisms are available respectively that the displayed interaction element
is purposeful.</p>
          <p>The reason why the MICRS condition is perceived as more stimulating than the comparative
condition may also be explained by the observation that participants discover aspects and
potential options that they had not considered before. Conversely, the provided options may
also lead to the MICRS condition being seen as more rigid. Some participants noted that it
would be nice to still be able to submit open text input if, for instance, none of the available
options apply. Although this was possible in the prototype video shown, the option may not
have been obvious enough. However, in a real system, a text-based input option should remain
available to avoid restricting users unnecessarily. This, in turn, can also ensure a more flexible
conversation and exploit the potential strengths of an open-ended CRS.</p>
          <p>While participants rated the text-only condition in the questionnaires as easier to learn, this
partially contradicts the comments provided. Although it may be obvious how to interact with
a purely text-based system, it may still be necessary to learn how the virtual agent interprets
the input to avoid misunderstandings.</p>
          <p>The ratings regarding comprehensibility and interpretability (“Messages from the chatbot
which prompt for user inputs are clear”; “It is easy to understand why the chatbot is showing me
the recommendations”) are consistent with the provided comments. Here, the MICRS condition
is rated better. Due to the highlighted features within the text along with the explicit inputs
via i.e. buttons or drop-down lists, it is clearly comprehensible which information the system
uses for providing recommendations. Compared to the text-only approach, users receive more
visual information throughout the entire conversation. However, we assume that the tendencies
observed in our video prototype based analysis could be manifested in an interactive prototype
with which participants can interact and thus leverage the features themselves.</p>
        </sec>
        <sec id="sec-4-4-2">
          <title>4.4.2. Free Text vs. GUI-Responses</title>
          <p>While comparing the two systems, we asked the participants to rate the individual forms
of interaction shown. Unlike in text-based interaction, an appropriate GUI option must be
presented by the system depending on the logical type of response requested. Since the system
demonstrated in the video prototype utilized a set of diferent GUI elements, we aggregated the
GUI response options (buttons, checkboxes and sliders) and compared them to the text-based
interaction. In the video prototype, we always provided the appropriate GUI response methods.</p>
          <p>In terms of enjoyability, the GUI based interaction methods were rated significantly higher,
which is in line with the results of the general evaluation. One explanation might be that the
use of various input methods is more interesting and thus the interaction is more enjoyable.</p>
          <p>Concerning the factor supportiveness, no significant diferences were found, although there
was a tendency for GUI responses to be rated better. The comments discussed in the previous
section support these findings: Non-textual interaction methods seem to support users better,
assuming that appropriate options are available.</p>
          <p>For the last two factors tested, there were also positive tendencies with regard to the GUI
responses. These are rated slightly better than free-text input in terms of eficiency and
simplicity, although these diferences are not significant. This also corresponds to the participants’
comments. As long as the appropriate answer choice can be provided directly, a GUI interaction
is considered more eficient since only one click is needed.</p>
          <p>Although it is reasonable to assume that appropriately displayed GUI elements would provide
more precise feedback, results were comparable. We suspect this is because free text input
allows requirements to be expressed that are not presented as options in GUI responses. In
case users already have a clear idea of the desired item and which requirements they intend to
communicate to the system, they may not require guidance in the form of GUI response options,
but can respond more precisely and flexibly with a free, textual interaction.</p>
        </sec>
        <sec id="sec-4-4-3">
          <title>4.4.3. Free Text vs. Critiquing Methods</title>
          <p>When comparing the text-based input with the two critiquing methods, results indicated that the
item-based method was rated better than the purely text-based input method in all tested factors.
With regard to the factors enjoyability, search support and ease of articulating requirements,
the diferences between these two interaction methods were significant.</p>
          <p>We assume that it is more enjoyable for users to give feedback directly based on specific
items than to articulate them in a text. Perhaps when critiquing features of a particular item,
the implications of that critique are less ambiguous.</p>
          <p>Although the inline critiquing was rated better than the text-only method in some aspects,
the diferences were rather minor. We suspect that the participants were not entirely aware of
how this interaction method was supposed to work. This may be due to the chosen visualization,
but also to the fact that this novel interaction method was not suficiently explained within the
video prototypes.</p>
          <p>While we did not find any significant diferences in the other factors tested, we suspect that
the alternatives to text-only interaction may still have advantages. The video prototype method
we chose was possibly not capable of identifying them clearly. Here, however, we must take
into account the fact that text only feedback may provide more accurate responses if the options
provided by the CRS are not what the user expects.</p>
        </sec>
        <sec id="sec-4-4-4">
          <title>4.4.4. Inline Critiquing Styles</title>
          <p>Finally, we discuss the results for the diferent inline critiquing styles. Although one might
assume that all three styles perform similarly, since they all have a prompting character, the
participants clearly preferred the drop-down button. We assume that this is mainly due to
familiarity with this technique and its clear afordance for changing values.</p>
          <p>In contrast, emphasizing active parts of the text by underlining may be confused with a
hyperlink—causing users to assume that clicking on it will forward them to another page. This
was also reflected in the participants’ comments.</p>
          <p>Highlighting text with a colored background may not convey clearly enough that it is possible
to interact with it and modify options. Rather, users might interpret this highlighting as an
indication of importance or as a reference to a help text that appears when hovering the mouse
pointer over the highlighted word.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions and Future Work</title>
      <p>We investigated a mixed modality interaction approach for CRS and could show in a user study
that is positively evaluated by the participants who appreciated the benefits of using diverse
interaction techniques within a CRS. Also, the possibility to criticize individual features of the
recommended items directly in a CRS as well as the proposed inline text critiquing method was
evaluated positively.</p>
      <p>Additionally, the non-textual interaction methods were particularly favorably evaluated.
Overall, this study suggests that text-only interaction might not be optimal for creating a positive
user experience in a CRS. Instead, a combination of diferent interaction methods is probably
preferable. By summarizing and emphasizing relevant item features and user preferences in the
text, the explanatory value of CRS responses are probably enhanced, increasing transparency of
the system. Enabling users to modify terms directly in the output may also increase the sense of
user control. Considering that CRS systems should be accessible from the very first use without
detailed instructions, we believe the approach is promising and aim to focus in future work
on an easy-to-understand embedding and traceability of changes resulting from applying the
inline text and feature critiquing mechanisms.</p>
      <p>As a limitation of this work, we are aware that evaluating video prototypes cannot substitute
interacting with realistic interactive prototypes. Therefore, we intend to investigate the use of
mixed-modality interaction in CRS by implementing a fully interactive prototypes in future work.
A particular challenge for building mixed-modality CRS is the question how the interactive
options ofered to the user can be derived automatically. Potential approaches might be based
on leveraging knowledge graph data or information extracted from item descriptions or reviews.
Also, suitable response generation techniques are needed that summarize the features the user
is likely to criticize in the next interaction step. Furthermore, it will be interesting to explore
techniques that can be applied to automatically decide which interaction method is most suitable
in a certain conversational context.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>A survey on conversational recommender systems</article-title>
          ,
          <source>ACM Computing Surveys</source>
          <volume>54</volume>
          (
          <year>2021</year>
          ). doi:
          <volume>10</volume>
          .1145/3453154.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Conversational Recommender System, Association for Computing Machinery, New York, NY, USA,
          <year>2018</year>
          , p.
          <fpage>235</fpage>
          -
          <lpage>244</lpage>
          . doi:
          <volume>10</volume>
          .1145/3209978.3210002.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Weizenbaum</surname>
          </string-name>
          ,
          <article-title>Eliza-a computer program for the study of natural language communication between man and machine</article-title>
          ,
          <source>Commun. ACM</source>
          <volume>9</volume>
          (
          <year>1966</year>
          )
          <fpage>36</fpage>
          -
          <lpage>45</lpage>
          . doi:
          <volume>10</volume>
          .1145/365153. 365168.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ravishankaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chandrasekaran</surname>
          </string-name>
          ,
          <article-title>A survey of design techniques for conversational agents</article-title>
          ,
          <year>2017</year>
          , pp.
          <fpage>336</fpage>
          -
          <lpage>350</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-10-6544-6_
          <fpage>31</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kolomvatsos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hadjiefthymiades</surname>
          </string-name>
          ,
          <article-title>Facing the cold start problem in recommender systems</article-title>
          ,
          <source>Expert Syst. Appl</source>
          .
          <volume>41</volume>
          (
          <year>2014</year>
          )
          <fpage>2065</fpage>
          -
          <lpage>2073</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.eswa.
          <year>2013</year>
          .
          <volume>09</volume>
          .005.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zha</surname>
          </string-name>
          ,
          <article-title>Functional matrix factorizations for cold-start recommendation</article-title>
          ,
          <source>in: Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          , SIGIR '11,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2011</year>
          , p.
          <fpage>315</fpage>
          -
          <lpage>324</lpage>
          . doi:
          <volume>10</volume>
          .1145/2009916.2009961.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>A comparative user study on rating vs. personality quiz based preference elicitation methods</article-title>
          ,
          <source>in: Proceedings of the 14th International Conference on Intelligent User Interfaces</source>
          ,
          <source>IUI '09</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2009</year>
          , p.
          <fpage>367</fpage>
          -
          <lpage>372</lpage>
          . doi:
          <volume>10</volume>
          .1145/1502650.1502702.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Personalizing recommendation diversity based on user personality, User Modeling and User-Adapted Interaction 28 (</article-title>
          <year>2018</year>
          )
          <fpage>237</fpage>
          -
          <lpage>276</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s11257-018-9205-x.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          , L. Chen,
          <article-title>Predicting user intents and satisfaction with dialogue-based conversational recommendations</article-title>
          ,
          <source>in: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization</source>
          , UMAP '20,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2020</year>
          , p.
          <fpage>33</fpage>
          -
          <lpage>42</lpage>
          . doi:
          <volume>10</volume>
          .1145/3340631.3394856.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M. F. McTear</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Allen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Clatworthy</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ellison</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Lavelle</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>McCafery</surname>
          </string-name>
          ,
          <article-title>Integrating flexibility into a structured dialogue model: Some design considerations</article-title>
          ,
          <source>in: 6th International Conference on Spoken Language Processing</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>Critiquing-based recommenders: Survey and emerging trends, User Modeling and User-Adapted Interaction 22 (</article-title>
          <year>2012</year>
          )
          <fpage>125</fpage>
          -
          <lpage>150</lpage>
          . doi:
          <volume>10</volume>
          .1007/ s11257-011-9108-6.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Shimazu</surname>
          </string-name>
          ,
          <article-title>ExpertClerk: Navigating shoppers' buying process with the combination of asking and proposing</article-title>
          ,
          <source>in: Proceedings of the 17th International Joint Conference on Artificial Intelligence -</source>
          Volume
          <volume>2</volume>
          , IJCAI'
          <fpage>01</fpage>
          , Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <year>2001</year>
          , p.
          <fpage>1443</fpage>
          -
          <lpage>1448</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>A comparative study of compound critique generation in conversational recommender systems</article-title>
          , in: V. P. Wade,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ashman</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          Smyth (Eds.),
          <source>Adaptive Hypermedia and Adaptive Web-Based Systems</source>
          , Springer Berlin Heidelberg,
          <year>2006</year>
          , pp.
          <fpage>234</fpage>
          -
          <lpage>243</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>Evaluating critiquing-based recommender agents</article-title>
          ,
          <source>in: Proceedings of the 21st National Conference on Artificial Intelligence -</source>
          Volume
          <volume>1</volume>
          , AAAI'
          <fpage>06</fpage>
          , AAAI Press,
          <year>2006</year>
          , p.
          <fpage>157</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>