<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Factors Influencing the Perceived Meaningfulness of System Responses in Conversational Recommendation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ahtsham Manzoor</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wanling Cai</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dietmar Jannach</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Trinity College Dublin &amp; Lero, College Green</institution>
          ,
          <addr-line>Dublin 2</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Klagenfurt</institution>
          ,
          <addr-line>Universitätsstraße 65-67, Klagenfurt am Wörthersee, 9020</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Conversational recommender systems (CRS) support users in finding satisfying items via multi-turn dialogs and have recently attracted increasing attention. Many research eforts have been made in producing quality responses including item recommendations, which can be jointly assessed in a user-centric manner using a subjective criterion, i.e., meaningfulness. However, human perceptions of meaningfulness are complex and can be nuanced, as individual users have their own preferences over conversations, and their perceptions can be afected by various factors, such as users' personal characteristics, system functionalities like informativeness of the generated responses, and dialog context. To better design CRS adapted to user needs and context, this work investigates the impact of three types of factors (i.e., userrelated, system-related, and context-related) on users' perceived meaningfulness of system responses by analyzing a within-subject user study (N=90) data. Results indicate that users' domain knowledge and the informativeness of system responses positively influence users' perceived meaningfulness of system responses. A deep investigation reveals that users with previous chatbot experience tend to expect highly informative conversations. Moreover, older users seem to be less satisfied with lengthy dialogs.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Conversational recommendation</kwd>
        <kwd>personal characteristics</kwd>
        <kwd>interaction context</kwd>
        <kwd>evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Conversational recommender systems (CRS) assist users in finding items of interest and support
their decision-making process while conversing with the system in natural language. Technically,
regarding the development of a CRS, we observe two streams of works, i.e., retrieval-based and
generation-based approaches to CRS [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, unlike traditional recommender systems that
mainly provide one-shot recommendations (e.g., a ranked list of items), both kinds of approaches
facilitate “free-style” multi-turn conversations between a user and system. From the interaction
standpoint, a user can express her preferences to the system (e.g., “Can you suggest any good
sci-fi movies?” ), and the system in response can proactively suggest recommendations or takes
the lead to elicit further user preferences, (e.g., “Are you looking for recent or old movies?” ),
thereby supporting the recommendation process via producing efective dialogs, see also [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Regarding the evaluation, a CRS can be evaluated using both computational (“ofline” )
experiments or studies involving humans, see e.g., a survey on the evaluation of CRS [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Given
that various user-centric metrics to evaluate item recommendations are available [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], human
evaluations of CRS predominantly centered around assessing only linguistic quality aspects
like fluency, consistency, or naturalness [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ], and if the quality of item recommendations is
adequate in the given dialog context seems missing. For example, the system-response “have
you seen Black Panther (2018)?” to a user utterance asking for romantic movie recommendations
appears obscure [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ]. In addition, approaches that determine a relative ranking of diferent
systems via human evaluators, as was done in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], do not inform us if the highly ranked system
would be useful in practice [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>
        To address such limitations, the concept of meaningfulness, as an evaluation criterion of
system responses, was introduced in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In this approach, the authors specifically relied on
a single subjective metric for assessing the overall quality of system responses, including
recommendations, as shown in Fig 1. Specifically, perceived meaningfulness refers to assessing
the recommendation quality, response appropriateness, and coherency in a given dialog context.
An example of when a response should be meaningful can be when a user expresses interest in
horror movies. In such cases, a response like “The Conjuring (2013) is a really good one” should
be considered meaningful. More such cases and examples can be found in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Recent works on CRS however have shown that human perceptions of quality may difer
between groups of people and can impact user trust and satisfaction [
        <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref17">13, 14, 15, 16, 17</xref>
        ]. For
example, users with a higher level of domain knowledge may benefit more from conversational
recommendation interactions due to their ability to express their preferences better than novices
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Similarly, more aware users may expect shorter but more eficient dialogs. In addition,
Araujo et al. [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] suggest that personal characteristics and context can be linked to diferent
perceptions of automated decision-making. Likewise, individual perceptions of meaningfulness
can be complicated and nuanced. For example, in a study on CRS responses regarding cases about
seekers’ specific questions [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], participants prefer system responses that provide appropriate
recommendations over responses that are entirely irrelevant. Yet, it remains unclear how
diferent aspects corresponding to the user, system, and the context in which a dialog happens
influence human perceptions of the meaningfulness of system responses.
      </p>
      <p>Therefore, in this work, we identify and categorize diferent aspects into one of the three
categories, i.e., user, system, or context, and investigate the following research questions (RQs).</p>
      <p>RQ1: How do user-related, system-related, and context-related factors afect users’
perceived meaningfulness of responses of a CRS?
RQ2: How do various factors interact to afect the perceived meaningfulness of responses
in a CRS?</p>
      <p>
        To address these questions, based on the literature on CRS, in particular works like [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]
and [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], we carefully derived a set of user-related characteristics such as gender, age, domain
knowledge, or prior chatbot experience. To study system-related factors, we rely on two CRS,
i.e., KBRD [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and CRB-CRS [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Finally, inspired by [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], we derived a contextual factor, i.e.
dialog discourse length, which might have an impact on meaningfulness. Additional factors like
dialog initiative strategy could be taken into account as well, but due to the non-availability of
ReDial [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] data annotations, we consider that to be beyond the scope of the current work.
      </p>
      <p>Overall, we examined user feedback data collected in an online study conducted via the MTurk
crowdsourcing platform with 90 participants. Specifically, the participants were presented with
a dialog situation curated from the ReDial dataset where two humans, one with a user role and
the other as a recommender, discuss movie recommendations. The dialog situation always ends
with a user (or: seeker) utterance followed by responses by two CRS that appeared in random
order, as shown in Figure 1. Diferent system responses are evaluated in parallel. We note that
this is in principle not necessary to understand the given RQs and may lead to implicit bias; but
here the intuition was to collect user feedback through a smaller number of user interactions.</p>
      <p>Our findings reveal that users with domain knowledge and informativeness of responses
positively influence their perceptions of meaningfulness. Also, users with prior chatbot
experience tend to have higher expectations and require more advanced features for satisfactory
CRS performance. In addition, we found that the user’s age afects the relationship between the
dialog discourse length and perceived meaningfulness. CRS practitioners should aim to provide
more concise responses for older users as relatively young users seem reasonably satisfied. To
our best knowledge, this is the first work investigating the factors that influence the perceived
meaningfulness of responses in the context of conversational recommendation. We believe our
ifndings will contribute to the research on AI dialog systems and facilitate improved CRS design
in terms of a more personalized and inclusive user experience.</p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORK</title>
      <p>
        Conversational Recommender Systems Conversational recommender systems (CRS) that
use natural language ofer free-style conversations with users and help them make online
choices. We observe two main types of CRS: language generation-based and retrieval-based.
Language generation-based systems are popular due to their ability to incorporate new contexts,
but recent studies have found that they often generate identical responses [
        <xref ref-type="bibr" rid="ref11 ref8">11, 8</xref>
        ].
Retrievalbased methods on the other hand fetch appropriate responses from recorded dialog datasets,
and adapt them to the ongoing dialog context. Such responses are usually informative, fluent,
and semantically meaningful as they were originally made by humans. However, retrieval
approaches struggle with unseen dialog contexts or open-ended questions, see also [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] for a
comparison between generation and retrieval-based approaches in NLP. To study system-related
aspects, we consider two recent open-source CRS (KBRD and CRB-CRS) as representatives of
both streams of work. These systems have shown significant performance compared to others
and were published in high-quality research venues.
      </p>
      <p>
        Technically, KBRD [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a generation-based CRS, relies on a sequence-to-sequence
encoderdecoder Transformer framework [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] while mapping user domain concepts to an additional
knowledge graph, named DBpedia [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] to aid informativeness in the generated responses.
The authors chose the Transformer framework over HRED [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] due to its better performance
in various NLP tasks, such as Q&amp;A or machine translation [
        <xref ref-type="bibr" rid="ref27 ref28 ref29">27, 28, 29</xref>
        ]. On the other hand,
CRB-CRS is a contextual retrieval-based system that utilizes a dialog corpus to fetch, adapt
and integrate item recommendations given the user utterance or dialog history as input. Note
that both CRS are developed in the context of the ReDial dataset, see more details about the
collection procedure of the dataset in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which provides further grounds to construct fair
analyses in our study.
      </p>
      <p>
        Evaluation A CRS can be evaluated based on various quality dimensions, including system
efectiveness, eficiency, conversation quality, and subtask efectiveness, see also a survey on
CRS evaluation in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Quality measurements can be assessed through ofline experiments or
studies involving humans. Computational experiments typically evaluate system efectiveness
using metrics like Recall or RMSE [
        <xref ref-type="bibr" rid="ref30 ref31 ref32">30, 31, 32</xref>
        ], and linguistic quality measures such as BLEU
and Perplexity [
        <xref ref-type="bibr" rid="ref33 ref34 ref35 ref5">33, 34, 5, 35</xref>
        ]. However, the representativeness of such measures for human
perceptions remains unresolved [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In the CRS domain, human studies mainly assess linguistic
quality aspects using measures like informativeness, naturalness, persuasiveness, or
engagingness [
        <xref ref-type="bibr" rid="ref30 ref36 ref5 ref6 ref7">7, 6, 5, 30, 36</xref>
        ]. However, studies show language generation-based CRS have limited
capability to generate new sentences, therefore evaluating linguistic aspects primarily assesses
the quality of human-produced utterances, see also [
        <xref ref-type="bibr" rid="ref11 ref8">11, 8</xref>
        ].
      </p>
      <p>
        From the overview of predominant evaluation approaches to CRS, it is evident that various
quality constructs as a proxy to the efectiveness of a CRS can be practically realized by human
perceptions, which can be influenced by various factors such as context, beliefs, emotions, and
personal diferences [
        <xref ref-type="bibr" rid="ref37 ref38 ref4">37, 38, 4</xref>
        ]. It poses further challenges for researchers to explore factors
that assert and predict users’ needs accurately. In that regard, the perceived meaningfulness of
conversations can represent a practical approach to assess the overall efectiveness of a CRS in
a human-centered fashion to estimate how human perceptions vary along various user, system,
and contextual dimensions.
      </p>
      <p>
        User, System and Context-related Factors Among user-related aspects, inspired by
previous works [
        <xref ref-type="bibr" rid="ref13 ref39 ref40">39, 40, 13</xref>
        ], we consider personal features such as gender, age, education, prior
chatbot experience, and domain knowledge about movies that refer to enduring characteristics
pertaining to individuals’ cognition, emotions, and attitude [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ]. For example, in [
        <xref ref-type="bibr" rid="ref41">41</xref>
        ], it is
revealed that younger users prefer less human-like chatbots, while older users prefer more
human-like characteristics. Therefore, we aim to identify in what ways and to what degree
users’ personal diferences have an impact on perceived meaningfulness.
      </p>
      <p>
        When considering system-related aspects, a CRS can be equipped with several attributes
[
        <xref ref-type="bibr" rid="ref21 ref3">21, 3</xref>
        ], for example, enhancing transparency via explanations [42, 43]. Similarly, informativeness
(amount of domain concepts [
        <xref ref-type="bibr" rid="ref30 ref35 ref36 ref7">7, 30, 36, 35</xref>
        ]) is critical for the consistency and relevance of
responses. Many proposals in CRS integrate additional domain knowledge to manifest
domainrelated concepts in the CRS output and thereby evaluate its efectiveness through ofline and
human evaluations [
        <xref ref-type="bibr" rid="ref30 ref35 ref36 ref7">7, 30, 36, 35</xref>
        ]. In this context, both KBRD and CRB-CRS incorporate
additional knowledge and metadata into CRS. Therefore, we rely on the item ratio in responses,
as was done in [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] to estimate the extent of informativeness. Moreover, response length (i.e.,
token count) can be an important system factor that can afect human perceptions in chatbot
interactions. For example, studies have shown that response length can impact engagement,
satisfaction, and perceived quality of responses [
        <xref ref-type="bibr" rid="ref36 ref5">5, 36</xref>
        ]. To this end, CRB-CRS implemented
Response length as a system design parameter, heuristically curated based on ReDial dataset
statistics, see also [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], while KBRD does not explicitly set a parameter for response length.
      </p>
      <p>
        Finally, we explore context-related factors that relate to the specific context in which a dialog
takes place, e.g., diferent stages of the dialog discourse. Dialog discourse length is an important
consideration when assessing the usability of CRS, since certain CRS may excel during the
initial phases of dialog, where chit-chat and preference elicitation interactions are expected
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Alternatively, other CRS may perform better during the middle stage, where we typically
expect item recommendations, see also an overview of user intent and system actions in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
In addition, users may expect concise dialog, as lengthy conversations might lead to users being
bored and dissatisfied [44].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiment Design</title>
      <p>
        Like previous studies, e.g., in [
        <xref ref-type="bibr" rid="ref10 ref6 ref7">6, 7, 10</xref>
        ], we compared the quality of responses from both
CRBCRS and KBRD through an online user study. The details of the study procedure and participants
are as follows.
      </p>
      <p>
        Procedure A web interface is particularly developed for this study, where after the informed
consent and data privacy statements, participants were presented with a dialog situation, starting
from the first utterance and always ending by the user utterance, see also Figure 1. The dialog
situations are automatically randomly curated from a set of 70 dialogs randomly selected from
the ReDial [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] dataset. Dialog continuations (or responses) to the last user utterance from
two diferent systems generated by KBRD and CRB-CRS were shown in random order. Study
participants were tasked to rate the quality of responses independently in terms of meaningfulness
of each response on an absolute 5-point scale ranging from “Entirely meaningless” to “Perfectly
meaningful”. Note that when considering perceived meaningfulness, it entails evaluating the
quality of recommendations, appropriateness of responses, and coherence within the context of
a conversation.
      </p>
      <p>Nonetheless, prior to the response rating task, participants were given instructions and
examples of when a response should be meaningful. Precisely, we provided the following
instructions to our study participants.</p>
      <p>1. A response should be logical by the chatbot, for instance, the chatbot should make a
recommendation when the user asks for one.
2. A response should be complete and grammatically correct.
3. If a recommendation is made by the chatbot, it has to match the user’s stated preferences.</p>
      <p>If, for example, a user is looking for a funny movie, mostly funny movie recommendations
are meaningful. Please make this assessment based on your knowledge, and expectations in
the context of mentioned movies in dialog situations. You may always look up services like
IMDb to check, e.g., about the movie genres and plots. Note that movie titles are shown in
double quotes in both the dialog situations and corresponding responses.
4. If the chatbot response is not a movie recommendation, you are supposed to rate the
meaningfulness of the response as a reply to the user’s last statement keeping in mind the context
of the dialog situation.</p>
      <p>
        Despite the provided instructions, participants may still have varied perceptions of response
meaningfulness. Nonetheless, we believe that allowing for subjective interpretation is more
suitable for evaluating user perceptions in natural language interactions, rather than imposing
strict evaluation metrics. During the experiment, each participant completed 10 trials, with
one (hardcoded) randomly ordered dialog situation serving as an attention check. The CRS
responses for this situation were also manually curated, with one required to select a particular
rating from the given scale as a criterion to pass the attention check. After the rating task,
participants were shown demographic questions to reveal their user-related aspects.
Participants We recruited 107 participants from Amazon Mechanical Turk with qualifications
as fluent in English and having an interest in movies. After the experiment, we analyzed the data
and found that 9 participants failed the attention check, and 8 participants provided inconsistent
ratings. For example, in one of the cases, we found diferent ratings for similar responses by
both systems. Other such cases can also be found in [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. To avoid cherry-picking, we removed
the entire data of such unreliable subjects, leaving us with valid data for 90 participants. In this
way, we collected a total of 810 (9 per subject, after removing 1 attention-check trial) valid user
ratings for each CRS algorithm. On average, participants took 9 minutes to complete the task,
and they were paid 1.5 USD each. The study data is available online.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>In this section, we explain the procedure of curating data for each identified factor (variable)
and the methodology for analyzing the research questions.</p>
      <p>Dependent Variable. In our experiment, study participants rated their perceptions of
meaningfulness as a proxy to the overall quality of responses from two CRS for a total of 810 dialog
situations. So overall, we collected 1620 rating scores for both KBRD and CRB-CRS. The rating
score for perceived meaningfulness was defined as a dependent variable in our analyses.</p>
      <p>
        Independent Variables. We consider six diferent user-related aspects in our research,
acquired from users’ self-reported responses. Four of them, i.e., Gender, Age, Education, and English
proficiency describe the users’ demographic backgrounds. Additionally, inspired by works in
[
        <xref ref-type="bibr" rid="ref13 ref41">45, 13, 41, 46</xref>
        ], two features, Domain knowledge and Prior chatbot experience were included in
our analyses. To study system-related aspects, we consider three variables, Informativeness,
Response length, and Algorithmic performance diferences between KBRD [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and CRB-CRS
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Informativeness refers to the extent of domain items included in their system responses.
Specifically, we compute the item ratio, as was done in [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], as a proxy for informativeness in
the responses for each system. Moreover, Response length (i.e., token count of each response)
might be an important factor, also introduced in [
        <xref ref-type="bibr" rid="ref36 ref5">5, 36</xref>
        ].
      </p>
      <p>Regarding contextual aspects, we consider the Dialog discourse stage. To investigate this, we
split the dialog situations in our study into three groups – Initial, Middle, and End – based on
the number of seeker utterances. The dialogs in this study varied in length, ranging from one to
nine seeker utterances. On average, each dialog had 2.5 seeker utterances. Dialogs in the Initial
group consist of up to two seeker utterances, the Middle group had three to five utterances, and
longer dialogs were classified into the End group.</p>
      <p>Method. Given the collected data, the goal of this study is to examine the relationship
between various independent variables and the dependent variable (i.e., user rating scores
for the meaningfulness of CRS responses). Therefore, we used a Mixed Linear Efect Model
regression (MLEM) to analyze the data using Python’s statsmodels library1. Specifically, we
used MLEM because (i) individual observations in our experiment are not independent, (ii) the
ratings for perceived meaningfulness from the same subject are likely to be more similar to
each other given similar dialog situations, and (iii) the model is appropriate for within-subject
study designs where each participant is measured multiple times [47, 48]. Our model, therefore,
includes a fixed efects term for the CRS method (i.e., KBRD and CRB-CRS) and a random
intercept to account for any variability between diferent participants. We set  to 0.05 for the
statistical significance threshold. We apply this model twice aiming to address two research
questions.</p>
      <p>Specifically, at first, MLEM is applied to identify relationships between the dependent variable
and independent variables, i.e., addressing RQ1. Second, our goal was to investigate the
interaction efects between the dependent variable and various interaction terms of the independent
variables, targeting RQ2.</p>
      <p>Before applying the model, we employed the Shapiro-Wilk test [49] for normality checks
on the dialogs from which dialog situations were curated. The skewness turned out to be 0.01,
which indicates a slightly right-skewed, probably due to the large sample size (810 rating trials),
but an approximately symmetric distribution.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Analyses and Results</title>
      <sec id="sec-5-1">
        <title>5.1. Descriptive Statistics</title>
        <p>The descriptive statistics of the categorical and continuous variables in our study are shown
in Table 1 and Table 2, respectively. Overall, we have an almost balanced gender distribution
making our findings inclusive. All participants are fluent in English and have varied levels of
domain understanding, suggesting they are intellectually and linguistically suitable for our
study. Note however that before fitting the models for our further analyses: (i) we do not
consider language fluency variable as all subjects are fluent in English, and (ii) discarded entire
ratings from a subject stated “Prefer not to say” for its gender disclosure, thus leaving us with
1https://pypi.org/project/statsmodels/
801 rating trials for model fit. Based on the average scores, CRB-CRS with an average score of
3.78 (STD= 1.48) outperformed KBRD having a score of 3.62 ( STD=1.75). Also, a deeper analysis
of the response length shows in user requests asking for a movie plot or explanation, the system
response gets longer, thus reflecting high variance.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Efects of User, System and Context-related Factors (RQ1)</title>
        <p>The goal of this analysis is to reveal what kind of diferent factors related to either user, system, or
context influence human perceptions of meaningfulness. By following the above methodology,
we report the results of our Model 1 in Table 3, which indicates that users’ gender, education,
and domain knowledge have a significant influence on the outcome variable (i.e., perceived
meaningfulness). Among system-related aspects, results indicate that the informativeness of
system responses, i.e., the extent of mentioned domain entities in responses, with a high
coeficient score, is a leading predictor of users’ perceptions of meaningfulness (coef: 0.336, 
&lt; 0.001). Individual algorithmic diferences are also significant (  &lt; 0.01) indicating that users
perceived CRB-CRS responses as meaningful more often than KBRD, however, the mutual
diference is minimal (coef: 0.000). In addition, regarding a contextual feature, i.e., dialog
discourse stage is not significant (  &gt; 0.05) when keeping the remaining variables constant. On
the other hand, response length was negatively correlated to meaningfulness with a coeficient
of -0.014 ( &lt; 0.01), meaning that users tended to rate shorter responses as more meaningful.</p>
        <p>The mixed efect model’s Scale value of 1.54 implies that there is a certain extent of variability
in user ratings supporting the validity of our model in light of the mean ratings and standard
deviations of both KBRD and CRB-CRS. Additionally, our model converged on the data,
indicating its reliability and the ability to successfully capture the underlying relationships between
the independent and dependent variables, see also e.g., [50, 51].</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Interaction Efects (RQ2)</title>
        <p>
          Regarding RQ2, inspired by prior research [
          <xref ref-type="bibr" rid="ref18">18, 52, 53</xref>
          ], we examine how user-related factors
(e.g., age, gender, and chatbot experience) interact with system-related (response length and
informativeness) and context-related (dialog discourse stage) factors to shape users’ perceptions
and attitudes toward two CRS. In addition, to understand how two systems interact in a particular
dialog context, we explore the efects of system-related and contextual aspects as well.
        </p>
        <p>Table 4 presents the overall results of our MLEM model, which successfully converged on
the data ( with a Scale value of 1.52). Model 2 revealed significant interaction efects between
users’ previous chatbot experience and informativeness ( &lt; 0.001), as well as between users’
age and dialog discourse stage on perceived meaningfulness. We further explain the findings
for significant (highlighted in bold) interaction terms below.</p>
        <p>Interaction Efects between Chatbot Experience and Informativeness. Model 2
revealed a significant yet negatively correlated two-way interaction between users’ prior chatbot
experience and the extent of informativeness in responses from two CRS algorithms, impacting
perceived meaningfulness. As shown in Figure 2, specifically, the increase in informativeness
somehow linearly influences users’ perceptions of meaningfulness for those with prior chatbot
experience ( &lt; 0.001). However, for users without prior chatbot experience, the results did
not provide a definitive conclusion on such interaction efects, though it appears participants
perceived higher meaningfulness when one-three items were mentioned in the responses.</p>
      </sec>
      <sec id="sec-5-4">
        <title>Interaction Efects between Age and Dialog Discourse Stage. Figure 3 displays the</title>
        <p>interaction efects between age and dialog discourse stage on perceived meaningfulness. Shorter
dialogs have a positive influence on perceived meaningfulness across all age groups, indicating
the good performance of KBRD and CRB-CRS in generating responses for initial chitchat
or preference elicitation-based user utterances. Middle-length dialogs show mixed efects
based on age group, likely due to subjective assessments of the made item recommendations.
Interestingly, longer dialogs have a positive correlation for perceived meaningfulness with
younger participants, but a negative correlation with older participants (35 years onward). On
the contrary, younger adults (especially 18-25 years old) perceived the system response at the
end stage of dialog as quite meaningful. This indicates that old users expect concise but efective
responses to appropriately serve the user’s specific queries too, like asking for explanations
about made recommendations.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and Findings</title>
      <p>In this section, we discuss the implications and limitations of our research.</p>
      <p>Findings and Implications Generally, we observe substantial progress in developing various
techniques for conversational recommendation tasks. In this research, we attempt to understand
the relationships between users’ perceptions of system quality estimated through
meaningfulness and related factors. Specifically, from the literature, we investigated various factors
(user-related, system-related, and context-related), and their influence on human perceptions of
meaningfulness. Based on a series of analyses, we highlight the findings of our research and
their implications as follows.</p>
      <p>
        Designing CRS for diferent user segments. The interaction efect between chatbot
experience and informativeness suggests that users having prior experience with chatbots have higher
expectations. Therefore, when designing (CRS) chatbots, it is important to consider the target
user base and its familiarity with technology and provide more advanced features like access
to item metadata, reviews proliferation, or explanations for item recommendations, to meet
their expectations. This could also include incorporating more sophisticated natural language
processing capabilities, personalized recommendations, or advanced system functionalities.
Tailoring dialog length based on user age group. The relationship between the dialog
discourse stage and perceived meaningfulness being dependent on the user’s age indicates that
diferent age groups prefer varying lengths of dialogs. Older users tend to prefer shorter dialogs,
while younger users find longer dialogs more meaningful. To optimize user engagement and
satisfaction, CRS designers should consider adapting the length and state of conversations to
align with the preferences of diferent age segments. This may involve using concise responses
for older users and providing explanatory, and interactive conversations for younger users.
Understanding gender diferences in response expectations. Our finding, i.e., female
users tend to have a higher need for informative responses, aligns with previous research
highlighting their higher expectations and critical evaluation of system responses, see also
e.g., [54]. This suggests that CRS developers should pay attention to providing accurate and
relevant information to meet their expectations. Incorporating robust knowledge bases, ensuring
accurate information retrieval, and emphasizing clarity and completeness in responses can help
address the needs of various gender groups to enhance their satisfaction with the system.
Limitations One potential limitation of our study is that we only examined two CRS as
representatives of both language generation and retrieval methods. It remains an open question
to what extent our findings are generalizable. Also, both analyzed CRS are developed using the
same dataset from the movies domain, therefore limiting the generalizability of our findings to
other domains and to other datasets e.g., [
        <xref ref-type="bibr" rid="ref34">34, 55, 56</xref>
        ]. However, we believe our results can be
considered representative of the current state-of-the-art due to their superior performance over
other baselines, and both systems were published in highly-ranked scientific venues. Second,
the reliability of study participants is a potential threat to validity. However, we applied multiple
quality-assurance measures, including participant selection qualifications like interest in movies
and English fluency, attention checks, and manual inspection, which increases our confidence
in the reliability of our results.
      </p>
      <p>In addition, it is important to note that our study involved N=90 subjects, which might be
considered relatively limited for analyzing user aspects. However, we believe that these subjects
represent a subset of CRS users in practice. Thus, enabling CRS researchers to explore further
directions with more comprehensive studies. Overall, our research sheds light on the diverse
dimensions of user expectations and system functionalities supported by a substantial number
of 1602 users’ ratings. This understanding may further enable us to efectively design CRS that
enhances system quality and ultimately lead to improved user satisfaction. In future research,
we intend to explore additional factors like dialog initiative strategy and users’ personal traits in
a high-powered study, potentially influencing human perceptions of quality in conversational
recommender systems.</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>Conversational recommender systems (CRS) that interact with users in natural language assist
users in finding relevant items via multi-turn dialogs. To design interactive CRS that fulfills
individual needs and expectations, it is vital to understand what factors influence users’
perceptions of system quality. To this end, based on our user study specifically designed to obtain
users’ fine-grained feedback on various parts of the dialog, we investigate which user-related,
system-related, and contextual factors are the main predictors of human quality perceptions of
CRS. We further delved into investigating the combined efects of such factors to reveal how
various terms interact to afect the perceived quality of CRS. We believe our findings will be
helpful in tailoring various dimensions of user expectations and system functionalities, thus
opening new research directions to understand how we design and evaluate today’s CRS.
naire study, Quality and User Experience 5 (2020) 3.
[42] N. Sonboli, J. J. Smith, F. Cabral Berenfus, R. Burke, C. Fiesler, Fairness and transparency
in recommendation: The users’ perspective, in: UMAP ’21, 2021, pp. 274–279.
[43] D. Elsweiler, C. Trattner, M. Harvey, Exploiting food choice biases for healthier recipe
recommendation, in: SIGIR ’17, 2017, pp. 575–584.
[44] W. Lei, X. He, M. de Rijke, T.-S. Chua, Conversational recommendation: Formulation,
methods, and evaluation, in: SIGIR ’20, 2020, pp. 2425–2428.
[45] M. Nourani, J. King, E. Ragan, The role of domain expertise in user trust and the impact of
ifrst impressions with intelligent systems, in: AAAI ’20, volume 8, 2020, pp. 112–121.
[46] P. B. Brandtzaeg, A. Følstad, Chatbots: changing user needs and motivations, interactions
25 (2018) 38–43.
[47] H. Quan, W. J. Shih, Assessing reproducibility by the within-subject coeficient of variation
with random efects models, Biometrics (1996) 1195–1203.
[48] D. J. Barr, R. Levy, C. Scheepers, H. J. Tily, Random efects structure for confirmatory
hypothesis testing: Keep it maximal, Journal of memory and language 68 (2013) 255–278.
[49] S. Shapiro, M. Wilk, A goodness-of-fit test for normality: The shapiro-wilk test, Journal of
the American Statistical Association 60 (1965) pp. 619–626.
[50] A. Gelman, J. Hill, Data analysis using regression and multilevel/hierarchical models,</p>
      <p>Cambridge University Press, 2006.
[51] H. F. Senter, Applied linear statistical models. michael h. kutner, christopher j. nachtsheim,
john neter, and william li, Journal of the American Statistical Association 103 (2008)
880–880.
[52] B. P. Knijnenburg, M. C. Willemsen, Z. Gantner, H. Soncu, C. Newell, Explaining the user
experience of recommender systems, UMUAI 22 (2012) 441–504.
[53] C. M. Myers, A. Furqan, J. Zhu, The impact of user characteristics and preferences on
performance with an unfamiliar voice user interface, in: CHI ’19, 2019, pp. 1–9.
[54] P. K. Mo, S. H. Malik, N. S. Coulson, Gender diferences in computer-mediated
communication: A systematic literature review of online health-related support groups, Patient
Education and Counseling 75 (2009) 16–24.
[55] A. Manzoor, D. Jannach, Inspired2: An improved dataset for sociable conversational
recommendation (2022).
[56] Z. Liu, H. Wang, Z. Niu, H. Wu, W. Che, T. Liu, Towards conversational recommendation
over multi-type dialogs, in: ACL ’20, 2020, pp. 1036–1049.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Generation-based vs. retrieval-based conversational recommendation: A user-centric comparison</article-title>
          , in: RecSys '
          <fpage>21</fpage>
          ,
          <year>2021</year>
          , pp.
          <fpage>515</fpage>
          -
          <lpage>520</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Ai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Croft</surname>
          </string-name>
          ,
          <article-title>Towards conversational search and recommendation: System ask, user respond</article-title>
          ,
          <source>in: CIKM '18</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>177</fpage>
          -
          <lpage>186</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Evaluating conversational recommender systems</article-title>
          ,
          <source>Artificial Intelligence Review forthcoming</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <article-title>A user-centric evaluation framework for recommender systems</article-title>
          ,
          <source>in: Proceedings of the fifth ACM conference on Recommender systems</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Hayati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          , Inspired:
          <article-title>Toward sociable recommendation dialog systems</article-title>
          ,
          <source>in: EMNLP '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>8142</fpage>
          -
          <lpage>8152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Improving conversational recommender systems via knowledge graph based semantic fusion</article-title>
          ,
          <source>in: KDD '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1006</fpage>
          -
          <lpage>1014</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>Towards knowledge-based recommender dialog system</article-title>
          ,
          <source>in: EMNLP-IJCNLP '19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1803</fpage>
          -
          <lpage>1813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Conversational recommendation based on end-to-end learning: How far are we?</article-title>
          ,
          <source>Computers in Human Behavior Reports</source>
          <volume>4</volume>
          (
          <year>2021</year>
          )
          <fpage>100139</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.-W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lowe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. V.</given-names>
            <surname>Serban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Noseworthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Charlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pineau</surname>
          </string-name>
          ,
          <article-title>How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation</article-title>
          ,
          <source>in: EMNLP '16</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>2122</fpage>
          -
          <lpage>2132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Ebrahimi</given-names>
            <surname>Kahou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schulz</surname>
          </string-name>
          , V. Michalski, L. Charlin,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Towards deep conversational recommendations</article-title>
          ,
          <source>NIPS '18</source>
          <volume>31</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <article-title>End-to-end learning for conversational recommendation: A long way to go?</article-title>
          , in: IntRS@ RecSys,
          <year>2020</year>
          , pp.
          <fpage>72</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <surname>INFACT:</surname>
          </string-name>
          <article-title>An online human evaluation framework for conversational recommendation</article-title>
          , in: KARS@ RecSys,
          <year>2022</year>
          , pp.
          <fpage>72</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. J. M.</given-names>
            <surname>Reijmer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          ,
          <article-title>Each to his own: how diferent users call for diferent interaction methods in recommender systems</article-title>
          , in: RecSys '
          <fpage>11</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Berkovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Taib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Conway</surname>
          </string-name>
          ,
          <article-title>How to recommend? user trust factors in movie recommender systems</article-title>
          ,
          <source>in: IUI '17</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          , L. Chen,
          <article-title>Predicting user intents and satisfaction with dialogue-based conversational recommendations</article-title>
          ,
          <source>in: UMAP '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>I.</given-names>
            <surname>Benbasat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Trust in and adoption of online recommendation agents</article-title>
          ,
          <source>JAIS</source>
          <volume>6</volume>
          (
          <year>2005</year>
          )
          <article-title>4</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>Trust building in recommender agents</article-title>
          ,
          <source>in: WPRS-IUI '05</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          , L. Chen,
          <article-title>Impacts of personal characteristics on user trust in conversational recommender systems</article-title>
          ,
          <source>in: CHI '22</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Araujo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Helberger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kruikemeier</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. H. De Vreese</surname>
          </string-name>
          ,
          <article-title>In ai we trust? perceptions about automated decision-making by artificial intelligence</article-title>
          ,
          <source>AI &amp; Society</source>
          <volume>35</volume>
          (
          <year>2020</year>
          )
          <fpage>611</fpage>
          -
          <lpage>623</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Towards retrieval-based conversational recommendation</article-title>
          ,
          <source>Information Systems</source>
          <volume>109</volume>
          (
          <year>2022</year>
          )
          <fpage>102083</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <article-title>Key qualities of conversational recommender systems: From users' perspective</article-title>
          , in: HAI '
          <fpage>21</fpage>
          ,
          <year>2021</year>
          , pp.
          <fpage>93</fpage>
          -
          <lpage>102</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Bauer, Escaping the McNamara Fallacy: Towards more impactful recommender systems research</article-title>
          ,
          <source>AI</source>
          Magazine
          <volume>41</volume>
          (
          <year>2020</year>
          )
          <fpage>79</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Qu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Croft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>A hybrid retrieval-generation neural conversation model</article-title>
          ,
          <source>in: CIKM '19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1341</fpage>
          -
          <lpage>1350</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>NIPS '17</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Isele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jakob</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jentzsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kontokostas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hellmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Morsey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Van Kleef</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Auer</surname>
          </string-name>
          , et al.,
          <article-title>Dbpedia-a large-scale, multilingual knowledge base extracted from wikipedia</article-title>
          ,
          <source>Semantic Web</source>
          <volume>6</volume>
          (
          <year>2015</year>
          )
          <fpage>167</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sordoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Vahabi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lioma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Grue</given-names>
            <surname>Simonsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-Y.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <article-title>A hierarchical recurrent encoder-decoder for generative context-aware query suggestion</article-title>
          ,
          <source>in: CIKM '15</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>553</fpage>
          -
          <lpage>562</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Edunov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Grangier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Auli</surname>
          </string-name>
          ,
          <article-title>Scaling neural machine translation</article-title>
          ,
          <source>in: MT '18</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>Towards knowledge-based personalized product description generation in e-commerce</article-title>
          ,
          <source>in: SIGKDD '19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>3040</fpage>
          -
          <lpage>3050</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. W.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>HotpotQA: A dataset for diverse, explainable multi-hop question answering (</article-title>
          <year>2018</year>
          ). arXiv:
          <year>1809</year>
          .09600.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          , CRFR:
          <article-title>Improving conversational recommender systems via flexible fragments reasoning on knowledge graphs</article-title>
          ,
          <source>in: EMNLP</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>4324</fpage>
          -
          <lpage>4334</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ma</surname>
          </string-name>
          , S. Cui,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Revcore: Review-augmented conversational recommendation</article-title>
          ,
          <source>in: ACL-IJCNLP '21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1161</fpage>
          -
          <lpage>1173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Knowledge-based conversational recommender systems enhanced by dialogue policy learning</article-title>
          ,
          <source>in: ICKG '21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <article-title>Analyzing and simulating user utterance reformulation in conversational recommender systems</article-title>
          ,
          <source>in: SIGIR '22</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>133</fpage>
          -
          <lpage>143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>Towards topic-guided conversational recommender system</article-title>
          ,
          <source>in: ICCL '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>4128</fpage>
          -
          <lpage>4139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Y.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Towards enriching responses with crowd-sourced knowledge for task-oriented dialogue</article-title>
          ,
          <source>in: MuCAI '21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Liu,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Miao</surname>
          </string-name>
          , Kecrs:
          <article-title>Towards knowledgeenriched conversational recommendation system (</article-title>
          <year>2021</year>
          ). arXiv:
          <volume>2105</volume>
          .
          <fpage>08261</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Abdollahpouri</surname>
          </string-name>
          ,
          <article-title>A survey on multi-objective recommender systems</article-title>
          ,
          <source>Frontiers in Big Data</source>
          <volume>6</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>A.</given-names>
            <surname>Beheshti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yakhchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mousaeirad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Ghafari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Goluguri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Edrisi</surname>
          </string-name>
          ,
          <article-title>Towards cognitive recommender systems</article-title>
          ,
          <source>Algorithms</source>
          <volume>13</volume>
          (
          <year>2020</year>
          )
          <fpage>176</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>R.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Harper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>Factors influencing perceived fairness in algorithmic decision-making: Algorithm outcomes, development procedures, and individual diferences</article-title>
          ,
          <source>in: CHI '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>N.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Chen,</surname>
          </string-name>
          <article-title>User bias in beyond-accuracy measurement of recommendation algorithms</article-title>
          , in: RecSys '
          <fpage>21</fpage>
          ,
          <year>2021</year>
          , pp.
          <fpage>133</fpage>
          -
          <lpage>142</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>A.</given-names>
            <surname>Følstad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. B.</given-names>
            <surname>Brandtzaeg</surname>
          </string-name>
          ,
          <article-title>Users' experiences with chatbots: findings from a question-</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>