<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Taxonomy of User Feedback Intents for Conversational Recommendations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wanling Cai</string-name>
          <email>cswlcai@comp.hkbu.edu.hk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Li Chen</string-name>
          <email>lichen@comp.hkbu.edu.hk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Hong Kong Baptist University</institution>
          ,
          <addr-line>Hong Kong</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>Understanding users' feedback on recommendation in natural language is crucially important for assisting the system to refine its understanding of the user's preferences and provide more accurate recommendations in the subsequent interactions. In this paper, we report the results of an exploratory study on a human-human dialogue dataset centered around movie recommendations. In particular, we manually labeled a set of over 200 dialogues at the uterance level, and then conducted descriptive analysis on them from both seekers' and recommenders' perspectives. The results reveal not only seekers' feedback intents as well as the types of preferences they have expressed, but also the reactions of human recommenders that have finally led to successful recommendation. A taxonomy for feedback intents is established along with the results, which could be constructive for improving conversational recommender systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Related Work</title>
      <p>
        Current dialogue-based conversational
recommender systems (DCRS) have
mostly focused on question generation
and selection before giving the
recommendation. For instance, an active
learning and bandit learning based
conversational framework was proposed
in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which aims to adjust question
selection strategy in real time. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] trained
a deep policy network to decide when the
system should conduct facet preference
elicitation. However, litle work on DCRS
has explicitly studied the feedback issue
that occurs when the user is not satisfied
with the current recommendation. In
the broader area, critiquing-based
recommender systems [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] have been developed
to elicit users’ feedback in graphical
user interfaces (GUI), for which several
major types of critiquing are supported
such as user-initiated critiquing and
system-suggested critiques. However, this
kind of system limits the way users can
post their feedback since their interactions
are restricted to traditional GUI elements
(e.g., menu, form, buton). The advantage
of dialogue systems is that the interaction
is not limited to a pre-defined procedure
or a fixed set of atributes. But to the
best of our knowledge, few studies have
investigated users’ goals, intents, and
ways of expressing preferences when they
interact with DCRS [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], not to mention
their feedback on recommendations.
      </p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        Dialogue-based conversational recommender systems; user feedback; intent taxonomy
In recent years, dialogue systems have become increasingly popular in our daily life, with applications
in various domains such as education, healthcare, e-commerce, business, etc. They often mimic
human-like behavior to converse with users for addressing their chit-chating or information-seeking
requirements [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Given that users often explicitly request recommendations when they communicate
with a task-oriented dialogue system [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], more eforts have been put in integrating recommending
approaches into the system, so called the Dialogue-based Conversational Recommender System (DCRS)
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, most of existing systems have provided one-shot recommendation, with the focus on
selecting most informative questions to ask users [
        <xref ref-type="bibr" rid="ref3 ref8">3, 8</xref>
        ]. The dialogue often ends when the system
produces one or a list of recommendations to the user (see related work in the left bar). But in
reality, users may not get the desired recommendation within a single turn, in which case it becomes
important to allow users to freely provide their feedback on the recommendation, so that the system
could help them to find the desired item in the subsequent interactions. Our work is actually motivated
by the real dialogue that can occur between two persons [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. For example, if a seeker does not like
the recommended movies from the recommender, s/he can give feedback such as “I don’t like any of
those movies, too much talking” to refine her/his preferences. The user feedback issue has been studied
in a broader area of recommender systems, such as critiquing-based recommender systems that elicit
users’ feedback in graphical user interfaces [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], but litle work has been done on user feedback in
natural language. Since the language-based feedback can be in diverse, free styles, it is meaningful to
investigate how users express it (e.g., what intents they may have and what kinds of preferences they
want to convey), which should be constructive for developing more dedicated preference elicitation
and intent prediction strategies for DCRS. Therefore, we manually labeled a set of human-human
dialogues (over 200) centered around movie recommendations [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] with our established taxonomy for
user feedback intents, and analyzed the uterances starting from the point when a seeker did not like
one recommendation till s/he accepted another one. The results analysis reveals not only the seeker’s
intents and preferences, but also the human recommender’s responses that eventually helped the
seeker find a satisfactory item. It is hence inspiring for boosting the human-like aspect of current
dialogue systems.
      </p>
    </sec>
    <sec id="sec-3">
      <title>DIALOGUE-BASED RECOMMENDATION DATASET</title>
      <p>In this section, we present our data selection, taxonomy definition, and data annotation procedure.</p>
      <sec id="sec-3-1">
        <title>1htps://redialdata.github.io/website/</title>
      </sec>
      <sec id="sec-3-2">
        <title>2One conversation turn denotes a consecu</title>
        <p>tive uterance-response pair: Uterance is from
seeker and response is from recommender.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Data Annotation</title>
      <p>
        Two annotators were involved into the labeling
work. They were instructed to carefully read
the taxonomy table before they started. For
each uterance, the annotator was encouraged
to choose all labels that s/he thinks can
represent the seeker’s intents. They first
independently labeled 143 random dialogues. The
interrater agreement across their intent labels is 0.87
(through Fuzzy Kappa [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), which indicates
satisfactory annotation quality and consistency.
They then labeled the remaining dialogues, and
met to discuss and resolve disagreements.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Data Selection</title>
      <p>
        The original ReDail1 dataset contains 11,348 human-human dialogues [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We first filtered out
dialogues with less than 3 conversation turns2. We also removed those with inconsistent answers from
seekers and recommenders regarding the post-conversation reflective questions, because it may be
due to carelessness or dialogue ambiguity [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We then selected the dialogues containing at least
two movies suggested by the recommender, among which one was not liked by the seeker while
another subsequent recommendation was accepted by her/him. This process was mainly to capture
the seeker’s feedback on recommendation in case s/he was not satisfied with it, as well as how the
human recommender responded to the seeker and helped her/him to find a satisfactory item later. As
a result, we got 225 dialogues (see Table 1 with the statistics of our selected dialogue data).
      </p>
    </sec>
    <sec id="sec-6">
      <title>Taxonomy for User Feedback Intents</title>
      <p>
        Based on literature survey, we first established an initial taxonomy to classify user feedback on
recommendations, which basically covers all of the feedback types, such as the three types of feedback
modality (i.e., similarity-based, quality-based, quantity-based) in critiquing-based recommender
systems [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the session-aware intents (i.e., add filter condition, see-more, negation) in task-oriented
dialogue systems [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and the follow-up query strategies (i.e., refine, reformulate, and start over) when
users ask for recommendations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Then, we refined the taxonomy by applying the open coding
and theme identification approaches [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to our dialogue data. Four new categories (i.e., Inquire, Seen,
Provide Details, Ask) were added through keywords-in-context method; and some existing categories
were modified or merged into seven categories (i.e., Reject, Critique-Add, Critique-Compare,
CritiqueFeature, Restate, Restate with Further Constraints, Restate with Clarification) based on the real dialogues
through constant comparison method. We went through the standard classification procedure (i.e.,
propose-annotate-refine) three times, and finally came up with the taxonomy for user feedback intents
(see Table 2). The data annotation work is shown in the left bar.
      </p>
    </sec>
    <sec id="sec-7">
      <title>DATA ANALYSIS &amp; RESULTS</title>
    </sec>
    <sec id="sec-8">
      <title>Seeker Feedback Intents and Preference Expression</title>
      <p>
        Feedback Intent Distribution. The feedback intent distribution is shown in Table 2, where we can see
Reject, Seen, Critique-Feature, Provide Details, and Inquire more frequently occur than others, which
suggests that the seeker may tend to explicitly express her/his negative opinions on a
recommendation, and atempt to explain why s/he dislikes it as well as providing more preference info to the
recommender. Relatively, some seekers are also inclined to critique the recommendation by adding
further constraints, or start a new query if they feel it is dificult to receive a satisfactory result with
the current query.
Reject (REJ) Seeker dislikes the recommended item. “I hated that movie. I did not even crack a smile once.”
Seen (SEE) Seeker has seen the recommended item before. “I have seen that one and enjoyed it.”
Critique-Feature (CRI-F) Seeker makes critique on specific features of the current recommendation. “That’s a bit too scary for me.”
Provide Details (PRO) Seeker provides detailed preferences for the item s/he is looking for. “I usually enjoy movies with Seth Rogen and Jonah Hill.”
Inquire (INQ) Seeker wants to know more about the recommended item. “I haven’t seen that one yet. What’s it about?”
Critique-Add (CRI-A) Seeker adds further constraints on top of the current recommendation. “I would like something more recent.”
Start Over (STO) Seeker starts a new query. “Anything that I can watch with my kids under 10.”
Neutral Response (NRE) Seeker does not indicate her/his preferences for the current recommendation. “I have actually never seen that one.”
Critique-Compare (CRI-C) Seeker requests something similar to the current recommendation. “Den of Thieves (2018) sounds amazing. Any others like that?”
Answer (ANS) Seeker answers the question issued by the recommender. “Maybe something with more action.” (Q: “What kind of fun movie you look for?”)
Ask (ASK) Seeker asks the recommender’s personal opinions. “I really like Reese Witherspoon. How about you?”
Restate with Further Constraints (RES-CO) Seeker restates her/his query with further constraints. “Do you have something that is a thriller but not too scary?”
Restate (RES) Seeker completely restates her/his query. “Maybe I am not being clear. I want something that is in the theater now.”
Restate with Clarification (RES-CL) Seeker restates her/his query with clarification. “I’m fine with any sort of horrors, jump scares, clowns, etc.”
Others (OTH) The uterance cannot be categorized into any other categories. “ Sorry about the weird typing.”
navigational3. We refined this classification scheme by linking them to the concepts that the seeker
may mention [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]: Entity (like a movie or a series of movies that can be with subjective or navigational
goal), atribute
      </p>
      <p>(with objective or subjective goal), and purpose (the general uses of the item, e.g.,
“Anything that I can watch with my kids under 10? ”). The results show that seekers more frequently
express their preferences at the atribute level, which is much more often than the mentions of entity
and purpose concepts (see Figure 2). Moreover, they like to express subjective opinions on entity when
they mention it, but have more objective criteria for atributes (slightly higher than the proportion of
atribute-level subjective goals).</p>
    </sec>
    <sec id="sec-9">
      <title>Recommender Actions</title>
      <p>From the human recommender’s perspective, we investigated what actions s/he may carry out in
response to the seeker’s feedback. We first identified five major types of actions (see Table 3), and then
0.6
)
2
6
40.5
=
N
(
s
ce0.4
n
a
r
e
ttu0.3
f
o
e
ng0.2
a
t
n
e
rc0.1
e
P
0.0
25.54%
Recommend (REC) oRneecoomr mmeonrederercpormovmideensdations.</p>
      <sec id="sec-9-1">
        <title>Explain (EXP)</title>
      </sec>
      <sec id="sec-9-2">
        <title>Respond (RES)</title>
      </sec>
      <sec id="sec-9-3">
        <title>Answer (ANS)</title>
      </sec>
      <sec id="sec-9-4">
        <title>Request (REQ)</title>
        <p>Recommender explains
why the item is recommended.</p>
        <p>Recommender responds to
any other queries by the seeker.</p>
        <p>Recommender answers
the question from the seeker.</p>
        <p>Recommender requests for
the seeker’s preferences.
asked annotators to label all recommenders’ responses. From Table 3, we can see that, in nearly half
of the cases, the recommender tends to recommend one or more other items when the seeker rejects
the current one. In the other cases, the recommender tries to explain why the new recommendation
would be good to the seeker, respond to the seeker’s requests, answer the seeker’s explicit question,
or ask for the seeker’s preferences.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>FUTURE WORK</title>
      <p>In this work, we established a taxonomy for user feedback intents and analyzed a set of human-human
dialogues centered around movie recommendations. As the next step, we plan to label more dialogues
to further validate the taxonomy. We also want to perform temporal analysis so as to reveal the
frequent conversation paterns that may occur between seekers and recommenders. Based on the
findings from our analysis, we intend to develop a dedicated user intent prediction model to predict
users’ intents given their uterances, which is believed as an important component that could help
DCRS to track users’ current states, refine their preference model, and then select an approporiate
action to respond to users.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Joyce</given-names>
            <surname>Yue</surname>
          </string-name>
          <string-name>
            <surname>Chai</surname>
          </string-name>
          , Malgorzata Budzikowska, Veronika Horvath, Nicolas Nicolov, Nanda Kambhatla, and
          <string-name>
            <given-names>Wlodek</given-names>
            <surname>Zadrozny</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Natural Language Sales Assistant - A Web-Based Dialog System for Online Sales</article-title>
          .
          <source>In Proceedings of the Thirteenth Conference on Innovative Applications of Artificial Intelligence Conference</source>
          .
          <volume>19</volume>
          -
          <fpage>26</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Li</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pearl</given-names>
            <surname>Pu</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Critiquing-based Recommenders: Survey and Emerging Trends. User Modeling and User-</article-title>
          <source>Adapted Interaction 22</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>April 2012</year>
          ),
          <fpage>125</fpage>
          -
          <lpage>150</lpage>
          . htps://doi.org/10.1007/s11257-011-9108-6
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Konstantina</given-names>
            <surname>Christakopoulou</surname>
          </string-name>
          , Filip Radlinski, and
          <string-name>
            <given-names>Katja</given-names>
            <surname>Hofmann</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Towards Conversational Recommender Systems</article-title>
          .
          <source>In Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '16)</source>
          .
          <fpage>815</fpage>
          -
          <lpage>824</lpage>
          . htps://doi.org/10.1145/2939672.2939746
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jie</given-names>
            <surname>Kang</surname>
          </string-name>
          , Kyle Condif,
          <string-name>
            <given-names>Shuo</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Joseph A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          , Loren Terveen, and
          <string-name>
            <given-names>F. Maxwell</given-names>
            <surname>Harper</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Understanding How People Use Natural Language to Ask for Recommendations</article-title>
          .
          <source>In Proceedings of the Eleventh ACM Conference on Recommender Systems (RecSys '17)</source>
          .
          <fpage>229</fpage>
          -
          <lpage>237</lpage>
          . htps://doi.org/10.1145/3109859.3109873
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Andrei</surname>
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Kirilenko</surname>
            and
            <given-names>Svetlana</given-names>
          </string-name>
          <string-name>
            <surname>Stepchenkova</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Inter-Coder Agreement in One-to-Many Classification: Fuzzy Kappa</article-title>
          .
          <source>PLOS ONE 11</source>
          ,
          <issue>3</issue>
          (
          <issue>03</issue>
          <year>2016</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          . htps://doi.org/10.1371/journal.pone.0149787
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Raymond</given-names>
            <surname>Li</surname>
          </string-name>
          , Samira Ebrahimi Kahou, Hannes Schulz, Vincent Michalski, Laurent Charlin, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Pal</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Towards Deep Conversational Recommendations</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          <volume>32</volume>
          .
          <fpage>9748</fpage>
          -
          <lpage>9758</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Gery</surname>
            <given-names>W</given-names>
          </string-name>
          <string-name>
            <surname>Ryan and H Russell Bernard</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Techniques to Identify Themes</article-title>
          .
          <source>Field Methods 15</source>
          ,
          <issue>1</issue>
          (
          <year>2003</year>
          ),
          <fpage>85</fpage>
          -
          <lpage>109</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Yueming</given-names>
            <surname>Sun</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Conversational Recommender System</article-title>
          .
          <source>In The 41st International ACM SIGIR Conference on Research &amp; Development in Information Retrieval (SIGIR '18)</source>
          .
          <fpage>235</fpage>
          -
          <lpage>244</lpage>
          . htps://doi.org/10.1145/3209978.3210002
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Zhao</surname>
            <given-names>Yan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nan Duan</surname>
          </string-name>
          , Peng Chen, Ming Zhou,
          <string-name>
            <surname>Jianshe Zhou</surname>
            , and
            <given-names>Zhoujun</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Building Task-oriented Dialogue Systems for Online Shopping</article-title>
          .
          <source>In Thirty-First AAAI Conference on Artificial Intelligence.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>