<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dataset for Sociable Conversational Recom mendation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ahtsham Manzoor</string-name>
          <email>ahtsham.manzoor@aau.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dietmar Jannach</string-name>
          <email>dietmar.jannach@aau.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Klagenfurt</institution>
          ,
          <addr-line>Universitätsstraße 65-67, Klagenfurt am Wörthersee, 9020</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Conversational recommender systems (CRS) that are able to interact with users in natural language often utilize recommendation dialogs which were previously collected with the help of paired humans, where one plays the role of a seeker and the other as a recommender. These recommendation dialogs include items and entities that indicate the users' preferences. In order to precisely model the seekers' preferences and respond consistently, CRS typically rely on item and entity annotations. A recent example of such a dataset is INSPIRED, wich consists of recommendation dialogs for sociable conversational recommendation, where items and entities were annotated using automatic keyword or pattern matching techniques. An analysis of this dataset unfortunately revealed that there is a substantial number of cases where items and entities were either wrongly annotated or annotations were missing at all. This leads to the question to what extent automatic techniques for annotations are efective. Moreover, it is important to study impact of annotation quality on the overall efectiveness of a CRS in terms of the quality of the system's responses. To study these aspects, we manually fixed the annotations in performance of several benchmark CRS using both versions of the dataset. Our analyses suggest that the improved version of the dataset, i.e., INSPIRED2, helped increase the performance of several benchmark CRS, emphasizing the importance of data quality both for end-to-end learning and retrieval-based approaches to conversational recommendation. We release our improved dataset (INSPIRED2) publicly at https://github.com/ahtsham58/INSPIRED2.</p>
      </abstract>
      <kwd-group>
        <kwd>Conversational Recommender Systems</kwd>
        <kwd>data quality</kwd>
        <kwd>annotations</kwd>
        <kwd>evaluation</kwd>
        <kwd>dialog systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Sociable conversational recommender systems (CRS) aim
to build rapport with users while interacting with them in
natural language [1, 2]. CRS that rely on natural language
processing (NLP) nowadays commonly utilize datasets of
previously recorded dialogs between humans, where one
plays the role of a recommendation-seeker and the other
as human-recommender, see e.g., [
        <xref ref-type="bibr" rid="ref34">3</xref>
        ]. However, due to a
certain lack of rich sociable interactions in such datasets
[4], it can be challenging to build a sociable CRS that
builds rapport with the users using such limited data.
      </p>
      <sec id="sec-1-1">
        <title>SPIRED [1], which includes dialogs that implement rich social communication strategies. Such rich datasets represent a solid basis to develop trustable CRS that are able</title>
      </sec>
      <sec id="sec-1-2">
        <title>Another key factor for building high-quality CRS lies in the proper recognition of the named entities and other</title>
        <p>WA, USA.
∗Corresponding author.
nEvelop-O
LGOBE</p>
        <p>0000-0001-9418-753 (A. Manzoor); 0000-0002-4698-8507
(D. Jannach)
generation-based CRS approaches as well as for
retrievalbased approaches to build natural language
conversational systems [17]. For both types of systems, the
question arises to what extent better data quality, i.e., having
correct annotations and noise-free conversations, leads
to better results in terms of the quality of the responses
terms of consistency and plausibility.</p>
        <p>
          Therefore, it is important to develop datasets like IN- economically expensive process [
          <xref ref-type="bibr" rid="ref1">14, 15</xref>
          ]. Human costs
to engage users in a natural and user-adaptive manner. ity of the resulting annotations is crucial, and factually
4th Edition of Knowledge-aware and Conversational Recommender Sys- therefore been in the focus of research for several years.
        </p>
        <p>
          In this work, we study the recent INSPIRED dataset, A number of new datasets for conversational
recomin which the items and entities that were mentioned in mendation were published in recent years, e.g., [3, 21,
the recorded utterances are explicitly annotated. These 22, 23]. Such datasets, which are commonly collected
annotations were created with the help of automatic ap- with the help of crowdworkers, can however have
limiproaches using keyword or pattern matching methods. tations and may be not fully representative in terms of
However, looking at the data, we observed a substantial what we would observe in reality. In some cases, for
number of cases where items and entities were either example, crowdworkers were instructed to mention a
wrongly annotated or missing annotations at all, e.g., minimum number of movies in the conversations. This
“My favorite [ MOVIE_GENRE_1] are Groundhogs Day, leads to mostly “instance-based” conversations, where
[ MOVIE_TITLE_2] and Borat”. In addition, there were crowdworkers rather mention individual movies they
several cases where the utterances included noise, e.g., like than their preferred genres, see also, [
          <xref ref-type="bibr" rid="ref34">3, 24, 25</xref>
          ].
“How did you like QUOTATION_MARKHustlersQUOTA- Another problem when creating such datasets lies in
TION _MARK?”. Finally, we found instances where regu- the recognition and annotation of named entities
appearlar words were identified as being named entities. In this ing in the conversations, as mentioned above.
Annotatlatter case, human annotations would in fact have been ing entities in textual data can be a tedious process that
required.1 Overall, such issues may limit the quality of may require a substantial amount of manual efort and
any CRS that is built on top of such data. time. To overcome this challenge, researchers sometimes
        </p>
        <p>
          To understand the severity of the problem and the po- adapt a semi-automatic approach or rely on NLP-assisted
tential efects of data issues on the quality of a CRS, we tools that visualize the entities in a text in order to
rehave manually corrected the dataset by fixing the annota- duce the required manual efort [
          <xref ref-type="bibr" rid="ref1">16, 15, 26</xref>
          ]. Generally,
tions and by removing noise from the utterances. Then, some automatic approaches may experience problems to
we conducted ofline experiments and human evalua- correctly create annotations because human judgments
tions to compare the performance of diferent benchmark and opinions are required. An automated approach was
CRS when using the original (INSPIRED) and improved used in the context of the INSPIRED dataset. Here, the
(INSPIRED2) datasets. Overall, the results of our anal- items and entities were annotated using keyword or
patyses indicate that all CRS showed better performance tern matching approaches. However, verifying the
outin diferent dimensions when built on INSPIRED2. In comes of such automatic or semi-automatic approaches
order to facilitate the design and development of future can again be laborious and require manual efort.
sociable CRS, we release the INSPIRED2 dataset online Today, structured annotations for items and entities
at https://github.com/ahtsham58/INSPIRED2. mentioned in the conversations are common in recent
datasets. For example, in the case of the ReDial dataset
[
          <xref ref-type="bibr" rid="ref34">3</xref>
          ], the mentioned movie titles were annotated with
2. Related Work unique IDs. However, the ReDial dataset has some
limitations. Various meta-data concepts (e.g., genres, actors,
In this section, we first discuss datasets and aspects of or directors) were not annotated. Moreover, the recorded
data quality in the context of CRS. Afterwards, we review dialogs include limited social interactions or explanations
diferent design paradigms for building CRS, followed by for the made recommendations. On the other hand, the
a discussion of predominant evaluation approaches for INSPIRED dataset includes rich sociable conversation
such systems. and explanation strategies for the recommended items.
Also, aspects like movie genres or actors were explicitly
Datasets and Data Quality Research interest in CRS annotated too. A comparison of these diferences can be
has experienced a substantial growth in recent years, found in [1]. The key statistics of the INSPIRED dataset
see [
          <xref ref-type="bibr" rid="ref19">18, 19</xref>
          ] for related surveys. Many current systems are shown in Table 1.
interact with users in natural language, and one impor- As mentioned earlier, the INSPIRED dataset has some
tant goal for such system is to enable them to engage in limitations. The keyword or pattern matching approach
conversations that reflect human behavior. Since many used for the annotations might for example not detect
of these recent systems are built on recorded dialogs misspelled keywords or concepts in an utterance.
Morebetween humans, the capabilities of the resulting CRS over, data anomalies such as noisy utterances or
illdepend on the richness of the communication in the formed language can deteriorate the performance of an
datasets, e.g., in terms of the user intents that can be annotating algorithm, leading to challenges for the
downfound in the conversations, see [
          <xref ref-type="bibr" rid="ref31">20</xref>
          ] for an detailed anal- stream use of the dataset [
          <xref ref-type="bibr" rid="ref1">15, 27, 28</xref>
          ]. In reality, the level
ysis of such intents. of noise can be substantial both in real-world
applications and in purposefully created datasets. Therefore,
1Consider the movie “It (2017)” as an example of a dificult case, e.g., data quality assurance is often considered a significant
when appearing in an utterance like “Have you seen It?”. and important step in NLP applications.
Table 1 CRS Evaluation Evaluating a CRS is a multi-faceted
Main Statistics of INSPIRED and challenging problem as it requires the consideration
Total of various quality dimensions. An in-depth discussion
        </p>
        <p>of evaluation approaches for CRS can be found in [37].</p>
        <p>
          Number of dialogs (conversations) 1,001 Like in the recommender systems literature in general,
AAvveerraaggee ttoukrnesnspeprerdiuatltoegrance 170.9.733 computational experiments that do not involve humans
in the loop are the predominant instrument to assess
Number of human-recommender utterances 18,339 the quality of a CRS. Common metrics to evaluate the
Number of seeker utterances 17,472 quality of the recommendations include Recall, Hit Rate,
or Precision [
          <xref ref-type="bibr" rid="ref16 ref24 ref5">8, 21, 38</xref>
          ]. Moreover, certain linguistic
aspects such as fluency or diversity are often evaluated
Building Conversational Recommender Systems with ofline experiments as well to assess the quality of
Research on CRS has made substantial progress in terms the generated responses. Common metrics in this area
of their underlying technical approaches. Some early include Perplexity, distinct N-Gram, or the BLEU score
commercial system such as Advisor Suite [29] for example [
          <xref ref-type="bibr" rid="ref34 ref36 ref5">3, 8, 13, 22</xref>
          ].
relied on an entirely knowledge-based approach for the Given the interactive nature of CRS, ofline
experidevelopment of adaptive and personalized applications. ments and the corresponding metrics have their
limiSimilarly, early critiquing-based systems were based on tations. Mainly, it is not always clear if the results
obdetailed knowledge about item features and possible cri- tained from ofline experiments are representative of the
tiques and had limited learning capabilities [30, 31]. user-perceived quality of the recommendations or
sys
        </p>
        <p>
          Technological advancements particularly in fields like tem responses in general [35]. For example, when using
NLP, speech recognition, and machine learning in general metrics like the BLEU score, usually a system response is
led to the design of today’s end-to-end learning-based compared with one particular given ground truth. Such
CRS. In such approaches, recorded recommendation di- a comparison has limitations when used to estimate the
alogs between paired humans are used to train the deep average quality of a system’s responses, because there
neural models, see, e.g., [
          <xref ref-type="bibr" rid="ref5">8, 9, 10, 12</xref>
          ]. Given the last user might be many diferent alternative responses that might
utterance and the history of the ongoigng dialog history, be suitable as well in an ongoing dialog. Still, ofline
these trained models are then used to generate responses evaluations have their place and value. They can for
in natural language. These responses can either include example be informative for assessing particular aspects
item recommendations, which are also computed with such as the number of items or entities that appear in an
the help of machine learning techniques, or other types utterance or conversation.
of conversational elements, e.g., greetings. Overall, given the limitations of pure ofline
experi
        </p>
        <p>
          In terms of the underlying data, the DeepCRS [
          <xref ref-type="bibr" rid="ref34">3</xref>
          ] sys- ments, researchers often follow a mixed approach where
tem was built on the ReDial dataset, which was created some aspects of the system are evaluated ofline and some
in the context of this work. Later on, systems were devel- with humans. Typical quality aspects in terms of human
oped which also relied on this dataset as well but included perceptions in such combined approaches include the
additional information sources, e.g., from DBPedia or Con- assessment of the meaningfulness or consistency of the
ceptNet [32, 12], to build knowledge graphs that are then system responses [
          <xref ref-type="bibr" rid="ref36 ref5">1, 8, 12, 13, 36</xref>
          ].
used to improve the generated utterances. A number
of works also makes use of pretrained language models
like BERT [33] and subsequently fine-tune them using 3. Data Annotation Methodology
the recommendation dialogs, see, e.g., [34]. A related
approach was adapted by the authors of INSPIRED, in During the creation of the INSPIRED [1] dataset, items
which they proposed two variants of a conversational and other entities were annotated in an automated way,
system, with and without strategy labels. as described above. For example, genre keywords were
        </p>
        <p>Unlike generation-based systems, in retrieval-based annotated using a regular expression to match a set of
CRS the idea is to retrieve and adapt suitable responses predefined tokens. Regarding actors and directors and
from the dataset of recorded dialogs. One main advan- other entities, a pattern matching technique was used,
tage of retrieval-based approaches is that the retrieved where words starting with a capital letter were searched
responses were genuinely made by humans and thus in the TMDB database2. A similar technique was used for
are grammatically usually correct and in themselves se- movie titles. However, as mentioned, we observe a large
mantically meaningful [35]. Recent examples of such number of cases where items and entities were either
retrieval-based systems are RB-CRS [17] and CRB-CRS wrongly annotated or missing annotations. To answer
[36], which we designed and evaluated based on the Re- our research question on the impact of the quality of the
Dial dataset in our own previous work.
underlying data on the quality of the responses of a CRS, Observed Issues During the annotation process, we
we fixed the annotations as follows. recorded the observed issues in the original annotations.
Since the original annotations were created using
autoProcedure To fix the annotations, we interviewed a matic techniques, many issues were related to the
liminumber of university students to assess their knowledge tations of the simple keyword or pattern matching
techin the movies domain and their ability to do the correction niques. Overall, we observed a number of cases where
task. Subsequently, we hired two students and instructed minor spelling mistakes or incomplete movie titles made
them on how to annotate and clean the dataset. First, the exact string matching approaches inefective.
they were briefed on the logical format of the original For example, in one of the utterances, “I think I am
annotations and how to retain that format. Second, they waiting for Star Wars The Rise of Skywalker”, the
annotawere asked to read each utterance individually, to detect tion was missing because the correct title is “Star Wars:
potential noise, and to analyze which items or entities Episode IX – The Rise of Skywalker ”. Similarly, we
ob(e.g., title, genre, actor, or director) are mentioned in it. serve a significant number of cases where an utterance</p>
        <p>In case of ambiguity or obscurity, they were allowed was only partially annotated, e.g., “ok is it scary like
into access online portals, e.g., IMDb3. Note that regarding cidious or [ MOVIE_GENRE_2] [ MOVIE_TITLE_5]”. In
the genres, a set of 27 keywords was provided to them, addition, at places where two entities were separated
which we curated and used in our earlier research [36]. with ‘/’ instead of a space, the automatic technique often
After the briefing, the dataset was split evenly for both failed to create proper annotations, e.g., “Since you like
annotators. On weekly basis, their performance and the [MOVIE_GENRE_1] drama/mystery, I’m going to send you
accuracy of the annotations was checked by one of the the trailer to the movie [MOVIE_TITLE_3]”.
authors. Finally, after annotating the complete dataset, a Also, the automatic approach used for INSPIRED
somenumber of additional validation steps were applied. times had dificulties to deal with ambiguity. We found</p>
        <p>First, using a Python script, we ensured that every a number of cases where a regular word was annotated,
placeholder is enclosed by ‘[’ and ‘]’ as was done origi- although such a word did not belong to any item or
ennally, e.g., [MOVIE_TITLE_1]. Second, another thorough tity. For example, in one of the cases, “Are you interested
manual examination of the entire improved dataset was in a current movie in the box ofice? ”, the utterance was
performed to fix any missing annotations or noise. In annotated as “Are you interested in a current movie in
that context, we also double-checked the consistency of the box [ MOVIE_TITLE_0]”, where the word ‘ofice ’ was
the format and of the annotations. mistakenly annotated as an item, i.e., The Ofice (2005) .
Overall, the main observed issues are the following.</p>
        <p>The INSPIRED2 Dataset In total, 1,851 new
annotations were added to INSPIRED, leading to the INSPIRED2
dataset. The most mistakes or inconsistencies were found
for the items, i.e., movie titles, which is the most pertinent
information for developing a CRS. We present the
statistics about new annotations in Table 2. Overall, we added
around 20% new annotations in INSPIRED2. The number
of issues that were fixed, e.g., duplicate annotations in
an utterance, noise or factually wrong information in
the original annotations, are not shown in the presented
statistics. We release the INSPIRED2 both in the TSV and
JSON format online.</p>
        <p>Total</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Evaluation Methodology</title>
      <sec id="sec-2-1">
        <title>We performed both ofline experiments as well as a human evaluation to assess the impact of data quality on the quality of the responses of a CRS.</title>
        <sec id="sec-2-1-1">
          <title>Ofline Evaluation of Recommendation Quality</title>
          <p>
            We included the following recent end-to-end learning
approaches in our experiments: DeepCRS [
            <xref ref-type="bibr" rid="ref34">3</xref>
            ], KGSF [12],
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>TG-ReDial [22], and the INSPIRED model without strat</title>
        <p>egy labels4 [1]. This selection of models covers
various design approaches for CRS, e.g., using an additional
knowledge graph or not. We used the open-source toolkit
CRSLab5 for our evaluations. This framework was used
in earlier research as well, for example in [10, 39, 40].
For our analyses, we first trained the aforementioned
CRS models using the original split ratio, i.e., 8:1:1, for
each dataset. Afterwards, given the trained models and
test data for each dataset, we ran three trials for each
CRS and subsequently averaged the results for ofline
evaluation metrics. Note that the same procedure was
adapted for both versions of the dataset, i.e., INSPIRED
and INSPIRED2.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Results</title>
      <p>Recommendation Quality Table 3 shows the
accuracy results for the evaluated CRS models. Specifically,
we provide the results for the diferent benchmark CRS
models in terms of the performance diference when
using the original and improved annotations. Overall, we
can observe an almost consistent gain in performance for
all models and on all metrics except Hit@50 when the
improved dataset is used. The obtained improvements can
be quite substantial, indicating that improved data
quality can be helpful for CRS of diferent types, including (i)
CRS, which do not rely on additional knowledge sources,
(ii) CRS that leverage additional knowledge sources, (iii)
CRS that are guided by a topic policy, and (iv) CRS that
rely on pre-trained language models like BERT.</p>
      <p>Interestingly, we see negative efects for two
measurements in which Hit@50 is used as a metric. A deeper
investigation of this phenomenon is needed, in particular
as the other metrics at this (admittedly rather uncommon)
list length, MRR@50 and NDCG@50, indicate that the
improved dataset is helpful to increase recommendation
accuracy. At the moment, we can only speculate that
the improved annotations in the ongoing dialog histories
led to more diverse or niche recommendations compared
to the original dataset. We might assume that the
missing annotations in many cases referred to less popular
movies, so that the recommendations without the
improved annotations will more often recommend popular
movies, which is commonly advantageous in terms of hit
rate and recall.</p>
      <sec id="sec-3-1">
        <title>User Study on Linguistic Quality We conduct a user</title>
        <p>study to compare the perceived quality of system
responses using either INSPIRED and INSPIRED2.
Specifically, we randomly sampled same 50 dialog situations
from each dataset. To create the dialog continuations,
we used the retrieval-based CRS approaches, RB-CRS and
CRB-CRS, which we proposed in our earlier work, see
[36].</p>
        <p>
          In order to obtain fine-grained assessments, three
human judges6 were involved. The specific task of the
judges was to assess (rate) the meaningfulness of a
system response as a proxy of its quality and consistency in a
dialog situation, see [
          <xref ref-type="bibr" rid="ref34">3, 41, 12</xref>
          ]. Note that in this study we
did not explicitly assess the quality of the specific item
recommendations. Instead, the focus of this study was
to understand the impact of the improved underlying
dataset on the linguistic quality and the consistency of
the generated responses. Linguistic Quality We recall that three human
eval
        </p>
        <p>We used a 3-point scale for these ratings, from ‘Com- uators assessed the linguistic quality of the system
repletely meaningless (1)’ to ‘Somewhat meaningless and sponses (dialog continuations), which were created either
meaningful (2)’ to ‘Completely meaningful (3)’. The hu- based on the INSPIRED or the INSPIRED2 dataset. As
unman judges were provided with specific instructions on derlying CRS systems, we considered the retrieval-based
how to evaluate the meaningfulness of a response, e.g., approaches RB-CRS and CRB-CRS, as mentioned above.
they should assess if a response represents a logical dialog For our analysis, we averaged the scores by the three
continuation and evaluate the overall language quality evaluators. Table 4 shows the mean ratings across all
of the given response. Overall, the human judges were dialog situations as well as the standard deviations. We
provided 50 dialogs (446 responses to rate) that were pro- find that also in the case of retrieval-based approaches,
duced using the INSPIRED and INSPIRED2 datasets. We improving the quality of the underlying dataset was
helpalso explained the meanings and purpose of various place- ful, leading to higher mean scores, without observing
holders contained in the responses to the human judges. larger standard deviations. A Student’s t-test reveals that
Moreover, to avoid any bias in the evaluation process, the the observed diferences in the means are statistically
judges were not made aware which response was created significant ( p&lt;0.001). 7
for which dataset by which CRS. Also, the order of the
dialogs and the system responses were randomized.</p>
        <sec id="sec-3-1-1">
          <title>4The INSPIRED with strategy labels model was not publicly available.</title>
          <p>5https://github.com/RUCAIBox/CRSLab
6These judges were PhD students and were diferent than the ones
who fixed the annotations.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Comparison of Knowledge Concepts in Responses</title>
        <p>To understand the impact of the new annotations on the
responses in terms of the richness of knowledge
concepts, we compute the number of items and entities that</p>
        <sec id="sec-3-2-1">
          <title>7We provide the data and compiled results of our study online.</title>
          <p>
            appeared in the system responses. Specifically, we
compute the number of placeholders in the responses, before
they would be replaced by the recommendation
component, see [36]. In Table 5, we present the statistics for
RB-CRS and CRB-CRS for both dataset versions.
Overall, we find that the responses for the improved dataset
contain between 20% and 27% more concepts and
entities. We note that an increase in concepts is expected, as
INSPIRED2 has almost 20% more annotations. However,
the important observation here is that the retrieval-based
CRS approaches actually surfaced these richer system
responses frequently.
BLEU Score Analysis Finally, in order to understand
to what extent (ofline) linguistic scores correlate with
the perceived quality of responses as was done in [
            <xref ref-type="bibr" rid="ref5">1, 8</xref>
            ],
we performed an analysis of the BLEU scores obtained
for the diferent datasets. Specifically, given a system
response and the corresponding ground truth response,
we preprocess both sentences and compute the BLEU
scores for  = {1, 2, 3, 4} grams. We provide the results
of this analysis online. In sum, the analysis shows that
the BLEU scores generally improve when the underlying
data quality is higher, i.e., in the case of the INSPIRED2
dataset. These findings are thus well aligned with the
outcomes of our human evaluation study, where using
INSPIRED2 as an underlying dataset turned out to be
favorable.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6. Conclusion</title>
      <sec id="sec-4-1">
        <title>Datasets containing recorded dialogs between humans</title>
        <p>are the basis for many modern CRS. In this work, we
have analyzed the recent INSPIRED dataset, which was
developed to build the next generation of sociable CRS.</p>
        <p>We found that automatic entity and concept labeling has
its limitations and we have improved the quality of the
dataset through a manual process. We then conducted
both computational experiments as well as experiments
with users to analyze to what extent improved data
quality impacts recommendation accuracy and the quality
perception of the system’s responses by users. The
analyses clearly indicate the benefits of improved data quality
across diferent technical approaches for building CRS.</p>
        <p>We release the improved dataset publicly and hope to
thereby stimulate more research in sociable
conversational recommender systems in the future.
prove conversational recommendation, in: ICCBR conversational recommendation, Information
Sys’10, 2010, pp. 480–494. tems (2022) 102083.
[31] L. Chen, P. Pu, Critiquing-based recommenders: [37] D. Jannach, Evaluating conversational
recomsurvey and emerging trends, User Modeling and mender systems, Artificial Intelligence Review
User-Adapted Interaction 22 (2012) 125–150. forthcoming (2022).
[32] Q. Chen, J. Lin, Y. Zhang, H. Yang, J. Zhou, J. Tang, [38] T. Zhang, Y. Liu, P. Zhong, C. Zhang, H. Wang,
Towards knowledge-based personalized product de- C. Miao, KECRS: Towards knowledge-enriched
scription generation in e-commerce, in: KDD ’19, conversational recommendation system, 2021.
2019, pp. 3040–3050. a r X i v : 2 1 0 5 . 0 8 2 6 1 .
[33] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: [39] Y. Zhou, K. Zhou, W. X. Zhao, C. Wang, P. Jiang,
Pre-training of deep bidirectional transformers for H. Hu, C2-crs: Coarse-to-fine contrastive
learnlanguage understanding, in: NAACL-HLT, 2019. ing for conversational recommender system, in:
[34] L. Wang, H. Hu, L. Sha, C. Xu, K.-F. Wong, D. Jiang, WSDM ’22, 2022, pp. 1488–1496.</p>
        <p>Finetuning large-scale pre-trained language models [40] Y. Li, B. Peng, Y. Shen, Y. Mao, L. Liden, Z. Yu,
for conversational recommendation with knowl- J. Gao, Knowledge-grounded dialogue generation
edge graph, 2021. a r X i v : 2 1 1 0 . 0 7 4 7 7 . with a unified knowledge representation, 2021.
[35] A. Manzoor, D. Jannach, Conversational recom- a r X i v : 2 1 1 2 . 0 7 9 2 4 .</p>
        <p>mendation based on end-to-end learning: How far [41] D. Jannach, A. Manzoor, End-to-end learning for
are we?, Computers in Human Behavior Reports conversational recommendation: A long way to
(2021) 100139. go?, in: IntRS Workshop at RecSys ’20, Online,
[36] A. Manzoor, D. Jannach, Towards retrieval-based 2020.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Grosman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Furtado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Rodrigues</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. G.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Schardong</surname>
            ,
            <given-names>S. D.</given-names>
          </string-name>
          <string-name>
            <surname>Barbosa</surname>
            ,
            <given-names>H. C.</given-names>
          </string-name>
          <string-name>
            <surname>Lopes</surname>
            , Eras: Improv[1]
            <given-names>S. A.</given-names>
          </string-name>
          <string-name>
            <surname>Hayati</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <article-title>Yu, IN- ing the quality control in the annotation process</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          systems, in: EMNLP '
          <fpage>20</fpage>
          ,
          <year>2020</year>
          . Systems 93 (
          <year>2020</year>
          )
          <fpage>101553</fpage>
          . [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pecune</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Callebert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Marsella</surname>
          </string-name>
          , A socially-aware [16]
          <string-name>
            <given-names>B. C.</given-names>
            <surname>Benato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Gomes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Telea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. X.</given-names>
            <surname>Falcão</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>ized recipe recommendations</article-title>
          ,
          <source>in: Proceedings of space projection, Pattern Recognition</source>
          <volume>109</volume>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>the 8th International Conference on Human-Agent</source>
          <volume>107612</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Interaction</surname>
          </string-name>
          , HAI '
          <volume>20</volume>
          ,
          <year>2020</year>
          , p.
          <fpage>78</fpage>
          -
          <lpage>86</lpage>
          . [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Generation-based vs</article-title>
          . [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Kahou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schulz</surname>
          </string-name>
          , V. Michalski, L. Charlin,
          <article-title>retrieval-based conversational recommendation: A</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <article-title>Towards deep conversational recommenda- user-centric comparison</article-title>
          , in: RecSys '
          <fpage>21</fpage>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          tions, in: NIPS '
          <fpage>18</fpage>
          ,
          <year>2018</year>
          , pp.
          <fpage>9725</fpage>
          -
          <lpage>9735</lpage>
          . [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Manzoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          , A survey [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          , L. Chen, Conversational Recommen- on
          <source>conversational recommender systems</source>
          , ACM
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>dation: A Grand AI Challenge</article-title>
          ,
          <source>AI Magazine 43 Computing Surveys</source>
          <volume>54</volume>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          (
          <year>2022</year>
          ). [19]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. de Rijke</surname>
            , T.-S. Chua, Ad[5]
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Di Bratto</surname>
            ,
            <given-names>M. Di</given-names>
          </string-name>
          <string-name>
            <surname>Maro</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Origlia</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>Cutugno, vances and challenges in conversational recom-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>Dialogue analysis with graph databases: Character- mender systems: A survey</article-title>
          ,
          <source>AI</source>
          Open 2
          <article-title>(</article-title>
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>ising domain items usage for movie recommenda- 100-126</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>tions</surname>
          </string-name>
          (
          <year>2021</year>
          ). [20]
          <string-name>
            <given-names>W.</given-names>
            <surname>Cai</surname>
          </string-name>
          , L. Chen,
          <article-title>Predicting user intents</article-title>
          and satis[6]
          <string-name>
            <surname>C.-M. Wong</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , C.-M.
          <article-title>Vong, faction with dialogue-based conversational recom-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , P. He,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          , mendations, in: UMAP '
          <fpage>20</fpage>
          ,
          <year>2020</year>
          , p.
          <fpage>33</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>Improving conversational recommender system</article-title>
          [21]
          <string-name>
            <given-names>D.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Crook</surname>
          </string-name>
          , Y.-L.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>ICDE '21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>2607</fpage>
          -
          <lpage>2612</lpage>
          . nication game:
          <article-title>Self-supervised bot-play for goal[7</article-title>
          ]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Unifying oriented dialogue</article-title>
          ,
          <source>in: EMNLP-IJCNLP '19</source>
          ,
          <year>2019</year>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <article-title>knowledge graph learning</article-title>
          and recommendation:
          <fpage>1951</fpage>
          -
          <lpage>1961</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <article-title>Towards a better understanding of user preferences</article-title>
          , [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>in: WWW '19</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>151</fpage>
          -
          <lpage>161</lpage>
          .
          <article-title>Towards topic-guided conversational recommender [8</article-title>
          ]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          , system,
          <source>in: ICCL '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>4128</fpage>
          -
          <lpage>4139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>Towards knowledge-based recommender</article-title>
          [23]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , G. de Melo,
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>dialog system</article-title>
          ,
          <source>in: EMNLP-IJCNLP '19</source>
          ,
          <year>2019</year>
          , pp.
          <article-title>COOKIE: A dataset for conversational recommen-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          1803-
          <fpage>1813</fpage>
          .
          <article-title>dation over knowledge graphs in e-commerce</article-title>
          ,
          <year>2020</year>
          . [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          , CRFR: Improv- arXiv:
          <year>2008</year>
          .09237.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <article-title>ing conversational recommender systems via flexi-</article-title>
          [24]
          <string-name>
            <given-names>K.</given-names>
            <surname>Christakopoulou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Radlinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          , To-
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>EMNLP '21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>4324</fpage>
          -
          <lpage>4334</lpage>
          . KDD '
          <volume>16</volume>
          ,
          <year>2016</year>
          , pp.
          <fpage>815</fpage>
          -
          <lpage>824</lpage>
          . [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Knowledge-based conversational</article-title>
          [25]
          <string-name>
            <given-names>X.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          learning,
          <source>in: IJCKG '21</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          . in conversational recommendation, in: SIGIR '
          <fpage>21</fpage>
          , [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. D. V.</given-names>
            <surname>Hoang</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kan</surname>
          </string-name>
          ,
          <year>Perspectives 2021</year>
          , pp.
          <fpage>808</fpage>
          -
          <lpage>817</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <article-title>on crowdsourcing annotations for natural language</article-title>
          [26]
          <string-name>
            <given-names>P.</given-names>
            <surname>Stenetorp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pyysalo</surname>
          </string-name>
          , G. Topić,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ohta</surname>
          </string-name>
          , S. Ana-
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>processing</surname>
          </string-name>
          ,
          <article-title>Language resources and evaluation 47 niadou</article-title>
          , J. Tsujii,
          <article-title>Brat: a web-based tool for nlp-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          (
          <year>2013</year>
          )
          <fpage>9</fpage>
          -
          <lpage>31</lpage>
          .
          <article-title>assisted text annotation</article-title>
          ,
          <source>in: ACL '12</source>
          ,
          <year>2012</year>
          , pp. [12]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <volume>102</volume>
          -
          <fpage>107</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          , Improving conversational recommender sys- [27]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Zhang</surname>
          </string-name>
          , Data Annotation and
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>tems via knowledge graph based semantic fusion</article-title>
          ,
          <source>Preprocessing</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>in: KDD '20</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1006</fpage>
          -
          <lpage>1014</lpage>
          . [28]
          <string-name>
            <given-names>P.</given-names>
            <surname>Röttger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Vidgen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hovy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Pierrehumbert</surname>
          </string-name>
          , [13]
          <string-name>
            <given-names>Y.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Towards en- Two contrasting data annotation paradigms for sub-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>riching responses with crowd-sourced knowledge jective nlp tasks</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2112</volume>
          .
          <fpage>07475</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <article-title>for task-oriented dialogue</article-title>
          ,
          <source>in: MuCAI '21</source>
          ,
          <year>2021</year>
          , p. [29]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          , ADVISOR SUITE -
          <article-title>A knowledge-based</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          3-
          <fpage>11</fpage>
          .
          <article-title>sales advisory system</article-title>
          ,
          <source>in: ECAI '04</source>
          ,
          <year>2004</year>
          , pp. [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Arjannikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sanden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Verifying 720-724.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <article-title>tag annotations through association analysis</article-title>
          , in: [30]
          <string-name>
            <given-names>K.</given-names>
            <surname>McCarthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Salem</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          , Experience-based
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <source>ISMIR '13</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>195</fpage>
          -
          <lpage>200</lpage>
          . critiquing: Reusing critiquing experiences to im-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>