<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>October</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>“I'm looking for something like ...”: Combining Narratives and Example Items for Narrative-driven Book Recommendation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Toine Bogers</string-name>
          <email>toine@hum.aau.dk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marijn Koolen</string-name>
          <email>marijn.koolen@di.huc.knaw.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Digital Infrastructure, Humanities Cluster, Royal Netherlands Academy of Arts and Sciences</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Science, Policy and Information Studies, Department of Communication &amp; Psychology, Aalborg University Copenhagen</institution>
          ,
          <country country="DK">Denmark</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>7</volume>
      <issue>2018</issue>
      <abstract>
        <p>Current-generation recommendation algorithms are often focused on generic ratings prediction and item ranking tasks based on a user's past preferences. However, many recommendations are more complex with specific criteria and constraints on which items are relevant. This paper focuses on a particular type of complex recommendation needs: narrative-driven recommendation, where users describe their needs in short narratives, often with one or more example items that fit that need, against a background of historical preferences that may not be spelled out in the narrative, but do play a role in their considerations. Previous work has shown that numerous examples of such complex needs exist on the Web, yet current-generation systems ofer limited to no support for these needs. In this paer, we focus on narrative-driven book recommendation in the context of LibraryThing users posting recommendation requests in the discussion forums. We propose several new algorithms that take advantage of these narratives and example items as well as hybrid systems, the majority of which significantly outperform classic collaborative filtering. We show that narrative-driven recommendation is indeed a complex scenario that requires further study. Our findings have consequences for system design and development not only in the book domain, but also in other domains where users express focused recommendation needs, such as movies, television, games, and music.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Research on recommendation algorithms for ratings prediction and
item ranking has resulted in a better understanding of these tasks
Knowledge-aware and Conversational Recommender Systems (KaRS) Workshop 2018
(co-located with RecSys 2018), October 7, 2018, Vancouver, Canada.
2018. Copyright for the individual papers remains with the authors. Copying permitted
for private and academic purposes. This volume is published and copyrighted by its
editors..
and an array of algorithms with state-of-the-art performance.
However, not all recommendation needs can be solved by providing a
‘generic’ set of recommendations based on a user’s past preferences.
In many cases, recommendation is a more complex problem and
often just a single stage in a user’s more complex background need.
These needs can place a variety of constraints on which
recommendations are interesting and appropriate to the user—and they are
unlikely to be satisfied well by the current generation of ratings
prediction and item ranking algorithms.</p>
      <p>
        Relatively little research has been done on these complex
recommendation needs and how much of a problem they still pose for
current-generation recommender systems. In 2017, Kang et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]
analyzed the composition of such complex needs by exploring how
people use natural language to ask for movie recommendations
using a chatbot. They coded 498 natural-language interactions with
the chatbot, revealing a complex mix of objective, subjective and
navigational aspects of the users’ recommendation needs. Kang
et al. argue that many of these recommendation needs are hard to
support using current systems and data sources. This argument
was echoed in a similar study by Bogers and Koolen [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], who
analyzed a set of 974 complex book requests from the LibraryThing
(LT) forums. They argue that these are examples of narrative-driven
recommendation (NDR), which they define as a scenario that
combines (1) a narrative description of the desired aspects of relevant
items provided by the user, and (2) user preference information,
either in the form of a transaction log or of a user-provided
miniprofile containing positive and/or negative examples of other items.
Analysis of these 974 requests revealed them to be similar in many
ways to the movie recommendation needs collected by Kang et al.
[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Bogers and Koolen estimate that over 25,000 of such complex
requests exist in the LT forums alone, with an order of magnitude
more requests available all across the Web for a variety of domains.
      </p>
      <p>
        However, neither Kang et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] nor Bogers and Koolen [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
propose any recommendation algorithms to address such needs
and provide focused recommendations. To the best of our
knowledge, no algorithms exist that can tackle these types of complex
recommendation needs. We attempt to remedy this by building on
the work by Bogers and Koolen with a first attempt at designing
different NDR algorithms that incorporate both the textual narratives
as well as the example books and authors provided by the users.
Our experiments show that most of our algorithms outperform
state-of-the-art collaborative filtering (CF). Hybrid recommenders
that combine complementary algorithms based on narratives and
examples provide a significant performance boost over our
individual algorithms. Nevertheless, recommendation accuracy in general
suggests there is still room for improvement with many dificult
subproblems that require further study.
      </p>
      <p>Because of the relative novelty of our specific book
recommendation scenario, we start by describing it and our methodology
in greater detail in the next section. Sections 3 and 4 respectively
describe algorithms that utilize the textual narratives and
example items to produce relevant book recommendations. Our eforts
at hybrid recommendation are described in Section 5. We discuss
related work in Section 6 and conclude in Section 7.
2</p>
    </sec>
    <sec id="sec-2">
      <title>METHODOLOGY</title>
      <p>
        Bogers and Koolen [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] introduced the notion of Narrative-Driven
Recommendation (NDR) and examined the prevalence, composition,
and complexity of such needs on the LibraryThing (LT) forums.1
They argued that narrative recommendation needs have elements
of both recommendation and information retrieval (IR). Our work
in this paper builds on this notion and uses the requests from the LT
forum as an experimental testbed for NDR. LT is a social cataloging
website that allows its users to add books to their profile and tag,
rate and review them. The LT forums support users in discussing
their reading experiences as well as ofer a venue for requesting
book recommendations.
      </p>
      <p>Figure 1 illustrates our complex NDR scenario on the LT forums.
Users can post requests for book recommendations to a forum
discussion group, describing their recommendation need containing
both a narrative description of the kinds of books they want to
read and examples of (un)suitable) books and/or authors. These
examples form a so-called mini-profile , a temporary representations
of that user’s current preferences. Other users can reply to these
posts and come with suggestions for relevant books or authors to
read and explore. Many of the users interacting in these discussion
threads have their own LT profile of their book preferences. On LT,
these books are represented with diferent types of metadata and
user-generated content. We wish to explore the role these user
preferences and book representations can play in generating accurate
book recommendations.</p>
      <p>We expect that tailored recommendations based on individual
requests will provide better recommendations than generating a
list of book recommendations using a state-of-the-art CF algorithm
based on a user’s entire profile.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Datasets</title>
      <p>
        To develop and evaluate our NDR algorithms, we use the
Amazon/LibraryThing (A/LT) dataset. The A/LT dataset has been used
in the INEX Interactive Track [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and the Social Book Search Lab
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and has 2.8M book records consisting of both traditional book
metadata as well as user-generated content (UGC) harvested from
Amazon (user reviews and ratings) and LT (user tags), as visualized
in the bottom-left corner of Figure 1. The same book (or work) can
have many diferent editions, each of them with a diferent ISBN. LT
uses its own internal ‘work ID’ to map books to all associated ISBNs.
The book records in the A/LT collection use ISBNs as identifiers,
so to avoid including duplicates in our recommendation lists, we
mapped these ISBNs to the LT book identifiers using a mapping
ifle provided by LT. Not all ISBNs could be mapped to a general
      </p>
      <sec id="sec-3-1">
        <title>1Available at https://www.librarything.com/groups.</title>
        <p>work on LT, so for all intents and purposes we are recommending
books from a collection of 2,654,082 book records, corresponding
to 1,904,950 distinct works.</p>
        <p>To enable recommendations based on the example authors in
mini-profiles, we assigned unique author IDs to each author in the
A/LT collection. This was necessary, because the A/LT dataset has
inconsistent author naming and no author identifiers. We matched
duplicate authors by converting their name strings to lowercase
and removing punctuation and then assigned IDs. This reduced
the list of authors from 968,703 author strings to 916,479 unique
author IDs. We also checked for matches between names with
and without middle names and/or initial and collapsed them into
the same author ID. While not perfect—John Smith and John D.
Smith need not refer to the same person—inspection on a random
sample of 100 automatic matches revealed an precision of 89%.
Applying this rule further reduced the number of unique authors
from 916,479 to 849,578. We then matched all example authors in
the mini-profiles to their A/LT author IDs.</p>
        <p>In addition to the A/LT dataset, we also use a crawl of LT user
profiles with cataloging dates, tags, and ratings. This crawl was
conducted in 2012 and 2013 and contains over 66K LT user profiles,
with over 29M cataloging transactions and 4.5M distinct books. The
LT user profiles dataset has 4,409,399 ratings by 38,174 users,
representing less than 15% of the total number of cataloging transactions.
Over 42% of users have assigned no ratings at all, and the 58% who
have, rate only a small fraction of the books in their catalog. We
therefore ignore the ratings and treat each transaction as implicit
positive feedback on user preferences.
2.2</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental setup</title>
      <p>
        We use the same subset of 974 narrative requests from the LT
forums collected by Bogers and Koolen [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] as a starting point for
our experiments. These are requests where a LT user started a
discussion thread requesting book recommendations and at least one
other member posted a reply with suggestions. In our evaluation,
we excluded all requests of those users not present in our LT user
profile crawl. This resulted in a set of 331 requests, of which 298
have ≥ 1 example book with a mean (median) of 2.6 (2) books,
and 121 have ≥ 1 example author with a mean (median) of 2.0 (1)
authors. All requests have at least one example book or author.
These examples make up the mini-profile for each request.
      </p>
      <p>When recommending based on these example books and authors,
we do not distinguish between positive and negative examples.
Instead, we consider all example books as equally relevant sources of
recommendation; we are aware of the bias this may introduce.
Sentiment analysis could conceivably be used to determine whether
examples are cast in a positive or negative light. However, this
would also add another layer of complexity and thereby make it
harder to tease apart the accuracy of our recommendation
algorithms from our example sentiment detection component.</p>
      <p>We split the 331 requests into a random 50-50 train/test split with
166 training requests and 165 test requests. Default train/test splits
may have more training items, but none of our proposed algorithms
have more than one parameter to be optimized. In addition, the
recommendation narratives show great variety in terms of narrative</p>
      <p>User preferences</p>
      <p>for
Books
represented by
has
has</p>
      <p>Thread
length as well as example counts and popularity; a larger test set
would allow for better analysis of our algorithms.</p>
      <p>The requests were posted in the period 2006–2012, with some
users posting multiple requests during this period. At the time of
these requests, the same user will have diferent user profiles. This
introduces a challenge in training CF models, as their diferent
profiles should not be represented as diferent users. For a user u
posting request r1 at time t1 with profile p1 and r2 at time t2 with
profile p2, we need to either update the model in between requests,
or use only the profile at p1. There are some users with multiple
requests in our dataset, so to keep our baseline setup simple for
training, we represent as user u with multiple requests at t1 and t2
by their profile at time t1. For 53 out of 331 requests, the pre-request
profile ended up being based on the earliest request.</p>
      <p>We realize that our evaluation uses a relatively low number of
requests and only a single train-test split, but wish to emphasize
that NDR is diferent from generic ratings prediction and top-k
recommendation, and instead a hitherto unaddressed problem. In
our scenario, we are dealing with recommendation needs with a
strong IR element, where evaluation using 50-100 topics is common.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>The book requests can be evaluated from diferent perspectives,
representing diferent recommendation tasks. First, there are the
user profiles with all the books that each user added to their
personal catalog, with a cataloging date that can be used to determine
which books were cataloged before a user posted their request (=
pre-request cataloged items) and which books were cataloged
afterwards (= post-request cataloged items). For a typical recommendation
scenario, the former represent the user’s profile, the latter the books
to recommend. The request also receives replies from other users
with suggestions of books relevant to the request. Other users
often report having consulted the requester’s profile to target their
suggestions, although they sometimes suggest books the requester
already cataloged. These suggestions represent a recommendation
task more focused on the recommendation need. A third perspective
is the intersection of the suggestions and the post-request cataloged
items. We refer to these as post-request cataloged suggestions (PCS).
They represent focused and successful recommendations.</p>
      <p>
        For NDR, the recommendation needs are central, so we use the
suggestions as relevant recommendations, with the PCS as the most
relevant ones. Post-request cataloged items that are not suggested
in the request thread are considered outside the focus of the
recommendation need. We therefore adopted the relevance grading used
in clef 2016 Social Book Search Lab [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and consider suggestions
relevant with relevance value rv = 1 and PCS as relevant with
rv = 8. To enable baseline CF and content-based filtering (CBF)
approaches, the training data for a user are the pre-request catalogued
items of that user prior to their request.
      </p>
      <p>We assume this is a high-precision task, so we focus on top-N
measures. We are using graded relevance judgements, so we use
nDCG@10 as our main metric and report MRR to provide further
insights into the rank where users can find the first relevant item.
All reported statistical significance tests are one-tailed bootstrap
with 100,000 re-samples at levels p &lt; 0.05 and p &lt; 0.01.
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Baseline</title>
      <p>As argued in Section 1, narrative-driven book recommendation
ofers a unique and complex recommendation scenario. Generic
book recommendation can straightforwardly be addressed using
state-of-the-art CF algorithms, such as matrix factorization.
However, applying CF to the problem of NDR is unlikely to provide
many relevant, on-topic recommendations. To provide evidence
for this assertion, we compare our narrative- and example-driven
algorithms to a CF baseline.</p>
      <p>
        We used the LightFM toolkit2 by Kula [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] to train a baseline CF
model. LightFM uses Stochastic Gradient Descent for training, with
four diferent loss functions: Logistic, Weighted Approximate-Rank
Pairwise (WARP, [
        <xref ref-type="bibr" rid="ref41">41</xref>
        ]), Bayesian Personalised Ranking (BPR, [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ])
and k-th order statistic loss (k-OS WARP, [
        <xref ref-type="bibr" rid="ref42">42</xref>
        ]). Since most users
do not rate books on LT, we assume the presence of an book in a
user’s catalog as implicit positive feedback. We optimized for top-N
recommendation with N = 10 and evaluated after every 10th epoch
      </p>
      <sec id="sec-6-1">
        <title>2Available at https://github.com/lyst/lightfm.</title>
        <p>1.
2.
3.</p>
        <p>(f) (a)
(b)
up to 100 epochs. Optimal performance was achieved with WARP
running for 50 epochs and 300 factors. The narrative requests focus
the user’s recommendation need on specific criteria. We expect that
basing recommendations on the entire user profile leads to poor
performance. However, given the modest number of examples in
the mini-profiles, we expect CF to have little signal to work with,
akin to the cold-start problem, resulting in poor performance.</p>
        <p>
          For the CBF approaches, we used the language modeling
approach as implemented in the Indri toolkit. Optimal retrieval
parameter values3 for the diferent versions of the A/LT collection
are based on earlier experiments that share a similar setup in that
the narrative was also used as a query [
          <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>NARRATIVE-DRIVEN RECOMMENDATION</title>
      <p>One source of information about the user’s recommendation need is
the narrative posted to the LT forums. As shown earlier in Figure 1,
each request consists of a narrative and, as part of that narrative, a
set of example books and authors. In this section, we propose a NDR
algorithm that uses the textual representation of the narrative to
identify and recommend related books. Figure 2 shows a schematic
illustration of our approach.</p>
      <p>For each request, we take the narrative description and use CBF
to match it against the collection of 2.6 million books. Here, the
requester’s narrative is treated as a temporary representation of
3An overview of these parameter values can be found at http://anon.ymiz.ed/url.</p>
      <p>MRR
0.024 0.059△
0.026 0.052
0A.u0t4h8o▲red 0.111▲
0b.0o6o7k▲(s) 0.164▲
0.074▲ 0.181▲
0.074▲ (a) (f0).180▲
their user profile. Each book in the A/LT colle(cct)ion is represented
using diferent types of metadata fields and m a(dt)ched against the
narratives using retrieval algorithms. In our(ee)xperiments, we
distinguish between six categories of metadata representations. One
Authored
category of metadata (a) contains the core m(eg)tadata, such asbotoitkl(es),
author, and publisher. Curated metadata (b), contains index terms
and Dewey codes form another category, while Tags (c) are the set
of distinct tag1s. assigned to each book by the LT users. The1A.mazon
reviews (d) fo2r.m another collection representation; the
c2o.mbination of reviews and tags (e) represents the UGC associated with
each book. Fi3n.ally, combining all metadata fields (f) is th3e. sixth
collection representation.
3.1</p>
    </sec>
    <sec id="sec-8">
      <title>Results &amp; Analysis</title>
      <p>Table 1 shows the results of the narrative-based recommendation
algorithm. Diferences are tested for statistical significance using
one-tailed bootstrap with 100,000 resamples. CBF based on
narratives outperforms CF, which is not surprising, given the focused
nature of the task. CF will recommend items from similar users,
but not necessarily within the confines of the information need.
Within the CBF approaches, user-generated content clearly
outperforms other metadata, possibly because it reflects the language
of requesters better than e.g. curated metadata. Reviews perform
significantly better than tags. The combination of reviews and tags
is also significantly better than tags alone, but not than reviews
alone. Reviews use a broader vocabulary than tags to match
narrative requests, and maybe cover book aspects that are both relevant
in reviewing and in searching. We found no correlation between
the performance of diferent NDR approaches and the length of
narratives or the number of examples. Furthermore, we looked
at the popularity of recommended items, in terms of how many
LT users cataloged them. CF tends to recommend items from the
more popular end of the spectrum than the suggestions (including
the successful suggestions) and the examples, while CBF tends to
recommend more items from the long tail, especially when using
curated metadata. Here another advantage of UGC is revealed, as
it favors more popular books compared to curated metadata, in
that popular books have more UGC than obscure books, so have
1.
2.
3.</p>
      <p>EDR-1
a higher probability of being retrieved, but favors more obscure
books compared to CF in that it penalizes the topic drift in the huge
amounts of UGC of popular books.
to calculate book-to-book similarities as option (g). Sections 4.1 and
4.2 describe the two algorithms in more details as well as their
efectiveness compared to traditional CF and NDR.
4</p>
    </sec>
    <sec id="sec-9">
      <title>EXAMPLE-DRIVEN RECOMMENDATION</title>
      <p>The second source of information about the user’s recommendation
need is the mini-profile they provide, i.e., the example books and/or
authors they mention as relevant to their current need. We proposed
two algorithms that each uses one of the example types as input: (1)
example-driven recommendation using similar books (EDR-1), and
(2) example-driven recommendation using similar authors (EDR-2).
Figure 3 shows a schematic illustration of both algorithms.</p>
      <p>Both algorithms can use diferent representations of the books
and authors, corresponding to the six metadata groups introduced
in Section 3: (a) traditional metadata, (b) curated metadata, (c) tags,
(d) reviews, (e) UGC (= tags and reviews combined), and (f) all
representations combined. In addition, we used the user-item matrix
4.1</p>
    </sec>
    <sec id="sec-10">
      <title>Recommendation using example books</title>
      <p>Our first example-driven algorithm EDR-1 uses the example books
in the mini-profiles and is visualized on the left-hand side of Figure 3.
As mentioned in Section 2, we do not distinguish between positive
and negative examples, but consider all example books as relevant
sources of recommendation. The EDR-1 algorithm consists of two
basic steps: (1) identifying unseen books that are similar to each of
the example books, and (2) merging these sets of similar books to
get a single ranked list of book recommendations.</p>
      <p>(1) Identifying similar books. In the first step of EDR-1 for each
example book mentioned in a mini-profile, we can identify similar
books using one of two data sources: metadata representations of
each book, or item-to-item similarities calculated on the user-book
matrix. These roughly correspond to the diference between CBF
and CF. With regard to metadata-based matching, the metadata
representations of each example book are matched against the 2.6
million book representations in the A/LT collection. Matching can
be done for each of the six representation collections (a)-(f)
introduced in Section 2. Each example book representation is converted
into an Indri-safe query and matched against each of the collection’s
book representations, producing ranked lists of book
recommendations. Since the collection is indexed at the level of individual ISBNs,
the same book can be represented by multiple XML representations,
one for each ISBN. In each case, we pick the representation that
contains the largest amount of text as the canonical representation
of a book. Not every LT work could be mapped to an ISBN identifier;
52 out of 740 (7.0%) of all of our mini-profile example books could
not be mapped. This means that for requests without mappable
books no recommendations could be generated.</p>
      <p>Another issue is that for the richer collections, such as (f) with
all metadata fields combined, these book representations quickly
become too long for eficient document retrieval. For instance, the
median word count of book records for the (f) representation is
27,988 words, while the longest representation contains 210,790
words. Query-document matching in IR toolkits like Indri is
optimized for relatively short queries, so to speed up matching against
2.6 million book representations, we capped each book
representation at an empirically determined upper limit of the 500 most
frequent words in each representation.</p>
      <p>To identify similar books based on user preferences we calculated
item-to-item similarities from the user-book matrix using the CF
model trained with LightFm, as described in Section 2.4.</p>
      <p>
        (2) Merging sets of similar books. Of the 298 requests with
example books, 120 requests (40.3%) contain only a single example book.
In these cases the ranked list of relevant book representations
directly corresponds to the list of book recommendations after being
mapped back from ISBNs to LT works. The remaining 59.7% contain
multiple example books, which means they yield multiple results
lists per request, with many books occurring on several lists. We
merge these lists using the CombSUM fusion method introduced
by Fox and Shaw [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which has been shown to be efective for
fusing recommendation runs [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. CombSUM sums the normalized
similarity scores for the same book across diferent results lists,
corresponding to a linear combination with equal weights. Before
applying CombSUM, we score-normalize each results list according
to the formula simnorm = sismimormigainxa−l−sismimmminin , where original, max, min,
and norm represent the original, highest, lowest and normalized
similarity scores respectively.
      </p>
      <p>One parameter that could influence recommendation accuracy
is the number of similar books to include from each results list.
Including a larger number of results—the top 500 vs. the top 10—is
likely to result in more books occurring in multiple lists. Another
perspective is to see this as the number of nearest (book) neighbors
to recommend. We denote this parameter as k and optimized its
value on our training set of 166 requests, separately for each of
our seven collection representations (a)-(g). We varied the value
of k in steps of 50 up to k = 1000. The line chart on the left-hand
side of Figure 4 shows how varying k influences recommendation
performance. Overall, performance tends to be fairly stable with
regard to k. The optimal values of k for EDR-1 are included in
Table 2.
4.2</p>
    </sec>
    <sec id="sec-11">
      <title>Recommendation using example authors</title>
      <p>Our second example-driven algorithm EDR-2 uses the example
authors in requests, and is shown on the right-hand side of Figure 3.
It consists of three stages: (1) identifying authors that are similar
to each of the example authors; (2) merging these sets of similar
authors to get a single ranked list of similar authors; and (3)
replacing each similar author by their authored books and merging these
to get a single ranked list of book recommendations.</p>
      <p>(1) Identifying similar authors. In the first stage of EDR-2 for each
example author mentioned in a mini-profile, we identify similar
authors by matching their metadata representations against those
of all authors in the A/LT collection. To make this possible, we
re-indexed the A/LT collection at an author-level by aggregating
all of an author’s book representations into a single author-level
representation. For each book, we again took the associated ISBN
with the largest amount of text. These book representations were
then aggregated into a single author representation, one for each
of the six metadata representations (a)-(f) introduced in Section
2. This resulted in an author-centric collection containing 849,578
author representations. Each example author’s representation is
converted into an Indri-safe query and matched against each of
the collection’s author representations, producing a ranked list of
author recommendations. As with EDR-1, we capped each author
representation at a limit of the 500 most frequent words.</p>
      <p>(2) Merging sets of similar authors. The 121 requests with example
authors contained an average of 2.0 author per request. As for
EDR1, we used CombSUM fusion to merge normalized author lists
belonging to the same request. We also examined the influence
of the number of similar authors k to include from each results
list; the results of this optimization can be found in Because there
are fewer authors than books in the A/LT collection and because
expanding authors to their authored books results in a much larger
number of recommended books, we varied k for EDR-2 in smaller
step sizes. The line chart on the right-hand side of Figure 4 shows
how k influences performance. There are more pronounced peaks
for EDR-2 than for EDR-1, but , performance tends to be fairly
stable with regard to k. The optimal values of k for EDR-2 are also
included in Table 2.</p>
      <p>(3) Inserting &amp; merging authored books. To arrive at lists of
recommended books, we expanded each recommended author by their
authored books and assigned each book the similarity score of that
author. Because of errors in mapping ISBNs to work IDs, some books
occurred multiple times for the same request; we again summed
all of a book’s individual scores to produce a single list of
recommended books, capped at 1000 results.
4.3</p>
    </sec>
    <sec id="sec-12">
      <title>Results &amp; Analysis</title>
      <p>Table 2 shows the results of EDR-1 and EDR-2, with scores over
all test topics and over only those that have example books and
example authors respectively. EDR-1 is less efective than NDR
but more than EDR-2, with UGC again more efective than other
metadata. This is possibly due to requests having fewer authors
than books or that (especially prolific) authors introduce more topic
drift. EDR-2 is more efective than CF for requests with example
authors when using UGC.
5</p>
    </sec>
    <sec id="sec-13">
      <title>HYBRID RECOMMENDATION</title>
      <p>
        In this section we look at hybrid systems that combine NDR and
EDR. More specifically, we apply results fusion, where the outputs
of diferent recommendation algorithms are combined into a single
results list, which is a form of weighted hybridization according to
the hybrid recommendation taxonomy by Burke [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This results
fusion is done by first score-normalizing the results of the individual
systems and then re-ranking results using a weighted combination
of the prediction scores, i.e., a linear combination or weighted
CombSUM [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. We range λ weights of runs between 0.1 and 0.9
where combined run weights sum to 1.
      </p>
      <p>Table 3 contains the results of the hybrid recommenders, which
we compare against the best single approach from the previous
sections, i.e. the NDR (g) system which uses UGC. We only report the
best performing weights. We discuss the results with a per-request
analysis to explain the diferences. Combining NDR and EDR-1 (f)
results in a significant improvement in nDCG@10, although not
in MRR. The narratives and examples complement each other as
representations of the information need. The combination increases
the number of requests with non-zero scores from 55 to 68, and
also improves more request scores than it hurts.</p>
      <p>Combining NDR with EDR-1 with user preferences ((g)) strongly
improves precision for 20 requests, but slightly hurts it for 35
requests, which is why scores are higher, but not significantly so.
Again, user preferences are good for early precision, but beyond
that the signal gets so weak that popular items pollute the ranking.
The combination with EDR-2 (f) run leads to only 17 requests having
a diferent score, but significantly hurts performance. The similar
author recommendations appear to cause topic drift. Combining
all three approaches leads to small but statistically insignificant
improvements, mainly reducing the number of requests scoring
zero. Combining EDR-1 and EDR-2 leads to small improvements
over the individual EDR approaches, again, mainly by reducing the
number of zero-scoring requests.</p>
      <p>EDR-1 (Recommendation using similar books)
0 50 100 150 200 250 300 350 400 450 500 550 600 650 700 750 800 850 900 950 1000
Metadata Curatedmetadata Tags Reviews Reviews+Tags Allfields</p>
      <p>EDR-2 (Recommendation using similar authors)
0 50 100 150 200 250 300 350 400 450 500 550 600 650 700 750 800 850 900 950 1000</p>
      <p>Metadata Curatedmetadata Tags Reviews Reviews+Tags Allfields</p>
      <p>Overall, the hybrids exploit the complementarity of the data
sources, leading to more requests with relevant recommendations,
although it is hard to improve over the competitive NDR
baseline. However, our experiments do show that combining
recommendation based on the narrative and the examples can lead to
significantly improved performance.</p>
    </sec>
    <sec id="sec-14">
      <title>6 RELATED WORK</title>
      <p>
        Our work on book recommendation is far from the first: numerous
book recommenders have been proposed over the past years. Most
of the related work has focused on exploring and comparing CF,
CBF and graph-based algorithms for generic book ranking [
        <xref ref-type="bibr" rid="ref10 ref27 ref31 ref33 ref39 ref40">10, 27,
31, 33, 39, 40</xref>
        ]. At least two shared tasks have focused on book
recommendation: the Semantic Web Evaluation Challenge [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and
the Social Book Search lab [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
      </p>
      <p>
        Our focus on expressing recommendation needs through textual
narratives is reminiscent of the work by Adomavicius et al. [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]
on query-driven recommendation. They proposed Reqest, a
structured query language for customizing recommendations, which
can be used to tailor recommendation needs beyond the traditional
“give me items I would like” task. However, it does not allow for
purely textual representations of recommendation needs. Hariri
et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] introduced a query-driven, context-aware recommender
system that provides recommendations based on a user’s
preferences and can be adapted to a given context that represents the
short-term interests or needs of a user. Drenner et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] perform
movie linking between a movie recommendation Web site and a
movie-oriented discussion forum. Through automatic detection
and an interactive component, the system recognizes references to
movies in the forum and adds recommendation data to the forums
and conversation threads to movie pages.
      </p>
      <p>
        Narrative descriptions of needs and interests are typical of
conversational recommendation, where one person describes the kind
of items they like and what they would be interested in, while
others provide suggestions and explanations. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] analysed 498
movie recommendation conversations with a chatbot and found
that such conversations elicit many complex aspects of
recommendation needs. As such, NDR is related to conversational and
critiquing-based recommender systems, which aim to elicit more
information about a user’s recommendation needs through
interaction and dialog [
        <xref ref-type="bibr" rid="ref11 ref29 ref30 ref34 ref38 ref9">9, 11, 29, 30, 34, 38</xref>
        ].
      </p>
      <p>
        The narrative requests we use for recommendation also bear
similarities to online product reviews in terms of their composition
and complexity. Reviews cover diferent aspects of a product with
the author describing likes and dislikes. The overall review score
of a product can be seen as personal ranking of the importance of
aspects. O’Mahony and Smyth [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] investigated ways of
recommending reviews that ofer contrasting views to help users make
informed choices. Dong et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] extracted topical and sentiment
information from reviews to identify the most informative reviews.
      </p>
      <p>
        Our work on example-driven recommendation also has a few
parallels with earlier work. For instance, Liu et al. [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] use seed
items for cold-start recommendation, which is similar to our use of
mini-profiles. Schnabel et al . [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] studied shortlists as an interface
component for recommender systems, where users iteratively build
up a list of candidate items to consume, and found that they support
the user’s decision process, and the elicited implicit feedback can
increase recommendation quality. The main diference with the
NDR scenario is that the items in our mini-profiles are not
candidate items for final recommendation, but examples to illustrate the
recommendation need. However, the interface and system design
proposed by Schnabel et al. [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] would support both scenarios. Our
LT ratings data was sparse with many of the transactions lacking
ratings. The work by Sharma et al. [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ] suggests a possible way of
remedying this by modeling the mini-profile as a set of ratings by
using any ratings available and predicting the missing ratings in the
set. Sharma et al. asked Movielens users to give a single rating for
0.074
sets of items (movies) that they had previously rated individually
and investigated techniques for predicting item ratings based on
the user’s set ratings and ratings of other items not in the set.
      </p>
      <p>
        Our focus on complex recommendation mirrors the increased
focus in IR on complex search tasks and how best to support them
[
        <xref ref-type="bibr" rid="ref17 ref24">17, 24</xref>
        ]. Some of these tasks are located on the–not necessarily
clear—edge between search and recommendation [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Complex
narratives are commonly used in IR evaluation and test collection
building to guide assessors [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. They have been also used in
interactive IR [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] to study how users perform complex search tasks.
The Social Book Search campaigns at inex [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and clef [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] found
complex, narrative-focused information needs to be common in
online book discussion forums, such as GoodReads and LibraryThing.
We build on their work in this paper. However, Koolen et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]
found that the narratives written to assess artificially constructed
topics for IR evaluation are of a diferent nature than the narratives
that users write when asking peers for recommendations. In
addition, the choices that users make are very diferent from traditional
relevance judgments used in IR.
7
      </p>
    </sec>
    <sec id="sec-15">
      <title>DISCUSSION &amp; CONCLUSIONS</title>
      <p>In this paper we explored techniques to address narrative-driven
recommendation needs. This is a novel, complex scenario in which
users have specific recommendation needs that they express in a
request containing both a natural language description as well as
a mini-profile containing relevant example of books and authors.
The combination of these complex user needs and the richness and
heterogeneity of the item descriptions poses many challenges for
designing recommender systems that support this task.</p>
      <p>For the narrative part of the requests, we experimented with CBF
techniques on a range of book metadata and user-generated content.
We found that the latter is more efective as it more closely matches
the vocabulary of the user’s narrative and strikes a balance between
popularity and specificity. For example-driven recommendation
(EDR) on the mini-profiles we experimented with both CBF and CF
approaches, and found that example books ofer more focus and
thereby better recommendations than example authors. Again, UGC
is more efective than curated metadata and CF-based item-to-item
recommendation. Finally, we looked at hybrid systems that combine
NDR and EDR and found that NDR uses the more important signal,
but that carefully weighted combinations can lead to significant
improvements, especially in providing more requests with at least
some relevant recommendations.</p>
      <p>Overall we found that NDR is a challenging task that requires
multiple data sources and algorithms to solve. Nevertheless, there
are several issues that future work should address. One of them
is the detection of and diferentiation between positive and
negative examples using sentiment analysis techniques; currently all
examples are treated as positive examples, but this likely afects
performance negatively. Other algorithms, such as those from work
on conversational recommendation, or graph-based algorithms on
multi-partite networks containing user, book, author and tag nodes
could perhaps provide a more integrated approach to EDR. Finally,
testing our algorithms on other domains, such as movies, games
and music would be required to determine the generalizability of
our findings.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Gediminas</given-names>
            <surname>Adomavicius</surname>
          </string-name>
          , Alexander Tuzhilin, and
          <string-name>
            <given-names>Rong</given-names>
            <surname>Zheng</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>RQL: A Query Language For Recommender Systems</article-title>
          . (
          <year>2005</year>
          ). http://hdl.handle.
          <source>net/2451/ 14109</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Gediminas</given-names>
            <surname>Adomavicius</surname>
          </string-name>
          , Alexander Tuzhilin, and
          <string-name>
            <given-names>Rong</given-names>
            <surname>Zheng</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>REQUEST: A Query Language for Customizing Recommendations</article-title>
          .
          <source>Information Systems Research</source>
          <volume>22</volume>
          ,
          <issue>1</issue>
          (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Beckers</surname>
          </string-name>
          , Norbert Fuhr, Nils Pharo, Ragnar Nordlie, and Khairun Nisa Fachry.
          <year>2010</year>
          .
          <article-title>Overview and results of the inex 2009 interactive track</article-title>
          .
          <source>In International Conference on Theory and Practice of Digital Libraries</source>
          . Springer,
          <fpage>409</fpage>
          -
          <lpage>412</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Toine</given-names>
            <surname>Bogers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marijn</given-names>
            <surname>Koolen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Defining and Supporting Narrative-driven Recommendation</article-title>
          .
          <source>In RecSys '17: Proceedings of the Eleventh ACM Conference on Recommender Systems. ACM</source>
          ,
          <volume>238</volume>
          -
          <fpage>242</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Toine</given-names>
            <surname>Bogers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vivien</given-names>
            <surname>Petras</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Tagging vs</article-title>
          . Controlled Vocabulary:
          <article-title>Which is More Helpful for Book Search?</article-title>
          .
          <source>In Proceedings of iConference 2015</source>
          . iDEALS.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Toine</given-names>
            <surname>Bogers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Vivien</given-names>
            <surname>Petras</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>An In-depth Analysis of Tags and Controlled Vocabulary for Book Search</article-title>
          .
          <source>In Proceedings of iConference 2017</source>
          . iDEALS.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Toine</given-names>
            <surname>Bogers and Antal Van Den Bosch</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Fusing Recommendations for Social Bookmarking Websites</article-title>
          .
          <source>International Journal of Electronic Commerce</source>
          <volume>15</volume>
          ,
          <issue>3</issue>
          (
          <year>2011</year>
          ),
          <fpage>31</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Hybrid Recommender Systems: Survey and Experiments. User Modeling and User-</article-title>
          <source>Adapted Interaction 12</source>
          ,
          <issue>4</issue>
          (
          <year>2002</year>
          ),
          <fpage>331</fpage>
          -
          <lpage>370</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Konstantina</given-names>
            <surname>Christakopoulou</surname>
          </string-name>
          , Filip Radlinski, and
          <string-name>
            <given-names>Katja</given-names>
            <surname>Hofmann</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Towards Conversational Recommender Systems</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM</source>
          ,
          <volume>815</volume>
          -
          <fpage>824</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Maarten</surname>
            <given-names>Clements</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arjen P. de Vries</surname>
            , and
            <given-names>Marcel J.T.</given-names>
          </string-name>
          <string-name>
            <surname>Reinders</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>The Influence of Personalization on Tag Query Length in Social Media Search</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>46</volume>
          ,
          <issue>4</issue>
          (
          <year>2010</year>
          ),
          <fpage>403</fpage>
          -
          <lpage>412</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Jefrey</surname>
            <given-names>Dalton</given-names>
          </string-name>
          , Victor Ajayi, and Richard Main.
          <year>2018</year>
          .
          <article-title>Vote Goat: Conversational Movie Recommendation</article-title>
          .
          <source>In SIGIR '18: Proceedings of the 41st International ACM SIGIR Conference on Research &amp;#38; Development in Information Retrieval</source>
          .
          <fpage>1285</fpage>
          -
          <lpage>1288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Tommaso</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Noia</surname>
          </string-name>
          , Iván Cantador, and Vito Claudio Ostuni.
          <year>2014</year>
          .
          <article-title>Linked open data-enabled recommender systems: ESWC 2014 challenge on book recommendation</article-title>
          .
          <source>In Semantic Web Evaluation Challenge</source>
          . Springer,
          <fpage>129</fpage>
          -
          <lpage>143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Ruihai</surname>
            <given-names>Dong</given-names>
          </string-name>
          , Markus Schaal,
          <string-name>
            <surname>Michael P O'Mahony</surname>
            ,
            <given-names>and Barry</given-names>
          </string-name>
          <string-name>
            <surname>Smyth</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Topic Extraction from Online Reviews for Classification and Recommendation</article-title>
          .
          <source>In Proceedings of the Twenty-Third international joint conference on Artificial Intelligence</source>
          . AAAI Press,
          <fpage>1310</fpage>
          -
          <lpage>1316</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Sara</surname>
            <given-names>Drenner</given-names>
          </string-name>
          , Max Harper, Dan Frankowski, John Riedl, and
          <string-name>
            <given-names>Loren</given-names>
            <surname>Terveen</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Insert Movie Reference Here: A System to Bridge Conversation and Item-oriented Web Sites</article-title>
          .
          <source>In CHI '06: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>951</volume>
          -
          <fpage>954</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Edward</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Fox</surname>
            and
            <given-names>Joseph A.</given-names>
          </string-name>
          <string-name>
            <surname>Shaw</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Combination of Multiple Searches</article-title>
          .
          <source>In TREC-2 Working Notes</source>
          .
          <volume>243</volume>
          -
          <fpage>252</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Maria</surname>
            <given-names>Gäde</given-names>
          </string-name>
          , Mark Michael Hall, Hugo C. Huurdeman, Jaap Kamps, Marijn Koolen, Mette Skov, Toine Bogers, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Overview of the SBS 2016 Interactive Track</article-title>
          . In Working Notes of CLEF 2016 -
          <article-title>Conference and Labs of the Evaluation forum</article-title>
          , Évora, Portugal,
          <fpage>5</fpage>
          -
          <lpage>8</lpage>
          September,
          <year>2016</year>
          .
          <fpage>1024</fpage>
          -
          <lpage>1038</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Maria</surname>
            <given-names>Gäde</given-names>
          </string-name>
          , Mark Michael Hall, Hugo C. Huurdeman, Jaap Kamps, Marijn Koolen, Mette Skov, Elaine Toms, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2015</year>
          . First Workshop on Supporting Complex Search Tasks.
          <source>In Proceedings of the First International Workshop on Supporting Complex Search Tasks co-located with the 37th European Conference on Information Retrieval (ECIR</source>
          <year>2015</year>
          ), Vienna, Austria, March
          <volume>29</volume>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Negar</surname>
            <given-names>Hariri</given-names>
          </string-name>
          , Bamshad Mobasher, and
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Query-driven Contextaware Recommendation</article-title>
          .
          <source>In Proceedings of the 7th ACM conference on Recommender systems. ACM</source>
          ,
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Donna</given-names>
            <surname>Harman and Ellen M. Voorhees</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>TREC: An Overview</article-title>
          .
          <source>ARIST 40</source>
          ,
          <issue>1</issue>
          (
          <year>2006</year>
          ),
          <fpage>113</fpage>
          -
          <lpage>155</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Jie</surname>
            <given-names>Kang</given-names>
          </string-name>
          , Kyle Condif,
          <string-name>
            <given-names>Shuo</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Joseph A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          , Loren Terveen, and
          <string-name>
            <given-names>F. Maxwell</given-names>
            <surname>Harper</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Understanding How People Use Natural Language to Ask for Recommendations</article-title>
          .
          <source>In RecSys '17: Proceedings of the Eleventh ACM Conference on Recommender Systems. ACM</source>
          ,
          <volume>229</volume>
          -
          <fpage>237</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Marijn</surname>
            <given-names>Koolen</given-names>
          </string-name>
          , Toine Bogers, Maria Gäde,
          <string-name>
            <given-names>Mark M.</given-names>
            <surname>Hall</surname>
          </string-name>
          , Iris Hendrickx, Hugo C. Huurdeman, Jaap Kamps, Mette Skov, Suzan Verberne, and
          <string-name>
            <given-names>David</given-names>
            <surname>Walsh</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Overview of the CLEF 2016 Social Book Search Lab</article-title>
          . In Experimental IR Meets Multilinguality, Multimodality, and Interaction - 7th
          <source>International Conference of the CLEF Association, CLEF</source>
          <year>2016</year>
          , Évora, Portugal, September 5-
          <issue>8</issue>
          ,
          <year>2016</year>
          , Proceedings.
          <fpage>351</fpage>
          -
          <lpage>370</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Marijn</surname>
            <given-names>Koolen</given-names>
          </string-name>
          , Toine Bogers, Jaap Kamps, Gabriella Kazai, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Preminger</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Overview of the INEX 2014 Social Book Search Track</article-title>
          . In Working Notes for CLEF 2014 Conference, Shefield, UK,
          <source>September 15-18</source>
          ,
          <year>2014</year>
          .
          <fpage>462</fpage>
          -
          <lpage>479</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Marijn</surname>
            <given-names>Koolen</given-names>
          </string-name>
          , Toine Bogers, Jaap Kamps, and Antal van den Bosch.
          <year>2015</year>
          .
          <article-title>Looking for Books in Social Media: An Analysis of Complex Search Requests</article-title>
          .
          <source>In ECIR '15: Proceedings of the 37th European Conference on Information Retrieval (Lecture Notes in Computer Science)</source>
          , Vol.
          <volume>9022</volume>
          . Springer,
          <fpage>184</fpage>
          -
          <lpage>196</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Marijn</surname>
            <given-names>Koolen</given-names>
          </string-name>
          , Jaap Kamps, Toine Bogers, Nicholas J.
          <string-name>
            <surname>Belkin</surname>
            , Diane Kelly, and
            <given-names>Emine</given-names>
          </string-name>
          <string-name>
            <surname>Yilmaz</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Current Research in Supporting Complex Search Tasks</article-title>
          .
          <source>In Proceedings of the Second Workshop on Supporting Complex Search Tasks co-located with the ACM SIGIR Conference on Human Information Interaction &amp; Retrieval (CHIIR</source>
          <year>2017</year>
          ), Oslo, Norway, March
          <volume>11</volume>
          ,
          <year>2017</year>
          . 1-
          <fpage>4</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Marijn</surname>
            <given-names>Koolen</given-names>
          </string-name>
          , Jaap Kamps, and
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Kazai</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Social Book Search: Comparing Topical Relevance Judgements and Book Suggestions for Evaluation</article-title>
          .
          <source>In CIKM '12: Proceedings of the 21st ACM International Conference on Information and Knowledge Management</source>
          ,
          <string-name>
            <surname>Xue-wen Chen</surname>
            , Guy Lebanon,
            <given-names>Haixun</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Mohammed J. Zaki</surname>
          </string-name>
          (Eds.). ACM,
          <volume>185</volume>
          -
          <fpage>194</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Maciej</given-names>
            <surname>Kula</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Metadata Embeddings for User and Item Cold-start Recommendations</article-title>
          .
          <source>In Proceedings of the 2nd Workshop on New Trends on ContentBased Recommender Systems co-located with 9th ACM Conference on Recommender Systems (RecSys</source>
          <year>2015</year>
          ), Vienna, Austria,
          <source>September 16-20</source>
          ,
          <year>2015</year>
          . (CEUR Workshop Proceedings),
          <source>Toine Bogers and Marijn Koolen (Eds.)</source>
          , Vol.
          <volume>1448</volume>
          . CEUR-WS.org,
          <volume>14</volume>
          -
          <fpage>21</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Dong</given-names>
            <surname>Kun</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Research of personalized book recommender system of university library based on collaborative filter</article-title>
          .
          <source>Data Analysis and Knowledge Discovery</source>
          <volume>11</volume>
          (
          <year>2012</year>
          ),
          <fpage>44</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Qi</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Biao Xiang, Enhong Chen, Yong Ge, Hui Xiong, Tengfei Bao, and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Zheng</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Influential Seed Items Recommendation</article-title>
          .
          <source>In RecSys '12: Proceedings of the Sixth ACM Conference on Recommender Systems. ACM</source>
          ,
          <volume>245</volume>
          -
          <fpage>248</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Tariq</given-names>
            <surname>Mahmood</surname>
          </string-name>
          and
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Ricci</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Improving Recommender Systems with Adaptive Conversational Strategies</article-title>
          .
          <source>In Proceedings of the 20th ACM Conference on Hypertext and Hypermedia. ACM</source>
          ,
          <volume>73</volume>
          -
          <fpage>82</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Lorraine</surname>
            <given-names>McGinty</given-names>
          </string-name>
          and
          <string-name>
            <given-names>James</given-names>
            <surname>Reilly</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>On the Evolution of Critiquing Recommenders</article-title>
          .
          <source>In Recommender Systems Handbook</source>
          , Francesco Ricci, Lior Rokach, Bracha Shapira, and Paul B.
          <source>Kantor (Eds.)</source>
          . Springer,
          <fpage>419</fpage>
          -
          <lpage>453</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Raymond</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mooney</surname>
            and
            <given-names>Loriene</given-names>
          </string-name>
          <string-name>
            <surname>Roy</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Content-based Book Recommending using Learning for Text Categorization</article-title>
          .
          <source>In DL '00: Proceedings of the Fifth ACM Conference on Digital Libraries</source>
          .
          <fpage>195</fpage>
          -
          <lpage>204</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Michael P O'Mahony</surname>
            and
            <given-names>Barry</given-names>
          </string-name>
          <string-name>
            <surname>Smyth</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>A Classification-based Review Recommender</article-title>
          .
          <source>Knowledge-Based Systems 23</source>
          ,
          <issue>4</issue>
          (
          <year>2010</year>
          ),
          <fpage>323</fpage>
          -
          <lpage>329</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Soledad</surname>
          </string-name>
          Pera and
          <string-name>
            <surname>Yiu-Kai Ng</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>BReK12: a book recommender for K-12 users</article-title>
          .
          <source>In Proceedings of the 35th international ACM SIGIR conference on Research and development in information retrieval. ACM</source>
          ,
          <volume>1037</volume>
          -
          <fpage>1038</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Rachael</given-names>
            <surname>Rafter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Barry</given-names>
            <surname>Smyth</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Conversational Collaborative Recommendation-An Experimental Analysis</article-title>
          .
          <source>Artificial Intelligence Review</source>
          <volume>24</volume>
          ,
          <fpage>3</fpage>
          -
          <lpage>4</lpage>
          (
          <year>2005</year>
          ),
          <fpage>301</fpage>
          -
          <lpage>318</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Stefen</surname>
            <given-names>Rendle</given-names>
          </string-name>
          , Christoph Freudenthaler, Zeno Gantner, and
          <string-name>
            <surname>Lars</surname>
          </string-name>
          Schmidt-Thieme.
          <year>2009</year>
          .
          <article-title>BPR: Bayesian personalized ranking from implicit feedback</article-title>
          .
          <source>In Proceedings of the twenty-fifth conference on uncertainty in artificial intelligence</source>
          . AUAI Press,
          <fpage>452</fpage>
          -
          <lpage>461</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>Tobias</given-names>
            <surname>Schnabel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Paul N.</given-names>
            <surname>Bennett</surname>
          </string-name>
          , Susan T. Dumais, and
          <string-name>
            <given-names>Thorsten</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Using Shortlists to Support Decision Making and Improve Recommender System Performance</article-title>
          .
          <source>In WWW '16: Proceedings of the 25th International Conference on World Wide Web</source>
          .
          <fpage>987</fpage>
          -
          <lpage>997</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>Mohit</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Maxwell Harper</surname>
          </string-name>
          , and
          <string-name>
            <given-names>George</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Learning from Sets of Items in Recommender Systems</article-title>
          .
          <source>In eKNOW 2017: Proceedings of the Ninth International Conference on Information, Process, and Knowledge Management</source>
          .
          <fpage>59</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>Yueming</given-names>
            <surname>Sun</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yi</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Conversational Recommender System</article-title>
          .
          <source>In SIGIR '18: Proceedings of the 41st International ACM SIGIR Conference on Research &amp;#38; Development in Information Retrieval</source>
          .
          <fpage>235</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>Anand</given-names>
            <surname>Shanker</surname>
          </string-name>
          <string-name>
            <surname>Tewari</surname>
          </string-name>
          , Abhay Kumar, and Asim Gopal Barman.
          <year>2014</year>
          .
          <article-title>Book recommendation system based on combine features of content based filtering, collaborative filtering and association rule mining</article-title>
          .
          <source>In Advance Computing Conference (IACC)</source>
          ,
          <source>2014 IEEE International. IEEE</source>
          ,
          <fpage>500</fpage>
          -
          <lpage>503</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>Paula</given-names>
            <surname>Cristina</surname>
          </string-name>
          <string-name>
            <surname>Vaz</surname>
          </string-name>
          , David Martins de Matos, Bruno Martins, and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Calado</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Improving a hybrid literary book recommendation system through author ranking</article-title>
          .
          <source>In Proceedings of the 12th ACM/IEEE-CS joint conference on Digital Libraries. ACM</source>
          ,
          <volume>387</volume>
          -
          <fpage>388</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <surname>Jason</surname>
            <given-names>Weston</given-names>
          </string-name>
          , Samy Bengio, and
          <string-name>
            <given-names>Nicolas</given-names>
            <surname>Usunier</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Wsabie: Scaling up to large vocabulary image annotation</article-title>
          .
          <source>In IJCAI</source>
          , Vol.
          <volume>11</volume>
          .
          <fpage>2764</fpage>
          -
          <lpage>2770</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>Jason</surname>
            <given-names>Weston</given-names>
          </string-name>
          , Hector Yee, and
          <string-name>
            <surname>Ron</surname>
          </string-name>
          J Weiss.
          <year>2013</year>
          .
          <article-title>Learning to rank recommendations with the k-order statistic loss</article-title>
          .
          <source>In Proceedings of the 7th ACM conference on Recommender systems. ACM</source>
          ,
          <volume>245</volume>
          -
          <fpage>248</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>