<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Addressing Overchoice: Automatically Generating Meaningful Filters from Hotel Reviews</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>ISTVÁN VARGA</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Megagon Labs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tokyo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Recruit Co.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan YUTA HAYASHIBE</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Megagon Labs</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tokyo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Recruit Co.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Close to the sea. @Shizuoka Relaxing atmosphere. @Nagano Very helpful staf. @Gunma Atmosphere that makes you feel at home. @Iwate Suitable for sightseeing. @Kyoto Near downtown. @Aichi Delicious dinner. @Okinawa Fashionable rooms. @Hokkaido Child friendly. @Chiba Hotel with good access. @Akita</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Close to the sea. @Shizuoka Relaxing atmosphere. @Nagano Very helpful staf. @Gunma Atmosphere that makes you feel at home. @Iwate Suitable for sightseeing. @Kyoto Near downtown. @Aichi Delicious dinner. @Okinawa Fashonable rooms. @Hokkaido Child friendly. @Chiba Hotel with good access. @Akita</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <abstract>
        <p>In this paper we present a hotel filter recommendation method designed to address the cognitive load users face in an overchoice scenario. As online products and services are continuously diversifying, user needs are also becoming increasingly sophisticated. However, with more items to choose from, grasping the entire choice set and differentiating among all matching options becomes increasingly difficult, leading to sub-optimal outcomes. Conventional hotel reservation platforms provide with a limited set of additional filters, but these can not accommodate all intricate user needs. Employing natural language processing and machine learning techniques, we provide a simple framework that identifies meaningful filters from customer reviews. We define criteria and scoring methods to acquire relevant and interesting filters that may help customers refine their needs or even identify hidden, previously unknown ones. Our simulated user experiments show that our proposal is capable of identifying intricate and useful filters, leading to increased customer satisfaction. CCS Concepts: • Computing methodologies → Machine learning; • Information systems → Content ranking; Recommender systems; Rank aggregation; Similarity measures. Additional Key Words and Phrases: overchoice, clustering, filter recommendation</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>Online services have become not only ubiquitous, but indispensable in almost every aspect of our life. Nearly every
imaginable product or service is available through e-commerce transactions, including online shopping, restaurant
or hotel reservation to matchmaking. In a conventional hotel reservation service the customer is provided with an
interface that facilitates search using some of the most crucial criteria, typically objective queries that are meant to
reduce the choice set to a manageable size (e.g., number of visitors, length of stay, location, etc.).</p>
      <p>
        The emergence of e-commerce systems or online reservation services brought forward the advantage of having an
increased selection at the convenience of only a few clicks away. Both classic economics and psychology emphasize the
benefits of a larger number of choices [
        <xref ref-type="bibr" rid="ref34 ref41 ref42">34, 41, 42</xref>
        ]. However, it also raised a number of important challenges as well. One
such challenge is that the size of the choice set can be a cognitive load in the decision making process. Overchoice, or
having too many choices, can be detrimental, leading to anxiety or depression [
        <xref ref-type="bibr" rid="ref25 ref45 ref48">25, 45, 48</xref>
        ]. Even though a larger number
of choices is initially appealing, the consumers may feel less satisfied or convinced that they actually made the best
decision available [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Recent studies even suggest an inverted U-shaped relationship between customer commitment
and the number of available choices [
        <xref ref-type="bibr" rid="ref46">46</xref>
        ], with customers being more likely to find an item to their liking with the
growing number of choices, but starting to have difficulties when multiple items fit their needs.
      </p>
      <p>
        Furthermore, with continuous product diversification, user queries are also becoming even more refined, contributing
to customer satisfaction and self-satisfaction being increasingly difficult to achieve [
        <xref ref-type="bibr" rid="ref16 ref44">16, 44</xref>
        ]. To address the customers’
refined expectations, hotel reservation services provide faceted search functions, e.g., additional filters (e.g., free breakfast
or late check-out ), sets of objective options that serve as potential additional queries to reduce the choice set. Such
iflters range from being just a handful of pre-defined, static options to sometimes even thousands of carefully curated
ones over the course of several years [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, continuously updating such a set of filters as a response to product
diversification and customer expectation can be extremely costly. Moreover, navigating through a large set of filters can
even become a burden, defeating its very own purpose [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>In our work we provide a simple framework for automatically acquiring filters related to the hotels that match
the customer’s initial query, by identifying useful mentions from customer reviews. We especially focus on customer
experiences that are potentially both relevant and interesting for other customers, while also having the capability of
reducing the choice set in an intuitive and natural way.</p>
      <p>The main contributions of this paper are the following:
(1) In order to address the overchoice problem, we present a simple clustering based approach to identify useful
iflters in a dynamic manner, from customer reviews.
(2) We define key concepts and strategies in scoring and ranking filters that are meaningful and natural for the
customer.
(3) We present simple but eficient methods to implement filter scoring and ranking.
(4) We validate our proposal through a series of user experiments. We found that subjective, experience based filters
that express quality judgements were especially useful for potential users to narrow down the search space.</p>
      <p>The paper is organized as follows: in Section 2 we discuss the related work, followed by the definition of key concents
of our approach in Section 3, data description in Sections 4 and details of our proposal in Section 5. We describe our
experiments in Section 6, followed by discussions with future directions in Section 7 and the concluding remarks in
Section 8.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Automatic facet generation is a closely related field to our task. Faceted search augments traditional search by presenting a
set of attributes or filters that are grouped into facets, allowing customers to narrow down the search results [
        <xref ref-type="bibr" rid="ref10 ref22 ref37 ref8">8, 10, 22, 37</xref>
        ].
Manual curation and continuous updating of facets can be extremely costly [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], thus automatic methods to identify
and rank filters have been proposed [
        <xref ref-type="bibr" rid="ref18 ref27">18, 27</xref>
        ]. Our work difers in three main aspects from automatic facet generation.
Firstly, facet generation methods employ knowledge bases to maintain a well organized structure of facets [
        <xref ref-type="bibr" rid="ref18 ref27">18, 27</xref>
        ]. Our
method does not employ structured knowledge bases, instead, we rely only on customer reviews. Secondly, faceted
search typically targets objective filters to populate facets. Our work, besides objective filters, identifies subjective filters
as well, crucial in expressing unique experiences that might be of value for new potential customers. Thirdly, compared
to faceted search, our method puts an emphasys on addressing overchoice. The explanatory search nature of faceted
search does address overchoice, but sometimes navigation through a large set of facets becomes a burden in itself,
defeating its very own purpose [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Our method has the option of providing only a handful of potentially meaningful
and unique filters that can reduce the choice set, without putting extra burden on the customer.
      </p>
      <p>
        As another method to reduce information overload, customer review summarization is also related to our field
[
        <xref ref-type="bibr" rid="ref11 ref23 ref38 ref9">9, 11, 23, 38</xref>
        ]. Our work mainly difers from review summarization in that we attempt to identify filters that are common
across multiple items, whereas review summarization mainly focuses on identifying main characteristics of individual
items.
      </p>
      <p>
        Customer reviews have also been the target of sentiment analysis [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], aspect based opinion mining [
        <xref ref-type="bibr" rid="ref43 ref54">43, 54</xref>
        ], feature
based ranking [
        <xref ref-type="bibr" rid="ref53">53</xref>
        ]. Similarly with review summarization, these methods focus on the reviews of single items, as
opposed to identifying common, but meaningful characteristics across multiple items.
      </p>
      <p>
        Also, customer reviews can be employed to generate recommendations [
        <xref ref-type="bibr" rid="ref17 ref47">17, 47</xref>
        ]. These methods rely on customer
logs and information extracted from reviews to recommend items that are similar to previously liked ones. Our work
does not imply the existence of previous customer logs.
      </p>
      <p>
        Published work on query suggestion and recommendation has been prominently focused on the web domain
[
        <xref ref-type="bibr" rid="ref2 ref24 ref31">2, 24, 31</xref>
        ], with recent focus on e-commerce product search [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] or news related content [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Typically these works
employ knowledge bases [
        <xref ref-type="bibr" rid="ref19 ref24">19, 24</xref>
        ] or customer action logs [
        <xref ref-type="bibr" rid="ref2 ref24 ref31">2, 24, 31</xref>
        ] to suggest queries that are relevant to the original
user query. Our method difers in two key aspects. First, our target is not to suggest similar queries or filters, instead,
we attempt to provide useful filters that are not restricted to being related to the original user query. Second, we only
utilize customer reviews, without the employment of knowledge bases or customer action logs.
      </p>
      <p>
        Related to query recommendation is the field of query rewriting, the task which aims to reformulate customer
queries into well-formed ones, in order to improve customer experience [
        <xref ref-type="bibr" rid="ref50 ref52">50, 52</xref>
        ]. It difers from our work in that query
rewriting does not attempt to recommend new filters or queries to the customer.
      </p>
      <p>
        Interestingness or uniqueness discovery, key concepts in our work, is another related field, with special focus on
news articles [
        <xref ref-type="bibr" rid="ref14 ref28 ref36">14, 28, 36</xref>
        ], but definitions of uniqueness are often contain heavily domain dependent elements, such
as article freshness [
        <xref ref-type="bibr" rid="ref14 ref28">14, 28</xref>
        ] or diferences in events that occur before and after publication [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], not applicable in our
domain. A more robust method is presented in [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ], where authors define interestingness of articles as a combination
of multiple features, such as topic relevancy, source reputation, writing style or freshness. The main diference from our
work is that our target for uniqueness are simple sentences, rather than full articles.
      </p>
      <p>
        On a note, the field of anomaly detection [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] is also related to the concept of interestingness. However, unique or
interesting in our context does not go as far as being abnormal, as in Hawkins’s [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] definition of outlier 1.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>KEY CONCEPTS</title>
      <p>Our goal is to automatically identify filters that are characterized by: (1) being appealing to the customer; (2) having the
potential of addressing the overchoice problem by reducing the choice set in an intuitive and natural way. We define a
iflter “appealing” as having the quality of being both relevant and unique. Also, we define a filter set “appealing” as
being diversified, without too much emphasis on a single topic or aspect. Furthermore, to perform choice set reduction
in an intuitive way, we introduce size control policies.</p>
      <p>Relevance, uniqueness, diversity and size control are key concepts of our proposal. We employ size control policies
and diversity rules as hard constraints to identify possible filters, while using relevance and uniqueness scores to
determine the final filter ranking.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Relevance</title>
      <p>Filters are required to hold enough decision power in order to be viable expressions of user intent. Pre-defined static
iflters of conventional hotel reservation platforms are good examples of high relevance (e.g., breakfast included, late
check-out). We attempt to assign relevance scores to all possible filters. While relevance is highly subjective, we can
1“an observation that deviates so significantly from other observations as to arouse suspicion that it was generated by a diferent mechanism”
argue that all else being equal, certain filters satisfy a larger audience than others (e.g., close to the city center versus
bright pink curtains). Detailed information about relevance scoring can be found in Section 5.2.1.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Uniqueness</title>
      <p>Filters are also required to be representative of the choice set that matches the customer’s original query, capturing
characteristics that are unique within the search results. The motivation behind uniqueness is to identify options that
are especially appealing within the hotels that already match the user query (e.g., next to the city aquarium), with the
added potential to ofer choices previously unknown by the customer (e.g., private hot-spring). Section 5.2.2 ofers
detailed description on uniqueness scoring.
3.3</p>
    </sec>
    <sec id="sec-6">
      <title>Diversity 3.4</title>
    </sec>
    <sec id="sec-7">
      <title>Size control</title>
      <p>
        The importance of diversity and serendipidity is well recognized in the context of recommender systems [
        <xref ref-type="bibr" rid="ref30 ref33 ref5">5, 30, 33</xref>
        ].
Studies also point out that decision making factors are sometimes not even part of the original query [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. As a result,
we argue that, especially in cold start situations, a diversified set of filters that covers a wide range of topics is more
suitable to accommodate customer needs, than filters biased towards one or more topics. More information about our
approach in acquiring a diversified set of filters can be found in Section 5.1.
      </p>
      <p>By definition, filters are designed to address overchoice and reduce the choice set, i.e., the number of matching hotels.
We argue that the degree of the size reduction is also important. Providing highly appealing, but too generic or too
specific filters might result in a too drastic or too shallow choice set reduction, leading to customer dissatisfaction.
Instead, our strategy is to identify and provide only the filters that are guaranteed to result in a “just right” window of
matching hotels, compared to the original number of matching hotels. Naturally, this implies that our filters are based
on availability, i.e., filters that obey size control rules are guaranteed to reduce the choice set.</p>
      <p>Intuitively, in practice this should provide a natural way in reducing the choice set, balancing between relevance
and uniqueness. However, more often than not, relevance and uniqueness work against each other. Highly relevant
iflters are often not very unique (e.g., free continental breakfast), while highly unique filters may not be relevant to a
large audience (e.g., stay at a buddhist temple). When the choice set is large, arguably it is more natural to select from
more generic, thus high relevance, low uniqueness filters, with the preference shifting towards high uniqueness, low
relevance filters with a decreasing choice set. With a large choice set, size control policies rule out filters that are not
frequent enough, thus disregarding long-tail, but unique filters, with higher relevance ones gaining more prominence.
With a decreasing choice set, long-tail, unique filters should gain more exposure at the expense of more generic, relevant
iflters. More information about size control policies can be found in Section 5.1.2.
4</p>
    </sec>
    <sec id="sec-8">
      <title>HOTEL REVIEWS AS DATA SOURCE</title>
      <p>As our data source we use over 20 million sentences extracted from hotel reviews, collected from one of the largest hotel
booking sites in Japan2. The hotel review corpus contains the customer review texts and the location data associated to
each hotel.</p>
      <p>Original predicate-argument structure
very delicious food, really delicious food, all food is delicious, food is of course delicious,
delicious food as advertised, more than delicious food
hotel is close to the station, really close to the station, close to the station as mentioned,
closest to the station, pretty close to the station
extremely clean rooms, very clean rooms, rooms clean as always, rooms are of course
clean, thoroughly cleaned rooms, rooms cleaned to the last detail
4.1</p>
    </sec>
    <sec id="sec-9">
      <title>Filter units</title>
      <p>
        An underlying assumption of our method is that a user friendly filter extracted from customer reviews can be represented
by a simple predicate-argument structure. To this end, we extracted over 20 million predicate-argument structures
from our corpus by using JUMAN++ (v2.0.0-rc3), a Japanese morphological analyzer [
        <xref ref-type="bibr" rid="ref49">49</xref>
        ] and KNP++ (v0.9-21cc58c),
a Japanese dependency and case structure analyzer [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. We modified the case structure analyzer in order to retain
only the core arguments of the predicates, discarding subtle nuances (e.g., modifiers, adverbs, adjectives, adverbial or
adjective phrases, etc.) that are not relevant in the context of user friendly filters. To this end, we retained the arguments
that mark the most essential Japanese grammatical cases: nominative, accusative, dative, instrumental, and the Japanese
topic marker3. Table 1 illustrates some examples of core predicate argument structures together with their original
form before the discarding process.
      </p>
      <p>
        Some of the resulting predicate-argument structures were unrelated to hotels or had negative polarities, unsuitable
for our filter policies. As a result, we employed a filtering method based on the automatic classification results of
two BERT-based classifiers fine-tuned with an annotated corpus 4 [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] to identify non-negative predicate-argument
structures relevant to the hotel or its services. As pre-trained model we used a BERT model trained on our hotel review
corpus. For more information about our BERT model refer to Section 4.2. Finally, we retain core predicate-argument
structures whose frequency is at least 5 in our corpus. As a result of the above processes, we retained 167,886 unique
non-negative core predicate-argument structures.
4.2
      </p>
    </sec>
    <sec id="sec-10">
      <title>Filter representation</title>
      <p>
        To represent filters, we pre-trained a BERT [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] model on our review corpus. Here we followed the methodology
described in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The authors in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] employ SentencePiece [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ], an unsupervised text tokenizer which learns sentence
units for a predetermined vocabulary size. We set the vocabulary size to 32,000. To train the BERT model, we used the
parameter values oficially distributed with BERTBase. We set the batch size to 512, the number of attention heads to
12, the number of layers to 12, and the number of hidden layers to 12. We trained the BERT model for 1,500,000 steps
using TPUs.
      </p>
      <p>
        To improve on BERT’s embeddings, we employed the sentence embedding framework described in [
        <xref ref-type="bibr" rid="ref40">40</xref>
        ], using the
triplet loss function to fine-tune our pre-trained model. The triplet-loss function requires a triplet of (anchor, positive,
negative) sentences where the (anchor, positive) tuple is a positive pair, while the (anchor, negative) tuple is a negative
pair. As input triplets for fine-tuning our pre-trained model, we employed a simple tf-idf based word2vec sentence
      </p>
      <sec id="sec-10-1">
        <title>3We retained the arguments that were marked by the Japanese particles ga, wo, ni, de and ha.</title>
        <p>
          4https://github.com/megagonlabs/jrte-corpus
representation described in [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] for each filter, randomly selecting 30,000 triplets whose (anchor, positive) pair had a
cosine similarity larger than 0.85, and whose (anchor, negative) cosine similarity was smaller than 0.20.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>PROPOSED METHOD</title>
      <p>We developed machine learning based methods to identify and rank filters. Given a set of hotels that match an initial
set of original user queries, we automatically extract the non-negative core predicate-argument structures described in
Section 4.1 from the customer reviews of the matching hotels. These core predicate-argument structures will act as
potential filters. First, using the sentence embedding representations of the filters, we apply a 2 staged hierarchical
clustering method to group them into semantically similar clusters. In this step we employ policies to identify clusters
that follow size control restrictions. Next, we score each cluster for relevance and uniqueness to determine the final filter
class ranking. In this step we employ diversity policies. Also in this step we label the top ranked filter clusters. Below is
a detailed description of each step.
5.1</p>
    </sec>
    <sec id="sec-12">
      <title>2-stage clustering</title>
      <p>5.1.1 Stage 1: main topic identification. In the first stage of clustering we attempt to group filters into main latent
topics, e.g., food, location, hot spring, etc. The purpose is to serve diversity by identifying such latent topics, with
the assumption being that filters from diferent clusters after stage 1 will roughly have diferent topics 5.</p>
      <p>
        In order to identify the main topics, we employ Ward’s agglomerative clustering method [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ] with complete linkage
and cosine similarity as metric. As feature representation for the filters, we used the sentence embeddings described
in Section 4.2. Empirical results showed that a similarity threshold of 0.5 resulted in latent topic clusters with a good
trade-of between inter-clusters homogeneity and intra-cluster variance.
      </p>
      <p>We recognize that a carefully curated knowledge-driven approach may have the advantage in accurately associating
pre-defined topics to filters. However, besides cost issues, our data-driven approach has the advantage of recognizing a
potentially infinite number of intrinsic topics that would be dificult to manually acquire.</p>
      <p>Also note that the purpose of the first stage is solely to associate filters with latent topics, thus this step can be
performed beforehand, independently of user queries.
5.1.2 Stage 2: filter identification with size control. In the second clustering stage we identify filter clusters from each
main topic from Section 5.1.1 whose size obey size control rules. The size of a filter cluster is defined as the total number
of hotels that the members of the cluster are linked to. Size control is governed by two parameters,  _ and
 _ that represent the lower bound percentage and upper bound percentage, respectively, in respect to the
size of the original choice set.</p>
      <p>To achieve this, for each topic output by the main topic identification step described in Section 5.1, we parse the
hierarchical subtree of each topic by incrementally moving up in the cluster hierarchy. During this process we retain
clusters that obey size control rules and stop where the linkage drops below a certain similarity threshold, empirically
set to 0.7. Empirically, we set  _ and  _ to 30% and 70%, respectively.
5.2</p>
    </sec>
    <sec id="sec-13">
      <title>Filter scoring</title>
      <p>We score and rank filter clusters retrieved in the 2-stage clustering step based on their relevance and uniqueness. After
ranking, we apply diversity rules and label the top  filter clusters as described below.</p>
      <p>5Note that we do not attempt to label the resulting topics. Instead, we only attempt to identify filter groups that belong to the same latent topic.
5.2.1 Relevance score. In Section 3.1 we stressed the importance of discovering filters that are crucial enough in the
decision making process. Determining relevance is a non-trivial task, since people’s preferences are obviously not
uniform. A filter that may be highly relevant for one customer, may be less relevant for another one (e.g., rich choice of
baby formula or free pair ticket to the city aquarium), depending not only on personal preferences, but also on situations
or even purpose of visit.</p>
      <p>Since this is a cold start scenario, personalized relevance estimators based on customer action logs are not feasible.
Instead, we define the relevance of a filter independently from the original user query, as the average of multiple
subjective relevance scores.</p>
      <p>We pre-computed filter relevance scores using a simple k-nearest neighbor classifier. For each filter we took the top
 = 5 similar filters from our training data, and computed their average relevance, weighted by the similarity score,
as shown in the below formula, where  denotes the target filters,  denote filters of the training data, relevancegold
denotes gold relevance scores of the training data.</p>
      <p>Í
relevance( ) = =1 cossiÍm(,  ) × relevancegold ( ) (1)</p>
      <p>=1 cossim(,  )</p>
      <p>As similarity score we employed cosine similarity, computed on the sentence embeddings described in Section 4.2.
We normalized the relevance score by scaling it to between 0 and 1.</p>
      <p>
        As training data we randomly selected 8000 filters and asked 5 crowd workers to label their degree of relevance from
5 to 16. The most relevant was labeled with 5, the least relevant being 1. Ungrammatical or semantically unsound filters
were labeled with 0. We calculated pairwise inter-annotator agreement using Weighted Cohen’s kappa [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Kappa
values were between 0.24 and 0.56, representing fair to moderate agreement, underlying the highly subjective nature of
the task.
      </p>
      <p>For our classifier we used the filters that were judged as grammatically correct by at least 4 out of the 5 workers. For
the grammatically correct filters we averaged the individual worker scores. We preferred to use truncated mean (e.g.,
ignoring the lowest and highest scores of the 5 workers) in order to counter for highly subjective relevance scores (e.g.,
rich choice of baby formula). Table 2 shows an excerpt of the filters and their respective averaged relevance scores.</p>
      <p>We evaluated our relevance classifier on a held-out data of 1000 samples by calculating the precision on increasing
error ranges. We achieved a precision of 60.40% when the error range between the estimated relevance and reference
relevance was less or equal than 0.1 points, and 85.50% precision at 0.2 points error range, as shown in Figure 1.
6We manually selected the crowd workers based on their demographic information (i.e., gender, age range) to ensure diversity.</p>
      <p>We compute relevance scores for filter clusters as the weighted average relevance of its member filters as shown
in the below formula, where  denotes filter clusters, relevance( ) denotes filter relevance and freq( ) denotes the
frequency of filter  in the target choice set.
5.2.2 Uniqueness score. We define the uniqueness of a filter as the property of being important within a selected group
of hotels. We employed term frequency–inverse document frequency (tf-idf) as uniqueness of each filter, where 
denotes the filter,  denotes the reviews of a specific hotel.</p>
      <p>uniqueness(, ) = tf (, ) × idf ( ) (3)</p>
      <p>
        Intuitively, a unique characteristic of a subset is more dominant in the subset than within the entire population. The
sparse nature of the filters makes it unfeasible to handle them individually, thus for the purpose of computing tf -idf, we
employed Ward’s agglomerative clustering method [
        <xref ref-type="bibr" rid="ref51">51</xref>
        ] with complete linkage and cosine similarity as metric, with a
similarity threshold of 0.7 to group together filters of similar semantic properties. As a result, we clustered the filters
into 5178 clusters and we computed the tf-idf scores on the resulting clusters. Cluster members inherited the tf-idf
values of their parent cluster.
      </p>
      <p>Similarly to relevance scores, we computed filter cluster uniqueness score as the weighted average uniqueness of
its member filters as shown in the below formula, where  denotes filter clusters, uniqueness(, ) denotes filter
uniqueness for filter  in hotel review set , freq( ) denotes the frequency of filter  in the target choice set.
uniqueness( , ) =
Í ∈ freq( ) × uniqueness(, )
Í ∈ freq( )
5.2.3 Filter ranking and employing diversity rules. Filter cluster ranking is determined by multiplying the filter cluster’s
relevance and uniqueness scores, weighted by their respective weights (i.e.,  and  for relevance and uniqueness,
relevance( ) =
Í ∈ freq( ) × relevance( )
Í ∈ freq( )
(2)
(4)
respectively), as shown in the below formula. During preliminary empirical evaluations, we found that a reasonable
value for  and  were 1 and 2, respectively.</p>
      <p>rank( ) = ( + relevance( )) × ( + uniqueness( ))</p>
      <p>To produce the final ranking, only the top   iflter clusters are retained for each main topic described in
Section 5.1. We set   to 1, e.g., we retain only the top filter from each topic 7.
(5)
5.2.4 Labeling filters. Finally we label the top  filter clusters that will be presented for the customer. We perform
cluster labeling by choosing the most representative member, i.e., the member closest to the cluster centroid. We utilize
cosine similarity to determine the representative member that will act as the label of the filter cluster.
6</p>
    </sec>
    <sec id="sec-14">
      <title>EXPERIMENTS</title>
      <p>We conducted a number of user experiments to evaluate: (1) the top   overall output and (2) the top  
individual filter outputs of our proposed method, described in Section 5. Particularly, we compared the filters output
by our proposed method against manually acquired filters. Also, we assessed the efect of uniqueness, relevance and
diversity policies. To this end, we performed pairwise comparison against the following baseline models:
• human: a manually compiled filter list described below in Section 6.1.
• relevant: proposed without the uniqueness score, i.e., filter ranking is determined only by relevance.
• unique: proposed without relevance score, i.e., filter ranking is determined only by uniqueness.
• non-diverse: proposed without diversity policies, i.e., output is not restricted to the top   iflters for
each main topic.</p>
    </sec>
    <sec id="sec-15">
      <title>6.1 Manually compiled filters</title>
      <p>To manually acquire filters in a simulated overchoice scenario, we randomly selected 10 &lt;original request, location&gt;
query tuples with the following conditions:
• the resulting hotel hit count is at least 30 in our hotel review corpus;
• the total number of corresponding reviews8is at least 3000 in our hotel review corpus.</p>
      <p>Table 3 shows the query tuples used in this process. From the resulting reviews we randomly selected 1000 reviews
for each query tuple. Next, using the selected reviews, we asked 3 crowd workers to extract all simple short phrases
which in their opinion contain meaningful information in further filtering the choice set. Such phrases were manually
grouped by each worker into clusters that share the same meaning. Finally, all clusters were aggregated by a fourth
crowd worker, registering the number of contributing workers and the number of hotels each cluster links to. As the
ifnal output, we considered filters that had a number of majority contributors (at least 2 out of 3), ranked in descending
order by the number of corresponding hotels. Table 4 shows an example of manually acquired filters.
7One important note is that before applying diversity policies, first we remove filters that are semantically too similar to the original user queries. We
perform this by using cosine similarity on their sentence embeddings described in Section 4.2. Empirically we set this similarity threshold to 0.8.
8We performed exact text match in retrieving hotels reviews that mention an original request.
in Section 6.1. For each query, we considered the top   = 5 filters from each method’s output.</p>
      <p>We crowdsourced the pairwise evaluation, asking 300 workers using Yahoo!Japan’s crowdsourcing service9 to choose
the filter list they find more suitable in further narrowing down the choice set. In randomized order, we showed the
iflter lists of the two methods (named lists A and B, respectively), asking the workers to choose exactly one of four
choices:
• list A is more useful than list B
• list B is more useful than list A
• list A and list B are both equally useful
• neither of the lists are useful</p>
      <p>We also asked the workers to motivate their choice for each filter list pair. After basic data quality check (i.e., removing
workers that (1) did not provide any explanation for their choices, (2) always choose the same option, or (3) working
time was too short), we retained the results of 224 workers. Table 5 shows the overall list comparison results for each</p>
      <p>Against the manually generated human filters, proposed was considered to be the significantly better 10overall choice
with 5 out of the 10 evaluation query tuples. In 2 out of 10 cases the human output was considered to be significantly
superior to the output of proposed. Overall, proposed was found to have a significant advantage over human with over</p>
      <p>Analysing the workers’ comments, we observed that the overall output of proposed was overwhelmingly preferred
over human when the filters ofered very specific choices or experiences (e.g.,
the parking lot is large, thus easy to park
the car, delicious food with local ingredients). At the same time, human was preferred by workers who value filters which
proposed considered as relevant, but not unique enough to rank high (e.g., large room, clean hotel). It is also worth
mentioning that when human outperformed proposed, the number of votes counted for either of the methods was
actually smaller than in average, both lists being equally preferred or unpreferred by a large number of workers. We can
also note that proposed was voted as the better overall choice by an overwhelming majority with a number of query
tuples (e.g., Close to the sea @Shizuoka). The reason for this vote diference is that proposed managed to identify filters
that are highly specific to the initial query tuple, and at the same time are also quite appealing to potential customers
(e.g., the splendid alphonsino was very delicious), while human failed to identify such filters with a high enough frequency.</p>
      <p>Against non-diverse as well, proposed exhibited a significant vote advantage (45.4 points), validating the efect of
the diversity policies. However, non-diverse did perform better with some query tuples that are related to locations
especially recognized or famous in relation with a specific main topic, which was captured and over-represented by the
non-diverse method (e.g., nature topic in Akita).</p>
      <p>Proposed also outperformed unique and relevant by over 6.3 and 29.2 points vote count diference, respectively,
validating that both relevance and uniqueness contribute significantly to proposed. It is worth mentioning that
relevant behaved very similarly to human against proposed, suggesting that the workers employed in acquiring the
manual filters may have had preference towards more relevant, rather than unique filters.</p>
      <p>10We checked for significance using binomial test of significance with  set to 0.05.</p>
    </sec>
    <sec id="sec-16">
      <title>6.3 Individual filter evaluation</title>
      <p>In the second set of experiments we performed pairwise comparison of the individual filters output by proposed against
the outputs of the target methods11. Here we attempt to counter the tendency some workers may have had in rejecting
certain filter lists during list based evaluation described in Section 6.2, for the reason of containing unappealing filters.
We used the same filters as with list based evaluation, merging and shufling the filters into a single list. In case of
duplicates, a single occurrence was retained. We crowdsourced the evaluation asking 300 workers using Yahoo!Japan’s
crowdsourcing service. As opposed to filter list evaluation, here workers were asked to award individual filters, by
selecting from 1 up to at most 5 filters they find appealing in further reducing the choice set. We also asked the workers
to motivate their choice for each selection set.</p>
      <p>After basic data quality check (i.e., removing workers that (1) did not conform with the rule regarding the number of
selected filters, or (2) working time was too short), we retained the results of 197 workers. For each query, we counted
the total number of votes each method received. Filters that were duplicates counted for both methods. The results for
each query tuple are summed up in Table 6.</p>
      <p>Against the manually acquired human method, proposed had a significantly larger vote share with 5 out of the 10
evaluation queries, while being outperformed in only one case. We also found that proposed received more votes
for filters that express experiences or quality judgements, positive opinions regarding a specific service or the hotel
in general (e.g., the open-air hot spring was excellent, the free breakfast was delicious), while the human filters had the
tendency to have more factual filters or presence/absence indicators (e.g., open-air hot spring was available, free breakfast).
Similarly to overall filter list evaluation, human received numerous votes for filters that are highly relevant, but were
ruled out by size constraints by proposed (e.g., clean rooms being too frequent, washing machine available being too
rare).</p>
      <p>Also similarly to filter list evaluation, proposed outperformed both relevant and unique. However, it must be
noted that both relevant and unique received numerous votes with filters that are very relevant, but less unique (e.g.,
clean rooms, excellent service, free wifi ) or unique, but arguably not relevant enough (e.g., karaoke machine is available,
dog run attached to the hotel). These results suggest that while relevance and uniqueness both contribute to proposed,
their importance is highly subjective.</p>
      <p>11Here we skip pairwise comparison against non-diverse, since the target of this experiment are filters, as opposed to filter sets.
7</p>
    </sec>
    <sec id="sec-17">
      <title>DISCUSSIONS AND FUTURE DIRECTIONS</title>
      <sec id="sec-17-1">
        <title>We found that proposed was preferred by crowd-workers in two distinct scenarios.</title>
        <p>Firstly, as observed during both list based and individual filter evaluations, workers preferred highly specific filters
as opposed to more generic ones (e.g. the splendid alphonsino was very delicious versus food was delicious, the parking
lot is large, thus easy to park the car versus parking lot available). This validates our assumption that the majority of
potential customers are interested in very specific details in attempting to reach a decision.</p>
        <p>Secondly, during the individual filter evaluation, we observed the worker’s tendency in preferring filters that were
formulated as an experience or quality judgement, rather than as a fact or presence/absence indicator, when both
options were available. (e.g., hot-bath was great versus hot bath is available; food was delicious versus food available at
hotel). This tendency was weak with topics in which experience itself may not be too relevant (e.g., the experience
expressing easy to park was not overwhelmingly preferred over the factual parking lot available, arguably because the
fact that parking is actually available is the crucial piece of information, rather than the ease of parking), but it was
very prominent with topics such as food, location or other service related ones, where previous user’s experiences
and reviews are more valuable than the simple availability of that specific option.</p>
        <p>We validated this assumption with a very simple experiment. We manually selected 30 (fact, experience) filter pairs
and for each filter pair we asked 20 crowd-workers to choose the filter list they find more suitable in further narrowing
down the choice set. We randomized the order of the two filters (named A and B, respectively), asking the workers to
choose exactly one of four choices:
• filter A is more useful than filter B
• filter B is more useful than filter A
• filter A and filter B are both equally useful
• neither of the filters are useful</p>
        <p>After basic data quality check (i.e., (1) always choose the same option, or (2) working time was too short), we retained
the results of 19 workers. In 28 out of 30 pairs the experience based filter received the higher share of votes, although
both fact and experience based ones received many votes. Most of these pairs had the topic of location, food, hot
spring or some other type of hotel service. In case of a single pair the diference was only minimally in favor of the
experience based filter, namely parking as topic. One pair was voted as being equally helpful, having received only a
few votes for either fact or experience based filters, in the topic of hotel amenities (amenities are available versus
very basic amenity).</p>
        <p>This result suggests that the balance between fact and experience based filters is both subjective and possibly topic
dependent. While customer reviews mainly ofer intricate experience-like details, fact based filters still remain valuable.
As current filters provided by conventional hotel reservation systems are largely fact based ones, undoubtedly intricate
experience based filters extracted from customer reviews could add significant value. In deploying such a customer
review based filter recommender, strategies need to be implemented to combine various types of filters from multiple
sources of information. In the future we are planning to investigate how various sources of information can complement
each other in providing meaningful filters.
8</p>
      </sec>
    </sec>
    <sec id="sec-18">
      <title>CONCLUSIONS</title>
      <p>In this paper we proposed a simple clustering based approach to address the overchoice problem in the hotel industry
domain. We introduced size control and diversity policies, together with scoring verticals, in order to identify and score
iflters that could reduce the search space in a natural and intuitive way. We validated our proposal through a series of
user experiments where we also showed that the filters identified by our method were more useful than the manually
acquired ones.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Susan</given-names>
            <surname>Auty</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Consumer choice and segmentation in the restaurant industry</article-title>
          .
          <source>Service Industries Journal</source>
          <volume>12</volume>
          ,
          <issue>3</issue>
          (
          <year>1992</year>
          ),
          <fpage>324</fpage>
          -
          <lpage>339</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Ricardo</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Hurtado</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Marcelo</given-names>
            <surname>Mendoza</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Query recommendation using query logs in search engines</article-title>
          .
          <source>In International conference on extending database technology</source>
          . Springer,
          <fpage>588</fpage>
          -
          <lpage>596</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Lucas</given-names>
            <surname>Bernardi</surname>
          </string-name>
          , Pablo Estevez, Matias Eidis, and
          <string-name>
            <given-names>Eqbal</given-names>
            <surname>Osama</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Recommending Accommodation Filters with Online Learning. (</article-title>
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Juergen</given-names>
            <surname>Bross</surname>
          </string-name>
          and
          <string-name>
            <given-names>Heiko</given-names>
            <surname>Ehrig</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Automatic construction of domain and aspect specific sentiment lexicons for customer review mining</article-title>
          .
          <source>In Proceedings of the 22nd ACM international conference on Information &amp; Knowledge Management</source>
          .
          <fpage>1077</fpage>
          -
          <lpage>1086</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Castells</surname>
          </string-name>
          ,
          <string-name>
            <surname>Neil J Hurley</surname>
            , and
            <given-names>Saul</given-names>
          </string-name>
          <string-name>
            <surname>Vargas</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Novelty and diversity in recommender systems</article-title>
          .
          <source>In Recommender systems handbook</source>
          . Springer,
          <fpage>881</fpage>
          -
          <lpage>918</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Raghavendra</given-names>
            <surname>Chalapathy</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sanjay</given-names>
            <surname>Chawla</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Deep learning for anomaly detection: A survey</article-title>
          . arXiv preprint arXiv:
          <year>1901</year>
          .
          <volume>03407</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Varun</given-names>
            <surname>Chandola</surname>
          </string-name>
          , Arindam Banerjee, and
          <string-name>
            <given-names>Vipin</given-names>
            <surname>Kumar</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Anomaly detection: A survey</article-title>
          .
          <source>ACM computing surveys (CSUR) 41</source>
          ,
          <issue>3</issue>
          (
          <year>2009</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>58</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Joseph</given-names>
            <surname>Chee</surname>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , Nathan Hahn, Adam Perer, and
          <string-name>
            <given-names>Aniket</given-names>
            <surname>Kittur</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>SearchLens: composing and capturing complex user interests for exploratory search</article-title>
          .
          <source>In Proceedings of the 24th International Conference on Intelligent User Interfaces (Marina del Ray</source>
          ,
          <article-title>California) (IUI '19)</article-title>
          . ACM, New York, NY, USA,
          <fpage>498</fpage>
          -
          <lpage>509</lpage>
          . https://doi.org/10.1145/3301275.3302321
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Li</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Guanliang</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Feng</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Recommender systems based on user reviews: the state of the art</article-title>
          .
          <source>User Modeling and User-Adapted Interaction 25</source>
          ,
          <issue>2</issue>
          (
          <year>2015</year>
          ),
          <fpage>99</fpage>
          -
          <lpage>154</lpage>
          . https://doi.org/10.1007/s11257-015-9155-5
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Li</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Feng</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Explaining recommendations based on feature sentiments in product reviews</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on Intelligent User Interfaces (Limassol, Cyprus) (IUI '17)</source>
          .
          <article-title>Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <fpage>17</fpage>
          -
          <lpage>28</lpage>
          . https: //doi.org/10.1145/3025171.3025173
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Li</surname>
            <given-names>Chen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feng</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luole Qi</surname>
            , and
            <given-names>Fengfeng</given-names>
          </string-name>
          <string-name>
            <surname>Liang</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Experiment on sentiment embedded comparison interface</article-title>
          .
          <source>Knowledge-Based Systems</source>
          <volume>64</volume>
          (
          <year>2014</year>
          ),
          <fpage>44</fpage>
          -
          <lpage>58</lpage>
          . https://doi.org/10.1016/j.knosys.
          <year>2014</year>
          .
          <volume>03</volume>
          .020
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Ilia</surname>
            <given-names>Cherniavskii</given-names>
          </string-name>
          , Alexander Perelygin, and
          <string-name>
            <surname>Russell</surname>
          </string-name>
          Lee-Goldman.
          <year>2016</year>
          .
          <article-title>Suggested Keywords for Searching News-Related Content on Online Social Networks</article-title>
          .
          <source>US Patent App. 14/592</source>
          ,
          <fpage>988</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>1968</year>
          .
          <article-title>Weighted kappa: nominal scale agreement with provision for scaled disagreement or partial credit</article-title>
          .
          <source>Psychology. Bulletin</source>
          ,
          <volume>70</volume>
          ,
          <fpage>213</fpage>
          <lpage>220</lpage>
          (
          <year>1968</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Gianna M Del Corso</surname>
            , Antonio Gulli, and
            <given-names>Francesco</given-names>
          </string-name>
          <string-name>
            <surname>Romani</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Ranking a stream of news</article-title>
          .
          <source>In Proceedings of the 14th international conference on World Wide Web</source>
          .
          <fpage>97</fpage>
          -
          <lpage>106</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Jacob</surname>
            <given-names>Devlin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Kristin</given-names>
            <surname>Diehl</surname>
          </string-name>
          and
          <string-name>
            <given-names>Cait</given-names>
            <surname>Poynor</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Great expectations?! Assortment size, expectations, and satisfaction</article-title>
          .
          <source>Journal of Marketing Research</source>
          <volume>47</volume>
          ,
          <issue>2</issue>
          (
          <year>2010</year>
          ),
          <fpage>312</fpage>
          -
          <lpage>322</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Ruihai</given-names>
            <surname>Dong</surname>
          </string-name>
          and
          <string-name>
            <given-names>Barry</given-names>
            <surname>Smyth</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>From more-like-this to better-than-this: hotel recommendations from user generated reviews</article-title>
          .
          <source>In Proceedings of the 2016 Conference on User Modeling Adaptation and Personalization</source>
          .
          <volume>309</volume>
          -
          <fpage>310</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Leila</surname>
            <given-names>Feddoul</given-names>
          </string-name>
          , Sirko Schindler, and
          <string-name>
            <given-names>Frank</given-names>
            <surname>Löfler</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Automatic facet generation and selection over knowledge graphs</article-title>
          .
          <source>In International Conference on Semantic Systems</source>
          . Springer, Cham,
          <fpage>310</fpage>
          -
          <lpage>325</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Mohammad</given-names>
            <surname>Al</surname>
          </string-name>
          <string-name>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Nish</given-names>
            <surname>Parikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Gyanit</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Neel</given-names>
            <surname>Sundaresan</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Query suggestion for e-commerce sites</article-title>
          .
          <source>In Proceedings of the fourth ACM international conference on Web Search and Data Mining</source>
          .
          <fpage>765</fpage>
          -
          <lpage>774</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Douglas</surname>
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Hawkins</surname>
          </string-name>
          .
          <year>1980</year>
          .
          <article-title>Identification of outliers</article-title>
          . Vol.
          <volume>11</volume>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Yuta</given-names>
            <surname>Hayashibe</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Japanese realistic textual entailment corpus</article-title>
          .
          <source>In Proceedings of The 12th Language Resources and Evaluation Conference</source>
          .
          <volume>6827</volume>
          -
          <fpage>6834</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Marti</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Design recommendations for hierarchical faceted search interfaces</article-title>
          .
          <source>In ACM SIGIR workshop on faceted search</source>
          . Seattle, WA,
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Minqing</given-names>
            <surname>Hu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bing</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Mining and summarizing customer reviews</article-title>
          .
          <source>In Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          .
          <volume>168</volume>
          -
          <fpage>177</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Zhipeng</surname>
            <given-names>Huang</given-names>
          </string-name>
          , Bogdan Cautis, Reynold Cheng, Yudian Zheng, Nikos Mamoulis, and
          <string-name>
            <given-names>Jing</given-names>
            <surname>Yan</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Entity-based query recommendation for long-tail queries</article-title>
          .
          <source>ACM Transactions on Knowledge Discovery from Data (TKDD) 12</source>
          ,
          <issue>6</issue>
          (
          <year>2018</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Sheena</surname>
            <given-names>S</given-names>
          </string-name>
          <string-name>
            <surname>Iyengar and Mark R Lepper</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>When choice is demotivating: Can one desire too much of a good thing?</article-title>
          <source>Journal of personality and social psychology 79</source>
          ,
          <issue>6</issue>
          (
          <year>2000</year>
          ),
          <fpage>995</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Sheena</surname>
            <given-names>S Iyengar</given-names>
          </string-name>
          , Rachael E Wells, and
          <string-name>
            <given-names>Barry</given-names>
            <surname>Schwartz</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Doing better but feeling worse: Looking for the “best” job undermines satisfaction</article-title>
          .
          <source>Psychological Science</source>
          <volume>17</volume>
          ,
          <issue>2</issue>
          (
          <year>2006</year>
          ),
          <fpage>143</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Zhengbao</surname>
            <given-names>Jiang</given-names>
          </string-name>
          , Zhicheng Dou, and
          <string-name>
            <surname>Ji-Rong Wen</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Generating Query Facets Using Knowledge Bases</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>29</volume>
          ,
          <issue>2</issue>
          (
          <year>2017</year>
          ),
          <fpage>315</fpage>
          -
          <lpage>329</lpage>
          . https://doi.org/10.1109/TKDE.
          <year>2016</year>
          .2623782
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Mozhgan</surname>
            <given-names>Karimi</given-names>
          </string-name>
          , Dietmar Jannach, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Jugovac</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>News recommender systems-Survey and roads ahead</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>54</volume>
          ,
          <issue>6</issue>
          (
          <year>2018</year>
          ),
          <fpage>1203</fpage>
          -
          <lpage>1227</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Daisuke</surname>
            <given-names>Kawahara</given-names>
          </string-name>
          , Yuta Hayashibe, Hajime Morita, and
          <string-name>
            <given-names>Sadao</given-names>
            <surname>Kurohashi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Automatically acquired lexical knowledge improves Japanese joint morphological and dependency analysis</article-title>
          .
          <source>In Proceedings of the 15th International Conference on Parsing Technologies</source>
          .
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Denis</surname>
            <given-names>Kotkov</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Shuaiqiang</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jari</given-names>
            <surname>Veijalainen</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A survey of serendipity in recommender systems</article-title>
          .
          <source>Knowledge-Based Systems</source>
          <volume>111</volume>
          (
          <year>2016</year>
          ),
          <fpage>180</fpage>
          -
          <lpage>192</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Udo</surname>
            <given-names>Kruschwitz</given-names>
          </string-name>
          , Deirdre Lungley,
          <string-name>
            <surname>M-Dyaa</surname>
            <given-names>Albakour</given-names>
          </string-name>
          , and
          <string-name>
            <given-names>Dawei</given-names>
            <surname>Song</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Deriving query suggestions for site search</article-title>
          .
          <source>Journal of the American Society for Information Science and Technology</source>
          <volume>64</volume>
          ,
          <issue>10</issue>
          (
          <year>2013</year>
          ),
          <fpage>1975</fpage>
          -
          <lpage>1994</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>Taku</given-names>
            <surname>Kudo and John Richardson</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing</article-title>
          . arXiv preprint arXiv:
          <year>1808</year>
          .
          <volume>06226</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>Matevž</given-names>
            <surname>Kunaver</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tomaž</given-names>
            <surname>Požrl</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Diversity in recommender systems-A survey</article-title>
          .
          <source>Knowledge-based systems 123</source>
          (
          <year>2017</year>
          ),
          <fpage>154</fpage>
          -
          <lpage>162</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Ellen J Langer and Judith Rodin</surname>
          </string-name>
          .
          <year>1976</year>
          .
          <article-title>The efects of choice and enhanced personal responsibility for the aged: A field experiment in an institutional setting</article-title>
          .
          <source>Journal of personality and social psychology 34</source>
          ,
          <issue>2</issue>
          (
          <year>1976</year>
          ),
          <fpage>191</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Joseph</surname>
            <given-names>Lilleberg</given-names>
          </string-name>
          , Yun Zhu,
          <string-name>
            <given-names>and Yanqing</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Support vector machines and Word2vec for text classification with semantic features</article-title>
          .
          <source>In 2015 IEEE 14th International Conference on Cognitive Informatics Cognitive Computing (ICCI*CC)</source>
          .
          <volume>136</volume>
          -
          <fpage>140</fpage>
          . https://doi.org/10.1109/ICCI-CC.
          <year>2015</year>
          .7259377
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Sofus</surname>
            <given-names>A Macskassy</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Foster</given-names>
            <surname>Provost</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Intelligent information triage</article-title>
          .
          <source>In Proceedings of the 24th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          .
          <volume>318</volume>
          -
          <fpage>326</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Noemi</surname>
            <given-names>Mauro</given-names>
          </string-name>
          , Liliana Ardissono, and
          <string-name>
            <given-names>Maurizio</given-names>
            <surname>Lucenteforte</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Faceted search of heterogeneous geographic information for dynamic map projection</article-title>
          .
          <source>Information Processing &amp; Management</source>
          <volume>57</volume>
          ,
          <issue>4</issue>
          (
          <year>2020</year>
          ),
          <volume>102257</volume>
          . https://doi.org/10.1016/j.ipm.
          <year>2020</year>
          .102257
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>Samuel</given-names>
            <surname>Pecar</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Towards opinion summarization of customer reviews</article-title>
          .
          <source>In Proceedings of ACL</source>
          <year>2018</year>
          , Student Research Workshop. 1-
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Raymond</surname>
            <given-names>K Pon</given-names>
          </string-name>
          , Alfonso F Cárdenas,
          <string-name>
            <surname>David J Buttler</surname>
          </string-name>
          , and
          <string-name>
            <surname>Terence</surname>
          </string-name>
          J Critchlow.
          <year>2007</year>
          .
          <article-title>iScore: Measuring the interestingness of articles in a limited user environment</article-title>
          .
          <source>In 2007 IEEE Symposium on Computational Intelligence and Data Mining. IEEE</source>
          ,
          <fpage>354</fpage>
          -
          <lpage>361</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>Nils</given-names>
            <surname>Reimers</surname>
          </string-name>
          and
          <string-name>
            <given-names>Iryna</given-names>
            <surname>Gurevych</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Sentence-bert: Sentence embeddings using siamese bert-networks</article-title>
          .
          <source>arXiv preprint arXiv:1908</source>
          .
          <volume>10084</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>Elena</given-names>
            <surname>Reutskaja</surname>
          </string-name>
          et al.
          <year>2009</year>
          .
          <article-title>Experiments on the role of the number of alternatives in choice</article-title>
          .
          <source>Universitat Pompeu Fabra.</source>
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>Richard</surname>
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Ryan and Edward L Deci</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Self-determination theory and the facilitation of intrinsic motivation, social development, and well-being</article-title>
          .
          <source>American psychologist 55</source>
          ,
          <issue>1</issue>
          (
          <year>2000</year>
          ),
          <fpage>68</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <surname>Amani</surname>
            <given-names>K Samha</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yuefeng</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Jinglan</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Aspect-based opinion extraction from customer reviews</article-title>
          .
          <source>arXiv preprint arXiv:1404</source>
          .
          <year>1982</year>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <surname>Benjamin</surname>
            <given-names>Scheibehenne</given-names>
          </string-name>
          , Rainer Greifeneder, and
          <string-name>
            <surname>Peter M Todd</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Can there ever be too many options? A meta-analytic review of choice overload</article-title>
          .
          <source>Journal of consumer research 37</source>
          ,
          <issue>3</issue>
          (
          <year>2010</year>
          ),
          <fpage>409</fpage>
          -
          <lpage>425</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>Barry</given-names>
            <surname>Schwartz</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>The paradox of choice: Why more is less</article-title>
          . Ecco New York.
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <surname>Avni</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Shah</surname>
            and
            <given-names>George</given-names>
          </string-name>
          <string-name>
            <surname>Wolford</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Buying behavior as a function of parametric variation of number of choices</article-title>
          . Psychological Science -Cambridge- 18,
          <issue>5</issue>
          (
          <year>2007</year>
          ),
          <fpage>369</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <surname>Koji</surname>
            <given-names>Takuma</given-names>
          </string-name>
          , Junya Yamamoto, Sayaka Kamei, and
          <string-name>
            <given-names>Satoshi</given-names>
            <surname>Fujita</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A hotel recommendation system based on reviews: What do you attach importance to?</article-title>
          .
          <source>In 2016 Fourth International Symposium on Computing and Networking (CANDAR)</source>
          . IEEE,
          <fpage>710</fpage>
          -
          <lpage>712</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>Alvin</given-names>
            <surname>Tofler</surname>
          </string-name>
          .
          <year>1970</year>
          . Future shock,
          <year>1970</year>
          . Sydney.
          <string-name>
            <surname>Pan</surname>
          </string-name>
          (
          <year>1970</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          [49]
          <string-name>
            <surname>Arseny</surname>
            <given-names>Tolmachev</given-names>
          </string-name>
          , Daisuke Kawahara, and
          <string-name>
            <given-names>Sadao</given-names>
            <surname>Kurohashi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Juman++: A morphological analysis toolkit for scriptio continua</article-title>
          .
          <source>In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations</source>
          .
          <fpage>54</fpage>
          -
          <lpage>59</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          [50]
          <string-name>
            <surname>Yaxuan</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Hanqing Lu, Yunwen Xu, Rahul Goutam, Yiwei Song, and
          <string-name>
            <given-names>Bing</given-names>
            <surname>Yin</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>QUEEN: Neural Query Rewriting in E-commerce</article-title>
          . (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          [51]
          <string-name>
            <surname>Joe</surname>
            <given-names>H Ward</given-names>
          </string-name>
          <string-name>
            <surname>Jr</surname>
          </string-name>
          .
          <year>1963</year>
          .
          <article-title>Hierarchical grouping to optimize an objective function</article-title>
          .
          <source>Journal of the American statistical association 58</source>
          ,
          <issue>301</issue>
          (
          <year>1963</year>
          ),
          <fpage>236</fpage>
          -
          <lpage>244</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          [52]
          <string-name>
            <surname>Rong</surname>
            <given-names>Xiao</given-names>
          </string-name>
          , Jianhui Ji, Baoliang Cui, Haihong Tang, Wenwu Ou, Yanghua Xiao,
          <string-name>
            <given-names>Jiwei</given-names>
            <surname>Tan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xuan</given-names>
            <surname>Ju</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Weakly supervised co-training of query rewriting andsemantic matching for e-commerce</article-title>
          .
          <source>In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining</source>
          .
          <fpage>402</fpage>
          -
          <lpage>410</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          [53]
          <string-name>
            <surname>Kunpeng</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Ramanathan Narayanan, and
          <string-name>
            <surname>Alok N Choudhary</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Voice of the Customers: Mining Online Customer Reviews for Product Feature-based Ranking</article-title>
          .
          <source>WOSN</source>
          <volume>10</volume>
          (
          <year>2010</year>
          ),
          <fpage>11</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          [54]
          <string-name>
            <surname>Jingbo</surname>
            <given-names>Zhu</given-names>
          </string-name>
          , Huizhen Wang, Muhua Zhu,
          <article-title>Benjamin K Tsou,</article-title>
          and Matthew Ma.
          <year>2011</year>
          .
          <article-title>Aspect-based opinion polling from customer reviews</article-title>
          .
          <source>IEEE Transactions on afective computing 2</source>
          ,
          <issue>1</issue>
          (
          <year>2011</year>
          ),
          <fpage>37</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>