<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Extraction of Atypical Aspects from Customer Reviews: Datasets and Experiments with Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>SmitaNannawar</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ErfanAl-Hossami</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>RazvanBunescu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University of North Carolina at Charlotte</institution>
          ,
          <addr-line>Charlotte, NC 28223</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A restaurant dinner may become a memorable experience due to an unexpected aspect enjoyed by the customer, such as an origami-making station in the waiting area. If aspects that are atypical for a restaurant experience were known in advance, they could be leveraged to make recommendations that have the potential to engender serendipitous experiences, further increasing user satisfaction. Although relatively rare, whenever encountered, atypical aspects often end up being mentioned in reviews due to their memorable quality. Correspondingly, in this paper we introduce the task of detecting atypical aspects in customer reviews. To facilitate the development of extraction models, we manually annotate benchmark datasets of reviews in three domains - restaurants, hotels, and hair salons, which we use to evaluate a number of language models, ranging from ifne-tuning the instruction-based text-to-text transformer Flan-T5 to zero-shot and few-shot prompting of GPT-3.5.</p>
      </abstract>
      <kwd-group>
        <kwd>surprise</kwd>
        <kwd>serendipity</kwd>
        <kwd>customer reviews</kwd>
        <kwd>language models</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>compare, particularly under time constrai3n].tsM[aking
When looking for a restaurant or a hotel, people are ofteunre1 illustrates an example where a user Jane is looking
faced with an overwhelming number of options matchfo-r a ramen restaurant in her locality. The system knows
ing their search constraints. Even when ranked by thtehirat she has been passionate about creating crafts from
average review scores, there may be numerous high quapla-per since childhood. Upon being asked for
recommenity choices that satisfy the basic search criteria, especiadlalytions, the system finds a number of highly rated ramen
in a metropolitan area. This may lead to choice overloarde,staurants, of which Nikita Ramen stands out because
or overchoice [1, 2], where an individual is presentedit has an origami making station, an atypical aspect for a
with a large number of choices that are too dificult torestaurant, in its waiting area. The system recommends
this restaurant to Jane, importawnitthlyout telling her
observed to follow the well-known Wundt curv6e],[an
self-regulation4[], decision paralysis, and anxiety5][.
a decision in the presence of overchoice becomes men-about the origami station. Upon entering the restaurant,
tally exhausting and can lead to subsequent impairsehde is very pleasantly surprised to see the origami
making station in the waiting area, which brings feelings of
The level of satisfaction that people experience whennostalgia and happy memories from childhood. She takes
faced with an increasing number of choices has beensome time making various origami figures, before being
seated at her table. This serendipitous experience was
tion initially increases and then decreas7e,s8][. In this
inverted U-shape curve originally relating stimulus intfeanc-ilitated by the fact that an origami making station is
sity with its pleasantness. According to this functionaanl atypical aspect for a restaurant, hence it would be
dependency, as the number of choices goes up, satisface-xperienced as surprising. After the dinner event, the
system further confounds her expectations by asking her
the number of consumer choices9][ or by making one
context, choice overload can be alleviated by reducinigf she enjoyed the origami making station, which
surprises her because she did not expect that the system was
option stand out and appear better than the o1t0h]e.rsr[esponsible for the initial serendipity.
To this end, we propose that recommender systems em- To enable such recommendations with potential for
https://webpages.charlotte.edu/rbune(sRc.uBunescu)
3. The user’s data, especially in terms of their interesctast,egories: restaurants, hotels, and hair salons. In
Sectheir likes and dislikes. This would be useful for de-tion4 we describe a number of extraction approaches
termining if an atypical aspect would be enjoyed btyhat rely on language models (LM), ranging from
Flanthe user, i.e. serendipitous. T5 [11, 12] and ChatGPT 1[3] in zero-shot or few-shot
setting, to fine-tuning of Flan-T5. Experimental
evaluaTo increase the chance of serendipity, the system would</p>
      <p>tions of these models in both extractive and abstractive
need to (i) have knowledge about the user preferences,</p>
      <p>settings are detailed in Secti5o.nThe paper ends with
and (ii) also ensure that the user notices / takes advantraeglaeted work in Sectio6nand concluding remarks.
of the atypical aspect, e.g. estimating that there will be
some wait involved in the origami example.</p>
      <p>In this paper, we introduce a more focused task wher2e. Task Definition and Guidelines
we assume that the category of items requested in the
user’s query is known, e.g. restaurants, and the taGsikven a domain category, e.gr.estaurants, and a
cusis confined to using the item’s data to extract aspecttosmer review of a particular item in that categoray, e.g.
that araetypical for its category, e.g. origami station forrestaurant, the task is to extract aspects of that particular
restaurants. Because users’ expectations are shapeditbeym that araetypical of items in its category. Throughout
what they think is common for the category of interestmionst of this paper we will use the category of restaurants
their query, atypical aspects are likely to confound expeacs- an example. All aspects that are related to the core
tations, and hence be perceived as surprising. Hencefortbhu,siness of a restaurant, including but not limited to food,
we will use the termsurprising aspects to refer solely to service, price, opening hours, parking, are considered
typatypical aspects. Whenever atypical aspects are observedical aspects and are not annotated. Conversely, we define
for an item, they tend to be more noticeable and oftenand annotate an aspect as atypical if it is not related to
lead to more memorable experiences. As such, they artehe core business orfestaurants, yet it belongs to or is a
likely to be mentioned in customer reviews of that itefmea. ture ofthe restaurant (restaurants refers generically
Therefore, we use customer reviews as the source of a[n14] to the restaurant category, whetrheearsestaurant
item’s data. At this time, no user data is used as inpruetfe, rs to a specific restaurant). Correspondingly, in
Tawhich means that the atypical aspects that are extrabcltee1d,we show samples taken from 4 reviews, illustrating
while surprising for the user searching for that partticwuo- types of manual annotatioenxst,ractive
andabstraclar category of items, cannot be said for sure to leadtitvoe, analogous to the extractiv1e5,[16] and abstractive
serendipity due to unknown user preferences. In shor[1t7, , 18] annotation schemes widely used in
summarizathe task is that of extracting atypical aspects fromtciuosn- datase1t.sA special case is made of aspects related
tomer reviews, where atypical is defined to be relative ttoo the ambience or atmosphere of a restaurant: while
ama predefined item category. To the best of our knowledgeb,ience might be considered as an important part of, and
no prior work has looked into extracting atypical aspectthsus subordinated to, the core business of restaurants,
from reviews or any other types of item data. there are cases where ambience aspects stand out and</p>
      <p>The rest of the paper proceeds as follows. Secti2on become an attraction on their own. When that happens,
introduces the task of atypical aspect extraction frwoemannotate them on a secondary, optional layer.
customer reviews. Sectio3ndetails the development of In the extractive annotation, only base noun phrases
benchmark datasets of customer reviews that are
manually annotated for atypical aspects with respect to th1hrteteps://duc.nist.go,vhttps://tac.nist.gov
▶ A group of work friends and I stumbled upon Upper Deck a little over a year ago and everyone from our ofice has turned
Upper Deck into our local watering hole ever since. Their happy hour special is unbeatable, they have a good selection of draft
beers, and the food is out of this world good. The stand out feature of Upper Deck is the ofering of life size beer pong at
their outside patio. This takes traditional beer pong and substitutes solo cups with garbage cans (painted to look like solo
cups) and Ping Pong balls with dodgeballs They also have a pool table and recently added arcade games (nfl blitz 99 beats
madden 15 all day). Get some friends and bring your appetites and some quarters, you won’t be disappointed.
The restaurant ofers life size beer pong at their outside patio. They have a pool table. They recently added arcade
games, such as nfl blitz 99 and madden 15.
▶ The big draw of this place is the excellent pizza, which you can have with beer on an outside deck with a view of the
parklands. It’s a nice place to hang out on a sunny afternoon. You can even go for a walk in the Goose Creek Park behind
the restaurant afterwards to burn of the calories you just consumed. The big minus is that if you go for lunch or in early
afternoon, the menu is really limited. This place also seems to attract a goodly number of families with kids at lunch times,
probably because it serves pizza and there’s a playground in the adjoining park,
The restaurant has an outside deck with a view of the parklands. Customers can go for a walk in the Goose Creek
Park behind the restaurant. There’s a playground in the adjoining park.
▶ Classic West Philly spot where you can see local wildlife. Everyone from moms to anarchists to hackers to organic
gardeners to activists hangs out there. The cofee is excellent, the baked goods are great as well, and if you’re working on
something, you might run into a possible collaborator there. If you’re thinking of moving to West Philly, definitely check
out Satellite and the farmer’s market. They’ve replaced the cracked and chipped cups with awesome new cups, which are
awesome.</p>
      <p>In this restaurant you can see local wildlife. Everyone from moms to anarchists to hackers to organic gardeners to
activists hangs out there.
▶ This is such a cool place! Three words was all it took to add this gem to my list of places to visit while in St. Louis - ”Good
Burger Car!!” YES! They have the car from the movie Good Burger! A movie I was obsessed with as a child &amp; have since gotten
my kids to love just as much! The place is made out of cool, colorful shipping containers with many neat decorations,
what looks like an alien spaceship from Toy Story with a laser on it adorned the top of the place along with a cow. Now,
on to the food. They have many diferent options to choose from including create-your-own burgers &amp; many of their own
creations, sandwiches, salads, sides, kid’s meals, &amp; shakes &amp; floats ... Such a unique place &amp; worth a visit! They also sell
souvenir T-shirts &amp; hats, &amp; my fiance had to get himself a ”HI AF” shirt. The shirts were heavily influenced by the movie
Good Burger &amp; there was one in particular I had my heart set on, unfortunately they did not have it in my size.
The restaurant has the car from the movie Good Burger. They sell souvenir T-shirts and hats, and a customer got a
”HI AF” shirt. The shirts are heavily influenced by the movie Good Burger.</p>
      <p>The place is made out of cool, colorful shipping containers, with many neat decorations, what looks like an alien
spaceship from Toy Story with a laser on it adorned the top of the place along with a cow.
referring to atypical aspects are annotated. If multtioplaennotate only the most important part of the phrase.
phrases refer to the same atypical aspects, we only anI-n the abstractive annotation, one or more sentences
notate the grounding instance of the coreference chairne. generated that enumerate the atypical aspects
menFor example, the noun phrase ”adjoining park” in thteioned in the review. The formulation is kept as close
second review is not annotated, as it refers to the ”Goaosepossible to the original text while maintaining
natCreek Park” which is already annotated. However, ifuaralness. The generated sentences are intended to be
review mentions an atypical category, such as ”arcacdoencise, usually maintaining details that are expressed
games” in the first review or ”local wildlife” in the thirdin the same sentence in the review, however keeping out
review, any category instance that is mentioned will aulsnoimportant information about the atypical aspect that
be annotated as atypical, such as ”nfl blitz 99” in the firstis mentioned in other sentences, or details that are vague
review or ”anarchists” in the third review, respectiveolry.uncertain. Sentiment words are maintained only if it
If the atypical aspect is a more complex noun phrase, wheelps keep the text natural and faithful to the original.
only annotate the base noun phrase that expresses Tthheeabstractive annotation is meant to be standalone and
semantic core (often the syntactic head), as in ”tGhoeose used without the original review in downstream tasks,
Creek Park behind the restaurant”. The extractive annoatsas-uch it may require some minimal rewriting of the
tion is meant to be used together with the original revoierwiginal review formulation, e.g. adding the phrase ”the
in downstream applications, which makes it acceptabrleestaurant”, or removing opinion words such as ”beats” in
the first example. Sometimes reviewers use metaphors taomple, out of the 43 restaurant reviews that contain the
refer to atypical aspects, in which case it is important tlehmatma ”poncho”, in only 1 review the word ”ponchos”
the abstractive version preserves the metaphorical mewaans- deemed to refer to an atypical aspect (the
restauing. This is the case for ”local wildlife” in the third reviewr,ant was selling them). The other 42 reviews contained
which refers metaphorically to types of customers thraetferences to ponchos that were not associated with the
are seen relatively less often in that context. restaurants itself, e.g. staf helping customers put their</p>
      <p>In many aspect-based sentiment analysis approachpeosnchos on a rainy day, or customers describing their
[19], identifyingtypical aspects that are mentioned in arrival at the restaurant on a rainy day. As we went
review is done explicitly aasspect term extraction. How- down the list of rare words, their frequency increased,
ever, the task of extractiantygpical aspects, as introduced resulting in a larger number of reviews to skim through
above, cannot be solved simply by first (a) identifying for each rare word. Overall, for the Restaurant dataset,
all typical aspects of restaurants that are mentionedwine uased as search words the rare words that appeared
review, followed by (b) extracting all other noun phrasews,ith a frequency of up to 187. Upon semi-automatically
i.e. phrases that do not refer to a typical aspect osfiftiang through the ∼97K reviews found to contain these
restaurant. In their reviews, people often mention entwi-ords, we were able to collect 114 reviews that contained
ties or events that are not associated with the restauartaynptic,al aspects. On average, one hour of following this
such as ”our ofice” in the first review, or ”the farmer’s process led to finding between 2 and 3 reviews
containmarket” in the third review, and these phrases shouinldg atypical aspects for the restaurant category, whereas
not be extracted either. Thus, it is important thatfotrhtehe hair salon category it took on average two hours
noun phrase refers to an aspect thaatssiosciated with the to find 1 atypical reviews. Henceforth the teramtypical
reviewed restaurant, and that at the same time is atypicraelview will be used to refer to a review that contains one
of restaurants (or unexpected for a restaurant). Finoarllmy,ore atypical aspects; analogously, the tteyrpmical
while Table 1 may induce the perception that atypicarelview will be used to refer to reviews that do not contain
aspects are common, the opposite is actually true. As wailnl y atypical aspect.
be detailed in Section3 below, it takes going through at As illustrated in the examples from Sect2io,nwe
orleast 50 reviews in order to find one review that mentiongsanize annotations of atypical aspects on two layers:
an atypical aspect. The dificulty of manually finding this • A primary layer that contains atypical aspects that
”needle in a haystack” further motivates the developmentare clearly not connected to any core feature of that
of automated approaches for surprising aspect extractiondo.main.</p>
    </sec>
    <sec id="sec-2">
      <title>3. Manually Annotated Datasets</title>
      <p>• A secondary layer that contains atypical aspects that
are related to a typical aspect, such as ambiance or
location, but that stand out and are interesting on their
own, separate from the core features of the domain.</p>
      <p>We used the Yelp dataset20[] as a source of reviews for
the 3 target categories: Restaur∼a5nMt (reviews),
Hotel (∼190K reviews), and Hair Salon∼(115K reviews). For example,’I was even encouraged to visit their petting
Because most aspects are expressed as nouns and lezssoo in the back’ would be considered a primary atypical
frequently as verbs, we use spaC[y21] to collect lemmas aspect in any of the 3 categories, where’Tahsere is an
of all nouns and verbs and compute their frequencies foirnteresting giant stufed spider that goes up and down
each domain. We rank words in ascending order basewdhen the door leading to the bathrooms opens and closes’
on their counts and filter out words that appear with vewroyuld be annotated as a secondary atypical aspect.
low frequency, e.g., less than 10 times for the RestaurantTable 2 shows summary statistics for the 3 datasets,
domain, as these tend to be spelling mistakes or interjoecn-e for each domain (category), split between data used
tions that are purposely misspelled for extra emphasfoisr, training and testing, and data used for development.
e.g., ”amaazzing”. We then consider the remaining rarUender the Primary column, we show the number of
atypwords in ascending order of their frequency as candidaticeal reviews and atypical aspect annotations contained
in them. The next column shows the same statistics for
atypical words, extract the reviews that mention them,
and read these reviews to determine which occurrencweshen both primary and secondary atypical aspects are
truly refer to an atypical aspect. When reading a reviceown,sidered. The total number of reviews in each dataset,
all atypical aspects are annotated, not only the ones schoorw-n in the third column, is about double the number
responding to the search word. Notwithstanding tohfeprimary atypical reviews, reflecting a balanced dataset
heuristic selection of reviews based on the occurrence wofhere the number of typical reviews was selected to be
rare words, overall this was still a very time-consuminagbout the same as the number of atypical reviews.
process, because rare words very often appear in a review We computed inter-annotator agreement (ITA) on both
without necessarily referring to atypical aspects. Fortehxe- extractive and abstractive annotations in the
development sets of the Restaurant and Hair Salon domains.</p>
    </sec>
    <sec id="sec-3">
      <title>4. LM-based Approaches</title>
      <p>Based on the following restaurant review, list aspects
that are atypical for a restaurant. Separate them using
commas. context: {{Review}}
This section describes our Language Model (LM) based▶ Fine-tuning FLAN-T5 Abstractive Prompt:
quesapproaches to detecting surprising aspects in customer</p>
      <p>tion: Based on the following restaurant review, what are
reviews. We experimented with 2 language models: the atypical aspects for a restaurant? context: {{Review}}
Flan-T5 and ChatGPT. The 3 billion parameter
FLANT5 is an encoder-decoder transformer based on the T5 In the 0-shot setup for ChatGPT, we include an
instrucmodel [22] that was further instruction-tuned on thioen to either extract lists of atypical aspects (extractive)
FLAN dataset1[1, 12]. We decided to use the FLAN-T5 or to generate naturally sounding text about the atypical
model due to its exposure to narratives in the styleaospfects in the review (abstractive):
reviews, e.g., blog posts, during pre-training on the C▶4 0-shot ChatGPT Extractive Prompt: Given the
folcorpus [22], and also due to its instruction-tuning onlowing restaurant review, can you list atypical aspects for
summarization and sentiment analysis tasks. We alsoa restaurant? Atypical aspects are not related to service,
experimented with zero-shot and few-shot promptingfood, drinks, location, price, menu, discounts, policies,
of the much larger ChatGPTgp(t-3.5) [13] in order to staf, customer satisfaction, or other items commonly
evaluate the performance of a state-of-the-art languagaessociated with a restaurant. Please be precise in your
remodel without any fine tuning. sponse; it should contain only atypical aspects associated</p>
      <p>LMs are known to be sensitive to word choice2s3][. with the restaurant that is reviewed. Extract base noun
In the case of zero-shot/few-shot experiments, we triedphrases in the output format: ’Atypical aspects: aspect 1,
between 10 and 20 diferent prompts for each type of aspect 2, aspect 3.’ Output ⟨None⟨ if there are no atypical
annotation and selected the ones that performed the besatspects. Please follow the output format strictly.
on the development data. We noticed that, given thePassage: {{Review}}
description of the task along with examples of typical
aspects in the prompt helped model to identify the multip▶le0-shot ChatGPT Abstractive Prompt: Which
asatypical aspects present in a review. We also observedpects mentioned in the review are atypical for a
restauthat the more examples of typical aspects were given inrant? Unlike common aspects such as service, food,
the prompt the easier it was for the model to identifydrinks, location, price, menu, discounts, policies, staf, or
atypical aspects. We also had to try multiple prompt forc-ustomer satisfaction, atypical aspects are not commonly
mulations to instruct the LM to align with an expecteadssociated with a restaurant. In the output, formulate
output format. For 5-shot experiments, various sets of 5each aspect as sentences, e.g., ”Atypical aspects: – The
examples from the development set were first selected restaurant has ⟨aspect 1⟩. – The restaurant has ⟨aspect
at random, then manually checked for diversity in terms2⟩. – The restaurant has ⟨aspect 3⟩.” If there are no
atypof dificulty, number, and types of atypical aspects. Of ical aspects, output ”None”.
these, we selected the 5-shot examples that yielded thePassage: {{Review}}
highest performance on the development set. In the few-shot setup for ChatGPT, we include in the</p>
      <p>With the exception of abstractive generation for Hotperlso,mpt both the instruction and 5 worked-out examples:
▶ Few-shot ChatGPT Extractive and Abstractive</p>
      <p>Prompt: Given the following restaurant review, can you
list atypical aspects for a restaurant? Atypical aspects
are not related to service, food, drinks, location, price,
menu, discounts, policies, staf, customer satisfaction or
other types of items that are commonly associated with
a restaurant. Please be precise in your response, which
should contain only atypical aspects that are associated
with the restaurant that is reviewed. Output &lt;None&gt; if
there are no atypical aspects.</p>
      <sec id="sec-3-1">
        <title>Input : Gold Phrases , Extracted Phras es</title>
        <p>Output : Precision (P), Recall (R), an d 1
1 //   = # True Positives w.r.t 
2 //   = # True Positives w.r.t 
3 //   = # False Positives
4 //   = # False Negatives
5 for gp in GP do
6 if EP is empty then
7   ←   + ||/||
Example 1: {{Example Review 1}} Atypical aspects: 8 else
{{comma-separated extractive annotations OR
bull9etlisted abstractive sentences}E}x...amples 2, 3, 4, 5</p>
      </sec>
      <sec id="sec-3-2">
        <title>Algorithm 1: PartialMatchMetrics(,  )</title>
      </sec>
      <sec id="sec-3-3">
        <title>Find  ∈  that has maximum Jaccard similarity wit h</title>
        <p>Can you try for the review below? {{Review}} 10    ←    + | ∩ |/||</p>
        <p>11    ←    + | ∩ |/||</p>
        <p>We use the Hugging Face Transformers packag2e4][ 12   ←   + | − |/||
for fine-tuning Flan-T5 with the following hyper-13   ←   + | − |/||
parameters: an efective batch size of 32, a number o1f4 Remove  from the se t
epochs of 30, a learning rate of 3e-5 for Restaurants and</p>
        <p>15 for ep in EP do
5e-5 for Hotels and Hair Salons, a weight decay of 0.00116,   ←   + ||/||
and a generation max length set to 512. Those hype17r- ←   /(   +   ) ,  ←   /(   +   )
parameter values were found through tuning on the
development portion of each dataset. We perform the fin1e8- return  , ,  1 ← 2 /( + )
tuning experiments on a high-performance computing
cluster using8 CPU cores,128 GB RAM, and 2 A100 80
GB GPUs, for around 96 hours. We use the OpenAI API For the abstractive evaluation, we follow prior work
Python package 2[5] for ChatGPT, where we do greedy in summarization2[6, 27, 28] and compare the generated
decoding by setting the temperature parameter to 0o.utput with the ground truth using BERT F1 Sc2o9r]e [
instantiated with DeBERT3a0][, Rouge-1, Rouge-2, and</p>
        <p>Rouge-L-Sum [31]. In the case of typical reviews, the
5. Experimental Evaluations prediction, i.e. the generated text, should be empty,
relfecting no atypical aspect. When using RougeLsum and
The LM-based approaches are evaluated in a 10-fold
sce</p>
        <p>BERTScore to evaluate the model output on a typical
nario where the Train+Test reviews dataset is partitiorneevdiew, 1 is calculated as 1.0 if the model generates an
into 10 folds, 9 folds are used for training and 1 fold is</p>
        <p>empty text, and 0.0 if it generates a non-empty text.
used for testing. This process is repeated 10 times untilThe overall experimental results are shown in Ta3b.le
each fold in the dataset is used as a test fold. The metrFicosr Restaurants, we show results on extracting primary
computed across the 10 folds are then micro-averaged</p>
        <p>atypical aspects as well as results on extracting both
priyielding the final evaluation metric. For the extractivmeary and secondary atypical aspects. Since fine-tuned
evaluation, we report the precision, recall,  a1nsdcores Flan-T5 and ChatGPT (5-shot) obtained the best results
for theexact andpartial matches of the extracted baseon Restaurants, they were selected to be evaluated on the
noun phrase (BNP) with the ground truth phrase. other two domains, using solely primary atypical aspects.</p>
        <p>In the exact match method, an extracted phrase is</p>
        <p>Fine-tuning FLAN-T5 yields the best performance in the
considered correct if and only if it matches exactly a</p>
        <p>extractive task across all domains. While we observe a
ground truth (gold) phrase. Precision (P) and Recall (R)</p>
        <p>big performance gap between ChatGPT and fine-tuned
are then computed as follows: FLAN-T5 in the extractive setting, that gap shrinks
con = # correct extracted BNPs / #extracted BNPs siderably in the abstractive setting for Hotels and Hair
Salons, where ChatGPT (5-shot) occasionally outperforms
 = # correct extracted BNPs / #gold BNPs FLAN-T5 on some of the metrics. For both LMs, the
HoIn the partial match method, we use the greedy methtoedl domain appears to be more challenging. Compared
shown in Algorithm1 to compute a bipartite matchintgo the other domains, atypical aspects are more
combetween gold phrases and extracted phrases that mon and more diverse in hotels, likely because hotels try
aimed at maximizing their word overl|a∩p| in total. to diferentiate themselves from other hotels more than
The overlaps between extracted and gold phrases arreestaurants or hair salons do. Ta2bslehows that indeed
then used to compute precision, recall, a n1d. there are more primary and secondary atypical aspects
per review in the hotel domain. atypical aspects; (F) Typical reviews for which the model</p>
        <p>To determine how well Flan-T5 generalizes to unseencorrectly generates an empty string. The first 4 types are
atypical aspects in the Restaurant domain, we manuailllluystrated on the example below, where A = 2, B = 1, C =
created groupings of atypical aspects where semantica1ll,yD = 1, and E = 1:
similar atypical aspects, e.g. greeting cards and
anniversary gifts, are grouped together, such that aspects i•n Gold = [”On the weekends kids can experience ’Mark
diferent groups are semantically very diferent. We then the Balloon Guy’ at the restaurant. They have stufed
partition the set of groups into 10 folds of groups, which animal/puppets for sale at the froBn”,t”.The restaurant
ensures that the atypical aspects that the language modheals therapeutic sketching at every table.”]
sees in the test fold have not been seen during training •(eiO-utput = [”The restaurant has a balloon artist on
ther literally or semantically similar). Upon fine-tuning weekends.A The restaurant has traditional wooden
and evaluating FLAN-T5 on this dataset, we observe abaseball stadium seats for waitiEn”g,”.The restaurant
similar precision as reported in Tab3l,ehowever, there is ofers therapeutic sketching at every taAbleIf. your
a significant drop in recall from 60.2 to 46.1 for primary sketch is good, it will be on the waDll.The restaurant
atypical aspects and from 58.6 to 49.3 when extractinghas limited menu options before 4:30 pmC”.]
both primary and secondary atypical aspects. Improving
generalization to semantically novel atypical aspectsThisese counts are then used to compute a lenient precision
therefore an interesting avenue for future work. as  = ( +  )/( +  + ) and a strict precision as</p>
        <p>To get a sense of the real performance in the abstract iv=e ( +  )/( +  +  +  + ) , whereas recall is
setting for primary aspects, we also performed a manucaomlputed as = ( +  )/( +  + ) . Correspondingly,
evaluation of the fine-tuned Flan-T5 and 5-shot ChatGPTthe fine-tuned Flan-T5 obtains a lenient and stri1cotf
outputs. To enable calculation of precision and recall, w7e7.5% and 76.5%, respectively, whereas ChatGPT obtains
manually label and count the following: (A) Atypical asa- lenient and strict1 of 71.3% and 65.9%, respectively.
pects from the model output that are semantically simLiloaorking at Table3, the strict1 is in between RLS and
to gold aspects; (B) Gold aspects that are missing froBmERT, further supporting the use of these automated
the model output; (C) Typical aspects or other typesmofetrics. Overall, manual evaluation shows that models
entities that should not have been extracted as atyppicearlf;orm better in the abstractive vs. extractive setting.
(D) Extra details about atypical aspects; (E) SecondaryError analysis reveals that fine-tuned Flan-T5 is more
succinct in its answers, leading it to sometimes ignore
atypical aspects in its response. Conversely, ChatGPaTspects 4[6]. Qiu et al.[47] only extract aspects that have
tends to be more verbose, often generating unnecessaryan expressed opinion, using expanded opinion lexicons
details about the atypical aspects that it extracts, orcomnitsa-ining adjectives. As detailed at the end of Sec2t,ion
taking typical for atypical aspects. For instance, ChatGePxTtracting atypical aspects cannot be addressed simply
extracted a food-related aspect in the second sentenceasinthe logical complement of ABSA.
”The restaurant has an annual Valentine’s Date Night with Prior work investigated using Language Models (LMs)
table service, a caricature artist, and a piano player. The in recommender systems. Zhang et a[4l8.] investigate
restaurant ofers in-house made chocolate-covered strawber- using LMs as a recommender system by formulating a
ries.”. ChatGPT also tends to extract one-time custommeorvie recommendation task as a multi-token cloze task
experiences which are not atypical aspects of a restaun-d find that LMs underperform traditional recommender
rant, as shown in the second sentence” Tinhe restaurant systems such as GRU4Rec in both zero-shot and
finehas a sister restaurant next door with a lively fun band. tuned settings. LMs have also been used to elicit user
The server, Peter, gave very helpful suggestions.”. On the preferences in natural language given historical
interacother hand, fine-tuned FLAN-T5 mistakenly considerstions and prior selections for more personalized
recom”farmer’s market” to be an atypical aspect in the thirmdendations4[9]. Such a model can benefit a complete
imreview from Table1. Also, when the sole atypical aspectplementation of the system illustrated in Fig1,uwrehere
of a business is in its proximity, the fine-tuned FLAN-T5it would automatically infer Jane’s interests.
Furthertends to not extract it, as in the model failing to genermaotree, LMs have been used for explaining why a specific
the ground truth abstractive sente”nTcheesy are close to recommendation was made to the use5r0][.
a dock area where customers can board paddle cruises.” , Atypical aspects are valuable due to their potential to
”Customers can catch an airboat ride down the road.” create serendipity, i.e. experiences that are both
unexpected (surprising) and relevant (positive) for the user.</p>
        <p>Unlike Kotkov et al[.40], we do not consider novelty to
6. Related Work be a required dimension of serendipity. Finally, surprise
can appear from other sources, be it the experience of
Addressing overchoice is a core focus of recommender</p>
        <p>a typical aspect that stands out (such as a unique dish
systems and it is typically addressed by recommending</p>
        <p>ofered by a restaurant), or the accidental discovery of a
between 5 to 20 attractive and diverse ite3m2]s, b[ased restaurant as shown in the first sentence in Ta1balned
on user preferences3[3, 34], user ratings, item attributes,</p>
        <p>studied in 5[1]. All of these diferent sources of surprise
or user reviews3[5, 36, 37]. Recommending items with have great potential for setting up serendipity.
potential for serendipity is one way of diversifying an
item set. In3[8], unexpectedness is defined as the
distance of an item from a set of obvious items for th7at. Conclusion and Future Work
user, relative to the user’s preferred level of
unexpectedness. Li et al.[39] recommended unexpected itemsWe introduced the new task of extracting atypical
asby modeling user interests as clusters of historical dapteacts from customer reviews. To enable training and
in a latent space and calculating the weighted distanevcaeluation of atypical aspect extraction models, we
manbetween a new item and the clusters of interests. Kotkuoavlly annotated two layers of atypical aspects in customer
et al.[40] crowd-sourced serendipity labels for a moviereviews from three domains. While experimental
evaludataset using multiple definitions of serendipity. ations using few-shot prompting of ChatGPT and
fine</p>
        <p>User reviews have been used to learn latent featutreusning of Flan-T5 show promising results, there is still a
of users [37], extract sentimen3t6][, derive user prefer- substantial gap relative to human performance, as shown
ences [34], or to perform aspect-based sentiment anablyy-the higher ITA. Future work includes enhancing
reprosis in order to recommend better quality products wdituhcibility by using open-source LMs5[2], prototyping
aspects relevant to the use3r5].[ Conversational rec-a recommender system that leverages atypical aspects,
ommender systems use reviews to provide explanationasnd a user study verifying their utility.
[41], to maintain fluency in conversation42[], or to un- To facilitate reproducibility and future progress, we
derstand the user’s requirements by asking questiomnaske the code and the datasets publicly availablhetattps:
about aspects mentioned in review4s3][. //github.com/smitanannaware/Xtr A.tA</p>
        <p>Aspect-based sentiment analysis (ABSA) is a
technique that aims to extract topic-specific aspects
mentioned in a review, together with any associated senAtic- knowledgments
ment. There are diferent approaches to extracting aspect</p>
        <p>This research was partly supported by the United States
terms, with some focusing on using term frequency in the
corpus [44, 45] or using semantic clustering of prominenAtir Force (USAF) under Contract No. FA8750-21-C-0075.
tics, Seattle, United States, 2022, pp. 2300–2344. New York, NY, USA, 2010, p. 63–70. URL: https:
URL: https://aclanthology.org/2022.naacl-main..167 //doi.org/10.1145/1864708.1864724. doi:10.1145/
doi:10.18653/v1/2022.naacl-main.167. 1864708.1864724.
[24] T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. De- [33] N. Jakob, S. H. Weber, M. C. Müller, I. Gurevych,
langue, A. Moi, P. Cistac, T. Rault, R. Louf, M. Fun- Beyond the stars: Exploiting free-text user
retowicz, et al., Transformers: State-of-the-art nat- views to improve the accuracy of movie
recural language processing, in: Proceedings of the ommendations, in: Proceedings of the 1st
In2020 Conference on Empirical Methods in Natu- ternational CIKM Workshop on Topic-Sentiment
ral Language Processing: System Demonstrations, Analysis for Mass Opinion, TSA ’09,
Associa2020, pp. 38–45. tion for Computing Machinery, New York, NY,
[25] OpenAI, Openai python library, 2022. URLh:ttps: USA, 2009, p. 57–64. URL: https://doi.org/10.1145/
//github.com/openai/openai-pytho.n 1651461.1651473. doi:10.1145/1651461.1651473.
[26] R. Tangsali, A. J. Vyawahare, A. V. Mandke, O. R.[34] Y. Zhang, G. Lai, M. Zhang, Y. Zhang, Y. Liu,
Litake, D. D. Kadam, Abstractive approaches to S. Ma, Explicit factor models for explainable
recommultidocument summarization of medical literature mendation based on phrase-level sentiment
analreviews, in: Proceedings of the Third Workshop ysis, in: Proceedings of the 37th International
on Scholarly Document Processing, Association for ACM SIGIR Conference on Research &amp;
DevelopComputational Linguistics, Gyeongju, Republic of ment in Information Retrieval, SIGIR ’14,
AssociKorea, 2022, pp. 199–203. URL:https://aclanthology. ation for Computing Machinery, New York, NY,
org/2022.sdp-1.24. USA, 2014, p. 83–92. URL: https://doi.org/10.1145/
[27] O. Ahuja, J. Xu, A. Gupta, K. Horecka, G. Dur- 2600428.2609579. doi:10.1145/2600428.2609579.
rett, ASPECTNEWS: Aspect-oriented summa[-35] C. Musto, M. de Gemmis, G. Semeraro, P. Lops,
rization of news documents, in: Proceedings A multi-criteria recommender system exploiting
of the 60th Annual Meeting of the Association aspect-based sentiment analysis of users’ reviews,
for Computational Linguistics (Volume 1: Long in: Proceedings of the Eleventh ACM Conference
Papers), Association for Computational Linguis- on Recommender Systems, RecSys ’17,
Associatics, Dublin, Ireland, 2022, pp. 6494–6506. URL: tion for Computing Machinery, New York, NY,
https://aclanthology.org/2022.acl-long.4.4d9oi:10. USA, 2017, p. 321–325. URL: https://doi.org/10.1145/
18653/v1/2022.acl-long.449. 3109859.3109905. doi:10.1145/3109859.3109905.
[28] K. Mrini, C. Liu, M. Dreyer, Rewards with nega[-36] S. Selmene, Z. Kodia, Recommender system
tive examples for reinforced topic-focused abstrac- based on user’s tweets sentiment analysis, in:
tive summarization, in: Proceedings of the Third Proceedings of the 4th International Conference
Workshop on New Frontiers in Summarization, As- on E-Commerce, E-Business and E-Government,
sociation for Computational Linguistics, Online ICEEG ’20, Association for Computing Machinery,
and in Dominican Republic, 2021, pp. 33–38. URL: New York, NY, USA, 2020, p. 96–102. URL: https:
https://aclanthology.org/2021.newsum-.1d.4oi:10. //doi.org/10.1145/3409929.3414744. doi:10.1145/
18653/v1/2021.newsum-1.4. 3409929.3414744.
[29] T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, [37] P. Li, A. Tuzhilin, Learning latent multi-criteria
ratY. Artzi, Bertscore: Evaluating text generation ings from user reviews for recommendations, IEEE
with bert, in: International Conference on Learn- Transactions on Knowledge and Data
Engineering Representations, 2020. URLh:ttps://openreview. ing 34 (2022) 3854–3866. doi:10.1109/TKDE.2020.
net/forum?id=SkeHuCVFD.r 3030623.
[30] P. He, X. Liu, J. Gao, W. Chen, Deberta: Decoding-[38] A. Panagiotis, T. Alexander, On unexpectedness in
enhanced bert with disentangled attention, arXiv recommender systems: Or how to expect the
unpreprint arXiv:2006.03654 (2020). expected, in: Workshop on Novelty and Diversity
[31] C.-Y. Lin, ROUGE: A package for automatic eval- in Recommender Systems, volume 2011, 2011. URL:
uation of summaries, in: Text Summarization https://ceur-ws.org/Vol-816/paper2.p.df
Branches Out, Association for Computational Li[n3-9] P. Li, M. Que, Z. Jiang, Y. HU, A. Tuzhilin,
guistics, Barcelona, Spain, 2004, pp. 74–81. URL: Purs: Personalized unexpected recommender
syshttps://aclanthology.org/W04-10.13 tem for improving user satisfaction, in:
Pro[32] D. Bollen, B. P. Knijnenburg, M. C. Willemsen, ceedings of the 14th ACM Conference on
RecM. Graus, Understanding choice overload in rec- ommender Systems, RecSys ’20, Association for
ommender systems, in: Proceedings of the Fourth Computing Machinery, New York, NY, USA, 2020,
ACM Conference on Recommender Systems, Rec- p. 279–288. URL: https://doi.org/10.1145/3383313.</p>
        <p>Sys ’10, Association for Computing Machinery, 3412238. doi:10.1145/3383313.3412238.
[40] D. Kotkov, J. A. Konstan, Q. Zhao, J. Veijalainen, In- https://aclanthology.org/D18-13.8d4oi:10.18653/
vestigating serendipity in recommender systems v1/D18-1384.
based on real user feedback, in: Proceedings[47] G. Qiu, B. Liu, J. Bu, C. Chen, Opinion Word
of the 33rd Annual ACM Symposium on Ap- Expansion and Target Extraction through
plied Computing, SAC ’18, Association for Com- Double Propagation, Computational
Linguisputing Machinery, New York, NY, USA, 2018, p. tics 37 (2011) 9–27. URL: https://doi.org/10.
1341–1350. URL: https://doi.org/10.1145/3167132. 1162/coli_a_00034. doi:10.1162/coli_a_00034.
3167276. doi:10.1145/3167132.3167276.
arXiv:https://direct.mit.edu/coli/article[41] C. Musto, P. Lops, M. de Gemmis, G. Semeraro, pdf/37/1/9/1810309/coli_a_00034.pdf.</p>
        <p>Justifying recommendations through aspect-base[d48] Y. Zhang, H. Ding, Z. Shui, Y. Ma, J. Zou, A.
Deosentiment analysis of users reviews, in: Proceed- ras, H. Wang, Language models as recommender
ings of the 27th ACM Conference on User Modeling, systems: Evaluations and limitations (2021).
Adaptation and Personalization, UMAP ’19, Asso[-49] Z. Chen, Palr: Personalization aware llms for
recomciation for Computing Machinery, New York, NY, mendation, arXiv preprint arXiv:2305.07622 (2023).
USA, 2019, p. 4–12. URL: https://doi.org/10.1145/ [50] Y. Gao, T. Sheng, Y. Xiang, Y. Xiong, H. Wang,
3320435.3320457. doi:10.1145/3320435.3320457. J. Zhang, Chat-rec: Towards interactive and
ex[42] Y. Lu, J. Bao, Y. Song, Z. Ma, S. Cui, Y. Wu, X. He, plainable llms-augmented recommender system,
RevCore: Review-augmented conversational rec- arXiv preprint arXiv:2303.14524 (2023).
ommendation, in: Findings of the Association for[51] Z. Fu, X. Niu, L. Yu, Wisdom of crowds and
fineComputational Linguistics: ACL-IJCNLP 2021, As- grained learning for serendipity recommendations,
sociation for Computational Linguistics, Online, in: Proceedings of the 46th International ACM
SI2021, pp. 1161–1173. URL: https://aclanthology. GIR Conference on Research and Development in
org/2021.findings-acl.99. doi:10.18653/v1/2021. Information Retrieval, SIGIR ’23, Association for
findings-acl.99. Computing Machinery, New York, NY, USA, 2023,
[43] Y. Zhang, X. Chen, Q. Ai, L. Yang, W. B. Croft, To- p. 739–748. URL: https://doi.org/10.1145/3539618.
wards conversational search and recommendation: 3591787. doi:10.1145/3539618.3591787.
System ask, user respond, in: Proceedings of the[52] A. Liesenfeld, A. Lopez, M. Dingemanse, Opening
27th ACM International Conference on Informa- up ChatGPT: Tracking openness, transparency, and
tion and Knowledge Management, CIKM ’18, Asso- accountability in instruction-tuned text generators,
ciation for Computing Machinery, New York, NY, in: Proceedings of the 5th International Conference
USA, 2018, p. 177–186. URL: https://doi.org/10.1145/ on Conversational User Interfaces, CUI ’23,
Associa3269206.3271776. doi:10.1145/3269206.3271776. tion for Computing Machinery, New York, NY, USA,
[44] K. Bauman, B. Liu, A. Tuzhilin, Aspect based 2023. URL: https://doi.org/10.1145/3571884.360431.6
recommendations: Recommending items with the doi:10.1145/3571884.3604316.
most valuable aspects based on user reviews, in:
Proceedings of the 23rd ACM SIGKDD
International Conference on Knowledge Discovery and
Data Mining, KDD ’17, Association for
Computing Machinery, New York, NY, USA, 2017,
p. 717–725. URL: https://doi.org/10.1145/3097983.</p>
        <p>3098170. doi:10.1145/3097983.3098170.
[45] L. S. d. Souza, M. G. Manzato, Aspect-based
summarization: An approach with diferent levels
of details to explain recommendations, in:
Proceedings of the Brazilian Symposium on
Multimedia and the Web, WebMedia ’22, Association for
Computing Machinery, New York, NY, USA, 2022,
p. 202–210. URL: https://doi.org/10.1145/3539637.</p>
        <p>3557002. doi:10.1145/3539637.3557002.
[46] Z. Luo, S. Huang, F. F. Xu, B. Y. Lin, H. Shi, K. Zhu,</p>
        <p>ExtRA: Extracting prominent review aspects from
customer feedback, in: Proceedings of the 2018
Conference on Empirical Methods in Natural Language
Processing, Association for Computational
Linguistics, Brussels, Belgium, 2018, pp. 3477–3486. URL:</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>