<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Generating Personalised and Opinionated Review Summaries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Khalil Muhammad</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aonghus Lawlor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rachael Rafter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barry Smyth</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Insight Centre for Data Analytics, University College Dublin Bel ed</institution>
          ,
          <addr-line>Dublin 4</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <fpage>2</fpage>
      <lpage>5</lpage>
      <abstract>
        <p>This paper describes a novel approach for summarising usergenerated reviews for the purpose of explaining recommendations. We demonstrate our approach using TripAdvisor reviews.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Product reviews, that are written by real users, are now mainstream online. Sites
like Amazon and TripAdvisor have collected thousands of reviews for all manner
of products, and users are increasingly relying on these reviews make better
choices [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, there are so many reviews, some of which are quite long,
and it is increasingly di cult for users to identify the relevant information for
their needs [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Recently researchers have begun to explore the potential of such
reviews in building recommender systems [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], identifying useful reviews [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and
using related techniques to help users write better reviews in the rst place [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In this paper we explore how to exploit textual reviews to summarise product
experiences that can explain recommendations. In particular we describe how we
can pro le a user based on the reviews that they have written, and identify
product features that matter to them. And we explain how we can represent products
in a similar fashion, by extracting opinions and sentiment information from their
reviews. We then describe an initial approach to generating personalised product
summaries for given user-product pairs on TripAdvisor data.
This section describes our approach for generating personalised summaries that
are tailored to a user`s preferences, and non-personalised summaries that re ect
the general opinion of users about a product.
2.1</p>
      <sec id="sec-1-1">
        <title>Feature Extraction and Sentiment Classi cation</title>
        <p>
          Inspired by the methods described in [
          <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
          ], we consider bi-gram features that
conform to one of two part-of-speech (POS) co-location patterns: a noun preceded
by an adjective (AN ) or by a noun (NN ). Single noun features that frequently
co-occur (&gt; 70% of the time) with sentiment words in the same sentence are
also considered [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          To evaluate the sentiment of a given feature Fi in sentence Sj of review Rk,
we identify the closest sentiment word wmin to Fi in Sj ; Fi is labelled as neutral
if no sentiment words are present. In this context, sentiment words are those
contained in the sentiment lexicon [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Next we extract the opinion pattern: the
POS tags for wmin, Fi and any words that occur between them. After a pass
over all features, the frequency of occurrence of all patterns is noted. For valid
patterns (those which occur more than once) we assign sentiment to Fi based
on that of wmin in the sentiment lexicon; sentiment is reversed if Sj contains a
negation term within a 4-word distance of wmin. Features associated with invalid
patterns are labelled as neutral.
2.2
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Dataset and Feature Representation</title>
        <p>Our corpus is taken from the hotel review site TripAdvisor.com. It contains
226,110 unique reviews written by 150,352 unique users for 2,500 hotels in 6
di erent cities. We mined over 270,000 unique features mention around 4,129,265
times using the process described in Section 2.1.</p>
        <p>We propose two levels of features to represent users and hotels; these levels
are determined by their similarity to amenities prede ned on the TripAdvisor
website. We obtain a set of base features, which are single-nouns and bi-gram
noun phrases from the set of amenities that TripAdvisor uses to described hotels
(e.g tness centre and wheelchair access ). We expect these features to be highly
meaningful and familiar to users. We hypothesise that there is an extended bag of
words that people use when talking about the same thing. For instance, a person
talking about `breakfast' may use words like `orange juice' or `bu et'. The key
thing to note is that these are related words, but not necessarily synonyms.
We apply k-Means to base features and their corresponding sentences to nd
other co-occurring, related features. The k most co-occurring features are used
to enhance the representation of each base feature as expanded features.
2.3</p>
      </sec>
      <sec id="sec-1-3">
        <title>User Preferences and Hotel Pro les</title>
        <p>Users normally talk about the things that matter to them in reviews. Therefore
we assume that the preferences of each user consist of the features they mention
in reviews and the relative frequency at which they mention them, which may
indicate their relevance to the user. Hence we de ne the pro le of a user as the
set of all features mentioned by the user in reviews. Each feature in the set is
tagged with the relative frequency with which it was mentioned in the user's
reviews.</p>
        <p>Similarly we de ne a hotel pro le as a set of features mentioned by users
about the hotel. Each feature in the set is tagged with its relative frequency and
average sentiment score (a value in the range [ 1; +1]).</p>
        <p>Generating Personalised and Opinionated Review Summaries</p>
      </sec>
      <sec id="sec-1-4">
        <title>Generating Summaries</title>
        <p>To construct non-personalised summaries for a user-hotel pair, two hotel pro les
are built using the base and expanded features respectively. Each feature is
assigned a ranking score that is the product of its average sentiment and its
normalised frequency. When the non-neutral features in the hotel pro le are
ranked by the ranking score, the top-n and bottom-n features form the pros and
cons parts of the explanation respectively.</p>
        <p>To generate a personalised summary for a user-hotel pair, two pro les are
built each for users and hotels using the base and expanded features respectively.
Each feature in the hotel pro le is assigned a ranking score that is the product
of its average sentiment and its normalised frequency. The ranking score of each
feature in the hotel pro le is updated by multiplying its original ranking score
with its normalised frequency in the user pro le. When the non-neutral features
in the hotel pro le are ranked by the updated ranking score, the top-n and
bottom-n features form the pros and cons parts of the explanation respectively.
In both explanation types, expanded features are mapped to their corresponding
base features that are familiar to the user.
2.5</p>
      </sec>
      <sec id="sec-1-5">
        <title>Examples</title>
        <p>Earlier in 2.4, we described how we can model a user and hotel pro le using
base and expanded features. In Fig. 1 we show a fragment of the user and hotel
pro le used to generate the example summary in Fig. 2.</p>
        <p>Fig. 2 shows a screenshot of a personalised summary generated for a
userhotel pair. The pro features are highlighted in green, and the cons features in red.
We always present base features in the summaries regardless of how we choose
to model the user and hotel pro le. This is to avoid having too ne a level of
granularity which might be unintuitive to users. Therefore `shuttle bus service'
is a pro feature of the hotel, ranked by the user's preferences. The tooltips (see
(c) in Fig. 2) display expanded features in snippets of sentences from reviews
that are associated with the base feature in the summary. Here the reviewers
have discussed the `printer' not working, and the closing hours of the `centre' ;
these have been summarised to `business centre'</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusion</title>
      <p>This paper presents a method for constructing personalised summaries of items
based on opinions from textual reviews. With TripAdvisor data we show how
the pros and cons of hotels can be explained to users using di erent feature
representations. Our technique focuses on those features that users write about
most frequently in their reviews. This forms the basis for prioritising features that
are likely to be of interest to the user compared to non-personalised explanations,
which focus on features that are commonplace for a hotel but may not be so
relevant for an individual user. This work builds on related work in the area of
opinion mining and recommender systems but considers a novel application in
the form of explanation generation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            , D.H., Han,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>The di erent e ects of online consumer reviews on consumers' purchase intentions depending on trust in online shopping malls: An advertising perspective</article-title>
          .
          <source>Internet research 21</source>
          (
          <year>2011</year>
          )
          <volume>187</volume>
          {
          <fpage>206</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>O</given-names>
            <surname>'Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Smyth</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          , Wilson,
          <string-name>
            <given-names>D.C.</given-names>
            ,
            <surname>McDonald</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Smeaton</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Improving the quality of the personalized electronic program guide</article-title>
          .
          <source>User Modeling and UserAdapted Interaction</source>
          <volume>14</volume>
          (
          <year>2004</year>
          )
          <volume>5</volume>
          {
          <fpage>36</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Mahony</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.P.</given-names>
            ,
            <surname>McCarthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Smyth</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Opinionated product recommendation</article-title>
          .
          <source>In: Case-Based Reasoning Research and Development. Volume 7969 of Lecture Notes in Computer Science</source>
          . Springer Berlin Heidelberg (
          <year>2013</year>
          )
          <volume>44</volume>
          {
          <fpage>58</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>O</given-names>
            <surname>'Mahony</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.P.</given-names>
            ,
            <surname>Smyth</surname>
          </string-name>
          ,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Learning to recommend helpful hotel reviews</article-title>
          .
          <source>In: Proceedings of the 3rd ACM Conference on Recommender systems, ACM</source>
          (
          <year>2009</year>
          )
          <volume>305</volume>
          {
          <fpage>308</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCarthy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>O'Mahony</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Towards an intelligent reviewer's assistant: Recommending topics to help users to write better product reviews</article-title>
          .
          <source>In: Proceedings of the 2012 ACM International Conference on Intelligent User Interfaces</source>
          ,
          <source>ACM</source>
          (
          <year>2012</year>
          )
          <volume>159</volume>
          {
          <fpage>168</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Mining and summarizing customer reviews</article-title>
          .
          <source>In: Proceedings of the Tenth ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. KDD '04</source>
          , New York, NY, USA, ACM (
          <year>2004</year>
          )
          <volume>168</volume>
          {
          <fpage>177</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Moghaddam</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ester</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Opinion digger: An unsupervised opinion miner from unstructured product reviews</article-title>
          .
          <source>In: Proceedings of the 19th ACM International Conference on Information and Knowledge Management. CIKM '10</source>
          , New York, NY, USA, ACM (
          <year>2010</year>
          )
          <year>1825</year>
          {
          <fpage>1828</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>