<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards an Automated Writing Assistant for Online Reviews</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>James Caverlee</string-name>
          <email>caverlee@tamu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Parisa Kaghazgaran Texas A&amp;M University College Station</institution>
          ,
          <addr-line>TX</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Texas A&amp;M University College Station</institution>
          ,
          <addr-line>TX</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Online reviews play a critical role in persuading or dissuading users when making purchase decisions. And yet very few users take the time to write helpful reviews. Encouragingly, recent advances in AI offer good potential to produce review-like natural language content. However, The main challenge is a lack of large, high-quality labeled data at both the aspect and sentiment level for training the automated review generators. Hence, we study the feasibility of a writing assistant framework in order to help users post online reviews and introduce a data-driven approach to label data required for training. We evaluate the effectiveness of our approach by launching a user study on how end-users perceive the quality of the labels and automated generated reviews.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>________________________________________________________
Workshop proceedings Automation Experience across Domains
In conjunction with CHI'20, April 26th, 2020, Honolulu, HI, USA
Copyright © 2020 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).
Website: http://everyday-automation.tech-experience.at</p>
    </sec>
    <sec id="sec-2">
      <title>Author Keywords</title>
      <p>Intelligibility; review generation; review writing assistant;
user-based evaluation</p>
    </sec>
    <sec id="sec-3">
      <title>Introduction</title>
      <p>
        There is a growing attention in creating new methods to
help users share their opinions on review platforms. For
example, Airbnb site require hosts and guests to write
mutual reviews [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, such a requirement may be an
impediment to customer engagement in other platforms
that feature products like movies and books. In another
promising direction, new tools based on natural language
generation have shown good success in some domains.
For example, carefully configured templates can transform
well-structured data into legible text, especially for domains
with consistent format and structure like weather forecast
reports [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Olympics stories [
        <xref ref-type="bibr" rid="ref7">6</xref>
        ], and corporate earnings
reports [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ]. However, such methods face challenges for
online reviews that typically cover a broad range of categories
(e.g., apps, products, restaurants) with multiple aspects
within each category (e.g., food, service, staff and so on in
the restaurant domain) and diverse opinions that do not fit
in a single template.
      </p>
      <p>
        Seed data [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. It is used to label the unlabeled reviews
containing 1,472 pairs of reviews and their labels.
      </p>
      <p>Yelp Dataset. It is used to train the review generator
containing about 4M reviews on 60k restaurants.</p>
    </sec>
    <sec id="sec-4">
      <title>Automatic Review Labeling</title>
      <p>Since users express opinion on different aspects of the
target in a single review, a review cannot be labeled in whole
to state a specific aspect (see Figure 1). We first split a
review into its topically coherent segments and propose to
label the resulting segments.</p>
      <sec id="sec-4-1">
        <title>Review Segmentation.</title>
        <p>
          The review segmentation algorithm traverses through each
review sentence by sentence in order to cluster coherent
sentences into one segment. At its core, review
segmentation is based on a sliding window technique with a window
size of two. Each sentence is compared with the right most
sentence in the previous segment and if their distance is
less than a specific threshold , then the sentence is added
to the segment, otherwise it forms a new segment. This
process continues until the end of the review in linear time.
We adopt the Word Mover’s Distance (WMD) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] to
measure the similarity between sequential sentences. Rather
than relying on keyword matching, it attempts to find an
optimal transformation from one sentence to another
sentence in the word embedding space. Briefly, word
embedding technique map each word into a numeric vector such
that co-occurred words are close in the new vector space.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Label Assignment.</title>
        <p>The label assignment algorithm is based on a small seed
set. The main intuition is to find semantically similar seeds
to the unlabeled segments and use their labels to identify
the aspect of the segments. For this purpose, we compare
each sample in the seed set against the unlabeled
segments using the WMD distance function. The unlabelled
segment receives the label of closest seed sample.</p>
        <p>Table 1 demonstrates examples of segments obtained from
the segmentation algorithm and their labels assigned by
the label assignment algorithm across different aspects and
steak sandwich was delicious and the caesar salad had an absolutely delicious dressing with a perfect amount of
dressing and distributed perfectly across each leaf . i know i m going on about the salad .
today was my second visit to the place after having a good first experience but i am so disappointed with the quality
of the food that i can say it has been my worst experience of food in months the sun dried tomatoes very absolutely stale
to an extent that they tasted bitter the pizza base was so thick that it was uncooked and soggy the four cheese blend tasted
completely different than the last time and so did the pesto sauce . no consistency with food quality .
the ambiance is nice too . it s a bit dark but they have this nice light display above on the ceiling made with mason jars .
there is a comfy seating area in the bar area that s nice too .
however the one thing that surprised me was how dirty the restroom was in this restaurant . the floor was really dirty
and toilet papers were unwell kept . the restaurant could at least have someone maintained the restroom in good shape
and clean because this will reflect on how one maintains the cleanliness of the place .
the price is very reasonable for a family of four with plenty of leftovers to take home .
my wife i had a groupon for this place and for the price it was very poor value quality .
i had a nice glass of california cabernet . the wine list while not expansive was good . the bartender i had seemed
to have a nice knowledge of what was going on with the wine that encompassed it .
i ordered a glass of Merlot that was delivered to me in a dirty glass . the waitress was very polite and went to
get me a new glass of wine but i was still unimpressed at that point .
highly recommend for lunch . even during lunch rush it was not super packed . this would be a good place for a lunch meeting . General (+)
i am not sure why anyone would like this place . the only thing it has going is location and that is simply not enough
not for me .
sentiments.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Evaluation.</title>
        <p>We evaluate if automated labels are comparable with
manual labeling. We set up a crowd-based user study to
verify if the labels are assigned truthfully according to human
readers. We post 100 surveys on Amazon Mechanical Turk
(AMT) each including a guideline and a set of reviews for
which we seek a label from Turkers. The guideline has two
major points: (i) it shows a sample of reviews along with
their labels from the seed set to provide a context on how
reviews and labels are paired with each other, (ii) it asks
Food (+)
Food (-)</p>
        <sec id="sec-4-3-1">
          <title>General (+)</title>
          <p>General (-)</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>Ambience (+)</title>
          <p>Ambience (-)</p>
        </sec>
        <sec id="sec-4-3-3">
          <title>Price (+)</title>
          <p>Price (-)</p>
        </sec>
        <sec id="sec-4-3-4">
          <title>Drink (+) Drink (-)</title>
          <p>Turkers to label the reviews through a series of multi-choice
questions. We design 100 surveys each with 10 reviews
to cover all the labels. Each unique survey is assigned to
three workers, i.e., 3 HITs (Human Intelligence Task) per
task, giving us a total of 300 surveys and 3,000 questions.
To ensure the quality of responses, we insert a trivial
question into each survey, which asks the Turker to check if a
mathematical equation is False or True. It helps to manage
the risk of blindly answered surveys. Furthermore, we only
accept surveys from Turkers with approval rating of at least
95% and those who dwell on the survey for at least 7
minutes. We also restrict our tasks to workers located in the
United States to guarantee English literacy.</p>
          <p>Table 2 demonstrates the performance of automatic
labeling against human judgment across various labels. The
majority of the labels are found accurate by human
evaluators with at least 80% accuracy. However, the accuracy
for labels Price/negative and Drink/negative is relatively low
and we can relate this to the fact that these labels do not
have a significant representation in the seed set.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Automatic Review Generation</title>
      <p>Here, we introduce a generative model based on language
neural networks to automate user review generation. The
main intuition is to propose a sequence to sequence model
where the first sequence encodes the labels of the
segments obtained from previous step into a context vector.
This context vector along with review words are the input to
the second sequence in order to train the review generator.
Formally, Given the aspect A as input attribute, we aim to
generate a review sequence R = (w1; ::; wjRj) constrained
by the input to maximize the conditional probability p(RjA):
jRj
p(RjA) = Y p(wtjw&lt;t; A)
t=1
(1)
where w&lt;t refers to the tokens seen until time step t.</p>
      <sec id="sec-5-1">
        <title>Evaluation.</title>
        <p>Similar to label assessment, we launch a crowd-based user
study by posting surveys on AMT. We follow similar
guidelines to ensure the quality of the answers. We design 100
surveys each with 10 generated reviews at various aspects.
We assign 3 HITs per task, giving us a total of 300 surveys
and 3000 questions. We ask Turkers to label the
modelgenerated reviews through a multi-choice questions based
on the aspect.</p>
        <p>From Table 3, we observe that generated reviews stay with
the desired aspect with higher than 90% accuracy for a
majority of the labels. For example, 93% and 97% of reviews
on Food (positive and negative) are perceived equally by
the model and the human evaluators while this number is
34% for drink/negative. We can relate this to the fact that
the label Food has a better representation in both seed set
and our expanded dataset as it is the main topic of
discussion when writing a review for a restaurant.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Discussion</title>
      <p>We have explored how to make use of automation to help
users writing online reviews in particular when sufficient
data required by neural networks are not available. We
automate the review writing process at two steps: (i) build a
ground truth of reviews at aspect and sentiment level and
evaluate the effectiveness of the labels by human readers,
(ii) propose a generative model that produces reviews
conditioned on input aspects and evaluate the quality of the
generated reviews through user study. In the next step, we
aim to study how users are willing to use the proposed
system and how we can incorporate their intention at design
level.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ibrahim</given-names>
            <surname>Adeyanju</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Generating weather forecast texts with case based reasoning</article-title>
          .
          <source>arXiv</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Matt</given-names>
            <surname>Kusner</surname>
          </string-name>
          , Yu Sun, Nicholas Kolkin,
          <string-name>
            <given-names>and Kilian</given-names>
            <surname>Weinberger</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>From word embeddings to document distances</article-title>
          .
          <source>In International conference on machine learning.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Dean</surname>
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Lehr</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>An analysis of the changing competitive landscape in the hotel industry regarding Airbnb</article-title>
          .
          <article-title>(</article-title>
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Maria</given-names>
            <surname>Pontiki</surname>
          </string-name>
          , Dimitris Galanis, Haris Papageorgiou, Ion Androutsopoulos, Suresh Manandhar,
          <string-name>
            <surname>AL-Smadi</surname>
            <given-names>Mohammad</given-names>
          </string-name>
          , Mahmoud Al-Ayyoub,
          <string-name>
            <given-names>Yanyan</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Bing</given-names>
            <surname>Qin</surname>
          </string-name>
          , Orphée De Clercq, and others.
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>Semeval-2016 task 5: Aspect based sentiment analysis</article-title>
          .
          <source>In Proceedings of the 10th international workshop on semantic evaluation (SemEval-2016)</source>
          .
          <fpage>19</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Roetzer</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>How the Associate Press and the Orlando Magic Write Thousands of Content Pieces in Seconds</article-title>
          , https://bit.ly/2HCyAiS, Last Access:
          <volume>08</volume>
          /14/
          <year>2019</year>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <surname>WashPostPR.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>The Washington Post experiments with automated storytelling to help power 2016 Rio Olympics coverage</article-title>
          , https://wapo.st/2G67Sg6, Last Access:
          <volume>08</volume>
          /14/
          <year>2019</year>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Yuanshun</given-names>
            <surname>Yao</surname>
          </string-name>
          and et al.
          <year>2017</year>
          .
          <article-title>Automated crowdturfing attacks and defenses in online review systems</article-title>
          .
          <source>In CCS.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>