<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Mining implicit data association from Tripadvisor hotel reviews</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vittoria Cozza</string-name>
          <email>vittoria.cozza@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marinella Petrocchi</string-name>
          <email>marinella.petrocchi@iit.cnr.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angelo Spognardi</string-name>
          <email>spognardi@di.uniroma1.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information, Engineering, University of Padua</institution>
          ,
          <addr-line>Padua</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dipartimento di Informatica, Sapienza Università di Roma</institution>
          ,
          <addr-line>Rome</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>IIT-CNR</institution>
          ,
          <addr-line>Pisa</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>56</fpage>
      <lpage>61</lpage>
      <abstract>
        <p>In this paper, we analyse a dataset of hotel reviews. In details, we enrich the review dataset, by extracting additional features, consisting of information on the reviewers' profiles and the reviewed hotels. We argue that the enriched data can gain insights on the factors that most influence consumers when composing reviews (e.g., if the appreciation for a certain kind of hotel is tied to specific users' profiles). Thus, we apply statistical analyses to reveal if there are specific characteristics of reviewers (almost) always related to specific characteristics of hotels. Our experiments are carried out on a very large dataset, consisting of around 190k hotel reviews, collected from the Tripadvisor website.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Social media, forums, and blogs are privileged vehicles for
posting and spreading online reviews. Among the goods and services
that are discussed every day on the Internet, we can find those
belonging to the most disparate categories, like, e.g., food, clothes,
music, toys, etc. Particularly, the practice of choosing and booking
preferred destinations has been greatly eased by the possibility
for users to consult previous feedback about hotels and
restaurants. According to comScore Media Metrix1, Tripadvisor is the
world’s largest travel e-advice site, providing advices as
reported by actual travellers. Tripadvisor counts more than 87 million
visitors per month2.</p>
      <p>
        Not only common users, but also service providers have strong
motivations to analyse the myriads of posts, tweets, and
comments available online. The latter will benefit by adjusting, e.g.,
their products lines and advertisement campaigns, while the
former by relying on previous experiences for addressing their needs
and matching their expectations. Furthermore, online reviews are
a precious source of information, e.g., to unveil implicit and/or
unexpected characteristics of the reviewers. As an example, in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
the authors investigate if and how the words —and their use— in
a review are linked to the reviewer’s gender, country, and age.
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], the authors present a novel approach to build
featurebased user profiles and item descriptions by mining user-generated
reviews. Such additional information can be integrated into
recommender systems to deliver better recommendations and an
improved user experience.
      </p>
      <p>
        In our previous work [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we exploited a Tripadvisor dataset
in order to investigate how subjectivity of reviewers afects the
scores assigned to hotels. Thus, we leverage sentiment analysis
1https://www.comscore.com/Products/Audience-Analytics/Media-Metrix - All sites
last accessed December 23, 2017.
2https://www.comscore.com/Insights/Rankings - Statistics updated to June 2017.
techniques to identify mismatches between the text and the score
in online review platforms.
      </p>
      <p>Since several aspects can influence the customer experience
(e.g., the hotel price, or the presence of restaurants, cafe, discos
in the hotel neighborhood, the connections with bus/train
stations and airports, etc.), in this work we propose an automatic
approach - based on association rules - to understand which
factors most influence consumers’ reviews. We consider a very
large dataset consisting of around 190k hotel reviews collected
from Tripadvisor, enriching the dataset by extracting a series of
hotel-centric and reviewer-centric features. We leverage these
features to list correlations among hotel properties, reviewer’s
characteristics, and the review score. The results are obtained
applying association rules techniques to our dataset. Findings are
both expected - such as that the hotels close to entertainment and
food areas are ranked with the highest scores - and less intuitive
- such as that those reviewers featuring a very low activity
(measured with a lower bound in term of given reviews), considering
their stay in a particular area, select - very often - hotels with a
low number of transportation means in the neighbourhood.</p>
      <p>We argue that, with our approach, sociologists and marketing
experts could analyse the results of the association rules to better
understand some extra reviewers’s characteristics and
connections with the reviewed service. This kind of analysis paves the
way for surveying a larger segment of the population than that
usually interviewed through standard polls.
2</p>
    </sec>
    <sec id="sec-2">
      <title>DATASET</title>
      <p>To conduct our study, we grounded it in a dataset composed
of real reviews taken from the Tripadvisor3 website. In
particular, our dataset contains all the reviews that can be accessed on
the website between the 26th of June 2013 and the 25th of June
2014 – date of the newest extracted review – for hotels in New
York, Rome, Paris, Rio de Janeiro, and Tokyo. With a
straightforward approach, we were able to collect the following pieces of
information for each review:
• the review date, text, and numeric score;
• the reviewer username, location, and triptype, being the
type of trip, one among the following five categories:
Family, Friends, Couple, Solo Traveler, and Businessman;
• the ID of the hotel which the review refers to.</p>
      <p>In addition to the above elements, we collected from
Tripadvisor all the hotels of the considered reviews and included in
our review dataset some additional data regarding the reviewed
hotels. In particular, leveraging the ID of the hotel which the
review refers to, we have gathered
• the hotel name and full address (where full address
includes the street address, the city, and the country);
3http://www.tripadvisor.com
• the category of the hotel (number of stars);
• the number of guest pictures for the hotel.</p>
      <p>It is worth noting like the above lists are not exhaustive, i.e., they
do not represent all the information accessible from
Tripadvisor. As an example, further information available for a review
are the scores assigned by reviewers to specific aspects of a
hotel, like location, cleanliness, sleep quality, rooms, and service.
However, for the scope of the current work, we focus on those
summarised for the reader’s convenience in Table 1. We exploited
such pieces of information to further expand the dataset, with
enriched features, as described in the next Section 2.1.</p>
      <sec id="sec-2-1">
        <title>Basic information</title>
      </sec>
      <sec id="sec-2-2">
        <title>Review</title>
        <sec id="sec-2-2-1">
          <title>Date</title>
          <p>Text
Score
Reviewer username
Reviewer location
Triptype
Hotel ID</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Hotel</title>
        <p>
          Name
Street address
City
Country
Guest pictures
We have discarded reviews by “Anonymous” users, since they
represent users of the platform http://www.daodao.com—the
Chinese version of Tripadvisor—where all the reviewers are
indifferently grouped in this single virtual username. We have further
limited our analysis on reviews whose textual part is in English,
following the language identification and analysis approach
presented in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. While the reviews accessible from Tripadvisor in
the year under investigation are 353,167, after the pre-processing
the resulting dataset is made up of 189,304 reviews in English,
provided by 142,583 Tripadvisor’s registered users that reviewed
4,019 hotels. Table 1 recaps the information extracted from the
dataset, while Table 2 shows the distribution of the reviews per
given score value. As shown, the values distribution is highly
unbalanced, being the highest score the most frequent in the
dataset (reflecting indeed the distribution usually featured by
review platforms).
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Rating Value Occurrences</title>
        <p>1 6,504
2 8,826
3 24,627
4 64,949
5 84,398
Table 2: Distribution of the given scores in the dataset
2.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Hotel-centric and reviewer-centric features</title>
      <p>Starting from the information collected in the basic dataset, we
have augmented it performing some further elaboration. In
particular, we enriched the data regarding the reviewed hotel with
the following features:
• the popularity, defined as the number of reviews for a given
hotel. While we have neither the list of actual bookings
available, nor Tripadvisor requires the reviewer to show a
proof to have been a guest in the hotel, this feature, when
computed on a large number of reviews per hotel, could
indirectly act as a quantification of the actual hotel clients;
• the hotel triptype, defined as the most frequent reviewer
triptype for a given hotel (whereas triptypes are Families,
Friends, Couples, Solo Travelers, and Businessmen);
• the geospatial coordinates (latitude and longitude);
• three points of interest (POI) features, defined as the
number of transportation services, restaurants, and attractions,
respectively, in a range of 300 meters around the hotel.</p>
      <p>Popularity and Hotel triptype have been computed looking at
how many and which kind of reviewers have reviewed the hotel.
The geospatial coordinates have been calculated with Google
Places APIs4, starting from the hotel name and full address. Then,
latitude and longitude, together with the parameter “ radius=300”,
have been given as input to the Google Radarsearch API5 to find
the number of points of interest (POI) related to transportation,
food, and entertainment.</p>
      <p>
        The data regarding a reviewer, instead, have been enriched
with the following features:
• the reviewers’ activity, defined as the number of reviews
they have written (under the observation period). Our
intuition is that this feature could be useful to discriminate
between frequent travelers and sporadic ones.
• the gender of the reviewer. This feature has been extracted
with the Namsor Onomastics6 machine learning tool, able
to recognise the language behind a name, thus identifying
the gender according to that language vocabulary with
high accuracy [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>After cleaning the username from numbers and symbols and
splitting it in two parts (where one is likely to be the name and
the other one, when available, the surname), we have called the
“onomastics/api/json/gendre” API. This service takes as input
name and surname and returns the recognised gender. We have
used regular expressions to clean the username from symbols
and numbers and for splitting the username. This was possible
since, in many cases, the name and surname were separated by
a space, or the surname started with an uppercase letter. Some
examples of username are: “Eldon S”, “MeganJones88”.</p>
      <p>Unfortunately, for a subset of reviewers, it was not possible
to derive the gender from their usernames. This happened for
9,507 reviewers (corresponding to 6% of the entire reviewers
set), which wrote 12,653 reviews. Examples of usernames for
which it was not possible to derive the gender are
Hope-and</p>
      <sec id="sec-3-1">
        <title>Dreams, mistyrabbit, A TripAdvisor Member, R W, E A, Nickeykol,</title>
      </sec>
      <sec id="sec-3-2">
        <title>NawakRed, FreeTravel81. We labeled with unknown the gender of</title>
        <p>such 9,507 reviewers.</p>
        <p>Hereafter, we will refer to this dataset as the basic dataset.
Indeed, in the following, we will extract hotel-centric and
reviewercentric features to enrich the basic set (see Section 2.1).
4https://developers.google.com/places
5https://maps.googleapis.com/maps/api/place/radarsearch
6http://api.namsor.com/onomastics/api</p>
      </sec>
      <sec id="sec-3-3">
        <title>Features Hotel</title>
        <p>It is worth noting that Popularity, Hotel triptype, and Activity
have been calculated as the result of queries to the basic dataset,
with the aim of making explicit some data that originally were
implicit in the information at disposition. A story apart deserves the
computation of the reviewer gender, the points of interest close
to the hotel, and its geospatial coordinates. As above described,
the latter have been computed relying on external data sources,
namely the Google Points of Interest and the Namsor database,
containing 800k names and statistical information about names
in each country of the world.</p>
        <p>Table 3 recaps the hotel-centric and reviewer-centric features
we used to enrich the basic dataset.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>ASSOCIATION ANALYSIS</title>
      <p>
        Association rule mining is a well known and widely applied
methodology for discovering frequent patterns, correlations, and
causal structures in transaction and relational databases, as well
as in other information repositories [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Thus, given a set of
items (or itemsets), association rule mining allows to define rules
predicting the occurrence of an item (or more), given the
occurrence of other items in the same itemsets.
      </p>
      <p>A popular application is basket data analysis, where itemsets
are transactions, representing lists of items in the consumers’
baskets. An example of transaction is: {Bread, Steak, Juice, Butter,
Chips, Beer}. When several others are collected, e.g., in a large
database, the methodology allows to automatically find
associations like, e.g., {Bread} ⇒ {Steak} (steaks are often purchased
with bread). Beside sales transactions, the basket analysis can be
applied to other situations like click stream tracking, spare parts
ordering and online recommendation engines - just to name a
few7.</p>
      <p>An association rule (AR) is generally defined as an implication
expression of the form X ⇒ Y , where X and Y are disjoint
itemsets. They represent, resp., the condition and the consequence
of the rule.</p>
      <p>The strength of an AR is commonly measured through the
two metrics support and confidence . Support gives the fraction of
itemsets in the dataset that contains both X and Y . Confidence
says how frequently items in Y appear in itemsets that contain X .
As an example, we want to known the strength of the rule {Bread}
⇒ {Steak} in a dataset with 100 transactions, corresponding to 100
consumers’ baskets. Suppose that itemset {Bread, Steak} occurs 30
times, and that itemset {Bread} occurs 40 times, than the support
of the rule is equal to 13000 , while its confidence is 3400 .</p>
      <p>
        As discussed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], rules with high values for confidence and
support do not always correspond to meaningful ARs, especially
when working with real datasets, due data can be unbalanced.
7http://pbpython.com/market-basket-analysis.html
(1)
(2)
(3)
For example, one rule could have a very high confidence, but
only due to the fact that the item in the consequence is very
frequent. In this case, the rule is not relevant. Instead, one rule
could have a low confidence, due to the fact that the item in
the consequence is very unfrequent in general, but it could still
be relevant. Considering the above observation, to evaluate the
statistical significance of the ARs, two other metrics are often
used: lift and convinction.
      </p>
      <p>Lift is defined as the confidence divided by the support of the
consequence:
li f t (X =⇒ Y ) =</p>
      <p>supp(X ∩ Y )
supp(X ) ∗ supp(Y )
With respect to confidence, the lift measures the importance of
the association considering also the dependence from the support
of the consequence.</p>
      <p>Convinction is defined by the ratio of the frequency of itemsets
that don’t contain the consequence, to the frequency of incorrect
predictions:
conv(X =⇒ Y ) =</p>
      <p>1 − supp(Y )
1 − con f (X =⇒ Y )</p>
      <p>Both lift and conviction values ranging over the (0,1)
interval mean negative dependence, values above 1 mean positive
dependence, and a value of 1 means independence.</p>
      <p>When items are also divided according to diferent classes, it
is possible to force the AR analysis to return a specific class in
the consequence. The obtained rule is called “class association
rule" (CAR). The CAR is an implication of the form:</p>
      <p>X =⇒ y , where X ⊆ I and y ∈ Y
where I stands for the itemsets and Y for the classes. The
definition of the aforementioned metrics holds also for CARs.</p>
      <p>
        The a priori algorithm [
        <xref ref-type="bibr" rid="ref16 ref2">2, 16</xref>
        ] is one of the most popular
algorithms to find frequent itemsets, i.e., itemsets whose support ≥
minsup.
      </p>
      <p>In this work, we apply the association rule mining to the hotel
reviews scenario. Each itemset corresponds to a distinguished
review, and it is a vector whose components are the values of
the features extracted and detailed in Section 2. The same
features are reported in Tables 4, 5, 6 for the reader’s convenience,
together with additional information that are useful here. CARs
analysis can be applied when considering also the class, that in
our scenario corresponds to the review score, a discrete value
with a range between 1 and 5.</p>
      <p>To enable the application of the a priori algorithm, we have
ifrst discretised those features that natively ranged over a large
set of values. As an example, in Table 5, a very low label for Guest
Pictures indicates a hotel with a number of pictures comprised
from 0 to 11. Still in that table, a medium label for Popularity
means a hotel that has been reviewed n times, where n ranges
over [433, 1156]. The values in Table 6 should be read as follows:
looking at the first line of the “Geo Food" part of the table, our
review set contains 37,851 reviews about a hotel, which has a
number of restaurants in the range [0, 37] within a radius of 300
mt. Indeed, many diferent reviews are on the same hotels, being
the number of hotels reviewed equal to 4,019, see Section 2.</p>
      <p>All the tables also report the Frequency indication, i.e., how
many reviews correspond to those values for those features, with
respect to the values and features in the tables (still quite
obviously, the sum on the values in the Frequency column equals to
the total number of reviews considered, 189,304).
value
solo
couple
family
friends
business
unspecified</p>
      <p>
        value
up to 5 reviews
6 or more
TripType
In order to find ARs and CARs, we applied the Weka
framework [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] implementation of the a priori. The Weka a priori
implementation allows to rank the rules according to diferent
metrics. Among them, we rely on confidence, lift, and conviction.
For AR analysis we generate a large number of rules with lift
above 1. For CAR analysis, we generate a large number of rules
with confidence above 0.2 and then we compute the lift (since,
for CAR, Weka does not natively include the ranking based on
lift). We finally select the rules with lift greater than 1.
      </p>
      <p>Both for the generated ARs and CARs, we then manually select
the most interesting rules, among those with the highest lift and
conviction. Table 7 and Table 8 report an excerpt of the results
for both scenarios.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>Association analysis results are reported in Table 7 and Table 8,
please notice we only consider those rules that lead to a lift and
conviction greater than 1. It is worth noting like |X ∩ Y |, when
divided by the size of the dataset, corresponds to the support of
the given rule.</p>
      <p>We summarise the main findings, as follows. Rule r1 states that
those reviewers featuring a very low activity, considering their
stay in France, select - very often - hotels with a low number of
transportation means in the neighbourhood. The rule holds for
19,199 reviews, over a total of 29,837 reviews, with equal
premises. Rule r2 says that males visiting US prefer hotels with a high
popularity. Rule r7 says that, when the hotel has low
transportation means in the neighbourhood, and the number of stars for
that hotel is unknown (this may corresponds to accommodation
facilities like hostels), its rating is equal to 3. Rule r10 states that
Japanese people staying in a 3 stars hotels rate those hotels with
a score equal to 4. Rule r14 in Table 8 states that hotels close to
entertainments, which are 37,998, are scored with the top score
5 the 50% of times.</p>
      <p>This kind of study provides a general approach for a
preliminary data exploration. While the explanation for certain rules is
very intuitive, well-grounded justification for others is left to
experts in the field. We argue that this kind of analysis corresponds
to a preliminary step, useful for suggesting which extra-features
could be exploitable to build an enhanced hotel
recommendation system. Also, we acknowledge that the analysis is based on
the available (direct or indirect) information, obtained from the
Tripadvisor’s website. More detailed features could consider
elements like price or number of guests. This would allow to obtain
other interesting rules, which remain an exclusive prerogative of
the hoteliers.
4</p>
    </sec>
    <sec id="sec-6">
      <title>RELATED WORK</title>
      <p>
        E-advice technology ofers a form of “electronic word-of-mouth”,
with new potential for gathering valid suggestions that guides
the consumer’s choice. Extensive and nationally representative
surveys have been carried out in the recent past, “to evaluate the
specific aspects of ratings information that afect people attitudes
toward e-commerce”. It is the case, e.g., of work in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which
highlights how people, while taking into accounts the average
of ratings for a product, still do not take care of the number of
reviews leading to that average. Recent work showed that,
instead of showing first to the users the reviews with the highest
scores, a diferent order, based, e.g., on the user profile, could be
considered [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]: that work integrates new features based on the
user profile into recommender systems, to deliver better
recommendations and provide an improved user experience. Similarly,
in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], the authors focus on score values given by previous
contributors whose preferences are close to the user’s preference.
Even almost one decade ago, the work in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] applies text mining
tools to online reviews to define rules sets, to identify contextual
information in the texts, which goes beyond a mere order of
numerical scores. Similarly to our work, they rely on Tripadvisor,
focusing however on text analysis only.
      </p>
      <p>
        However, the cited literature proposes systems that
recommend a service based on the intrinsic characteristics of that
service (e.g., characteristics of the hotel and its facilities). Other
works, similar to ours, investigate if, and how, the review data
hide social and/or economic information of the reviewers. One
example is mining reviews to exploit them as a textual resource
for sociolinguistic studies at a large-scale, as done in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This
work leverages the size of the reviews corpus as a more
statistically solid base for the analysis, with respect to manually-collected
corpora. Since reviews sites, such as Trustpilot8, may contain
reviewer metadata like, e.g., age, gender and location, the work
highlights gender-specific lexical diferences, the the distribution
of regional markers, spelling variations and the use of
grammatical constructions across the reviewers.
      </p>
      <p>
        The work from [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], which focused on reviews manipulation,
exploits reviewer-centric and hotel-centric features to identify
outliers: the work compares hotels reviews and related features
across diferent review sites, outperforming the detection of
suspicious hotels with respect to check the reviews on sites in isolation.
Relying on visualization tools, the authors of [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] highlight
suspicious changes on reviews scores, while work in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] proposes
new score aggregators to let review systems robust with respect
to injection of fake scores.
      </p>
      <p>
        Research efort has also being spent to understand which are
the factors that let a review perceived as useful: in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], the
authors highlight how the reviewer history is a dominant factor
to let a review be voted as useful or not. In [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] propose to use
the reviews as a source for demographic recommendations.
      </p>
      <p>
        In this work we enhance the review dataset with additional
features based on characteristics of the reviewer (e.g., gender)
and the hotel (e.g., popularity and the neighbourhood). On the
contrary, work in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] studies how, independently from the type
of service or the type of reviewer, the scores may be afected by
external factors, such as the whether conditions and the daylight
length of the service cities. We leverage an extensive
experimental campaign, addressing around 190k real reviews, which leads
to the provision of statistically sound results. Addressing a large
scale of data has been done also in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], which already has targeted
users’ reviews as a rich source of information for sociolinguistic
studies. While they achieve correlations between metadata in the
reviewers’ profile and the review text to let writing styles emerge,
we highlight association evidence among hotels and reviewers
features and the reviewer’s attitude to score the hotel.
5
      </p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSIONS</title>
      <p>We focused on hotel reviews to investigate which factors could
impact the scores that reviewers assign to hotels throughout the
world.First of all we have enriched review data with with novel
hotel-centric and reviewer-centric features, obtained for example
through linked data information available from the web, then we
have applied association rule mining to focus on these features
possibly motivating the classification scores.</p>
      <p>The approach can help both consumers and providers: the
former could achieve a better awareness on how to read the
reviews (consumers), the latter on how to improve their services
(providers). The providers also can query a very large segment of
population, in an automatic way and without relying on standard
interviews.</p>
      <p>The proposed technique is also applicable to a various range
of services: accomodation, car rental, food services, to cite a few.
Being association rule mining parametric with respect to the
itemsets in input, the approach is easily extensible to further
features not considered here, such as, e.g., the service price.
6</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENTS</title>
      <p>This research is partly supported by the EU H2020 Program, grant
agreement #675320 (NECS: European Network of Excellence in
Cybersecurity). Funding has also been received by Fondazione Cassa
di Risparmio di Lucca that partially finances the regional project
ReviewLand. Vittoria Cozza is also supported by the Starting
Grants Project DAKKAR (DAta benchmarK for Keyword-based
Access and Retrieval) promoted by University of Padua, Italy and
Convinction
r2
r4
r6</p>
      <sec id="sec-8-1">
        <title>Condition</title>
      </sec>
      <sec id="sec-8-2">
        <title>Confidence Convinction 0.64 0.59</title>
        <p>26,505
20,926
20,4316
34,192
20,362</p>
        <p>Lift
r7
r8
r9
high ==&gt; rating=5}
{memberActivity=very
low
guestPics=very
high
geoenter=very high ==&gt; rating=5}
{star=3 geofood=very high ==&gt; rating=4}
{geoenter=very high ==&gt; rating=5}
0.25
Fondazione Cariparo, Padua, Italy. The first author would like to
thank Giorgio Maria Di Nunzio, for his helpful support.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Silvana</given-names>
            <surname>Aciar</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Mining context information from consumer's Reviews. Prooceedings of the Context-Aware Recommender Systems</article-title>
          (CARS)
          <string-name>
            <surname>Workshop</surname>
          </string-name>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Rakesh</given-names>
            <surname>Agrawal</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ramakrishnan</given-names>
            <surname>Srikant</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>Fast Algorithms for Mining Association Rules in Large Databases</article-title>
          .
          <source>In Proceedings of the 20th International Conference on Very Large Data Bases (VLDB '94)</source>
          . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA,
          <fpage>487</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Brin</surname>
          </string-name>
          , Rajeev Motwani,
          <string-name>
            <given-names>Jefrey D.</given-names>
            <surname>Ullman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Shalom</given-names>
            <surname>Tsur</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Dynamic Itemset Counting and Implication Rules for Market Basket Data</article-title>
          .
          <source>In Proceedings of the 1997 ACM SIGMOD International Conference on Management of Data (SIGMOD '97)</source>
          . ACM, New York, NY, USA,
          <fpage>255</fpage>
          -
          <lpage>264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Elian</given-names>
            <surname>Carsenat</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Onomastics and Big Data Mining</article-title>
          .
          <source>CoRR abs/1310</source>
          .6311 (
          <year>2013</year>
          ). http://arxiv.org/abs/1310.6311
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Celli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. Marta L.</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Lascio</surname>
          </string-name>
          , Matteo Magnani, Barbara Pacelli, and
          <string-name>
            <given-names>Luca</given-names>
            <surname>Rossi</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Social Network Data and Practices: The Case of Friendfeed</article-title>
          .
          <source>In Advances in Social Computing. LNCS</source>
          , Vol.
          <volume>6007</volume>
          . Springer Berlin Heidelberg,
          <fpage>346</fpage>
          -
          <lpage>353</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Colantonio</surname>
          </string-name>
          , Roberto Di Pietro, Marinella Petrocchi, and
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Spognardi</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Visual detection of singularities in review platforms</article-title>
          .
          <source>In 30th Annual ACM Symposium on Applied Computing</source>
          ,.
          <volume>1294</volume>
          -
          <fpage>1295</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Pietro</surname>
          </string-name>
          , Marinella Petrocchi, and
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Spognardi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A Lot of Slots - Outliers Confinement in Review-Based Systems</article-title>
          .
          <source>In Web Information Systems Engineering Part I. 15-30.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Ruihai</given-names>
            <surname>Dong</surname>
          </string-name>
          and
          <string-name>
            <given-names>Barry</given-names>
            <surname>Smyth</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>From More-Like-This to Better-ThanThis: Hotel Recommendations from User Generated Reviews</article-title>
          .
          <source>In Proceedings of the 2016 Conference on User Modeling Adaptation and Personalization (UMAP '16)</source>
          . ACM, New York, NY, USA,
          <fpage>309</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Michela</given-names>
            <surname>Fazzolari</surname>
          </string-name>
          , Vittoria Cozza, Marinella Petrocchi, and
          <string-name>
            <given-names>Angelo</given-names>
            <surname>Spognardi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Study on Text-Score Disagreement in Online Reviews</article-title>
          .
          <source>Cognitive Computation 9</source>
          ,
          <issue>5</issue>
          (
          <issue>01</issue>
          <year>Oct 2017</year>
          ),
          <fpage>689</fpage>
          -
          <lpage>701</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>AndrewJ</surname>
            . Flanagin, MiriamJ. Metzger, Rebekah Pure, Alex Markov, and
            <given-names>Ethan</given-names>
          </string-name>
          <string-name>
            <surname>Hartsell</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Mitigating risk in e-commerce transactions: perceptions of information credibility and the role of user-generated ratings in product quality and purchase intention</article-title>
          .
          <source>Electronic Commerce Research</source>
          <volume>14</volume>
          ,
          <issue>1</issue>
          (
          <year>2014</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Hall</surname>
          </string-name>
          , Eibe Frank, Geofrey Holmes, Bernhard Pfahringer,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Reutemann</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ian H Witten</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>The WEKA data mining software: an update</article-title>
          .
          <source>ACM SIGKDD explorations newsletter 11</source>
          ,
          <issue>1</issue>
          (
          <year>2009</year>
          ),
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Jochen</surname>
            <given-names>Hipp</given-names>
          </string-name>
          , Ulrich Güntzer, and
          <string-name>
            <given-names>Gholamreza</given-names>
            <surname>Nakhaeizadeh</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Algorithms for Association Rule Mining: a General Survey and Comparison</article-title>
          .
          <source>SIGKDD Explor. Newsl. 2</source>
          ,
          <issue>1</issue>
          (
          <year>June 2000</year>
          ),
          <fpage>58</fpage>
          -
          <lpage>64</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Dirk</surname>
            <given-names>Hovy</given-names>
          </string-name>
          , Anders Johannsen, and
          <string-name>
            <given-names>Anders</given-names>
            <surname>Søgaard</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>User Review Sites As a Resource for Large-Scale Sociolinguistic Studies</article-title>
          .
          <source>In 24th International Conference on World Wide Web (WWW '15)</source>
          .
          <fpage>452</fpage>
          -
          <lpage>461</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Nikolaos</given-names>
            <surname>Korfiatis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marios</given-names>
            <surname>Poulos</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Using online consumer reviews as a source for demographic recommendations: A case study using online travel reviews</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>40</volume>
          ,
          <issue>14</issue>
          (
          <year>2013</year>
          ),
          <fpage>5507</fpage>
          -
          <lpage>5515</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Asher</given-names>
            <surname>Levi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Osnat</given-names>
            <surname>Mokryn</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The Social Aspect of Voting for Useful Reviews</article-title>
          . In Social Computing,
          <string-name>
            <surname>Behavioral-Cultural Modeling</surname>
            and
            <given-names>Prediction. LNCS</given-names>
          </string-name>
          , Vol.
          <volume>8393</volume>
          . Springer International Publishing,
          <volume>293</volume>
          -
          <fpage>300</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Bing</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Wynne Hsu, and Yiming Ma.
          <year>1998</year>
          .
          <article-title>Integrating Classification and Association Rule Mining</article-title>
          . In KDD.
          <volume>80</volume>
          -
          <fpage>86</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Amanda</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Minnich</surname>
            , Nikan Chavoshi, Abdullah Mueen, Shuang Luan, and
            <given-names>Michalis</given-names>
          </string-name>
          <string-name>
            <surname>Faloutsos</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>TrueView: Harnessing the Power of Multiple Review Sites</article-title>
          .
          <source>In 24th International Conference on World Wide Web (WWW '15)</source>
          .
          <fpage>787</fpage>
          -
          <lpage>797</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>SyedA. Rahman</surname>
            , Tazin Afrin, and
            <given-names>Don</given-names>
          </string-name>
          <string-name>
            <surname>Adjeroh</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Determinants of User Ratings in Online Business Rating Services</article-title>
          . In Social Computing,
          <article-title>BehavioralCultural Modeling, and</article-title>
          <string-name>
            <surname>Prediction. LNCS</surname>
          </string-name>
          , Vol.
          <volume>9021</volume>
          . Springer International Publishing,
          <volume>412</volume>
          -
          <fpage>420</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Koji</surname>
            <given-names>Takuma</given-names>
          </string-name>
          , Junya Yamamoto, Sayaka Kamei, and
          <string-name>
            <given-names>Satoshi</given-names>
            <surname>Fujita</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A Hotel Recommendation System Based on Reviews: What Do You Attach Importance To?</article-title>
          .
          <source>In Fourth International Symposium on Computing and Networking</source>
          ,
          <string-name>
            <surname>CANDAR</surname>
          </string-name>
          <year>2016</year>
          , Hiroshima, Japan,
          <source>November 22-25</source>
          ,
          <year>2016</year>
          .
          <fpage>710</fpage>
          -
          <lpage>712</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>