<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>mendations⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Discussion Paper</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Massimo</string-name>
          <email>davmassimo@unibz.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Ricci</string-name>
          <email>fricci@unibz.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Free University of Bozen-Bolzano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>We here focus on Points of Interest (POIs) Recommender Systems (RSs), aimed at helping users visiting a city to discover new and relevant POIs. RSs are often assessed in ofline settings, hence, measuring the system's precision in predicting previously observed user behaviour. However, when deployed, the system produced recommendations are often of limited use, because they lack novelty. We conjecture that this phenomenon is primarily due to the limited capability of RSs in extracting from the observed behaviour general characteristics of POIs that are relevant for diferent classes of users (tourist types). We compare an Inverse Reinforcement Learning (IRL) based RS algorithm with more traditional Nearest Neighbour and Popularity-based ones. Through an ofline evaluation, we show that the nearest neighbour and popularity-based RSs excel in precision (ofline) and are perceived as not novel by users of a live-user study. On the contrary, despite a lower ofline precision, the IRL-based RS, which learns the preferences of tourists for POIs characteristics, can give a better support to a tourist.</p>
      </abstract>
      <kwd-group>
        <kwd>recommender systems</kwd>
        <kwd>tourism</kwd>
        <kwd>user behaviour learning</kwd>
        <kwd>evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        This paper focuses on Recommender Systems (RSs) in the tourism domain and summarises the
results of Massimo and Ricci on Points of Interest (POIs) recommendation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is introduced a next POI-visit recommendation strategy, named Q - B A S E , which uses
Inverse Reinforcement Learning. It consists of three steps. Firstly, it identifies diferent tourist
clusters based on POI-visit sequence observations. Then, by analysing the sequential
consumption pattern of clustered users’ POI visits, it learns their preferences for POI features and the
reward a (generic) user belonging to a cluster obtains in conducting certain POI visits. Finally,
Q - B A S E computes the state action-value function  which tells how much total reward the user
will gain if she selects to visit any POI and keeps choosing POIs according to the optimal
visit selection policy which is optimal for their cluster. In an ofline evaluation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] Q - B A S E has
been compared to the sequence aware recommendation strategy S K N N [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Given a set of users’
POI-visit sequences and a target user (partial) POI-visit sequence, S K N N recommends to the target
⋆This paper presents some of the results published in Massimo, D., Ricci, F. Popularity, novelty and relevance in
user next POI visits, which are prevalent in the POI visit sequences of the most similar users.
An ofline evaluation showed that S K N N ofers recommendations that more precisely predict the
expected user behaviour, but are not novel. Besides, Q - B A S E excelled in suggesting novel POI
visits that are also more rewarding than S K N N . This comes at the cost of a lower precision. It was
conjectured that in an online RS, when Q - B A S E is again compared to S K N N , its recommendations
will be perceived as relevant because they contain novel POIs with a larger reward. Besides, it
was supposed that S K N N ’s higher precision is due to its bias in suggesting popular items. We
therefore investigated two research questions. (RQ1) If S K N N achieves higher precision by being
biased towards popular items, can Q - B A S E be modified, by biasing its recommendations towards
more popular items, to achieve a similar precision of S K N N ? (RQ2) Will online users like the
precise recommendations of S K N N more than those generated by Q - B A S E , which are more novel
and yet relevant?
      </p>
      <p>
        In order to investigate these questions we also proposed a novel and adaptable RS, called Q - P O P
P U S H , derived from Q - B A S E and that generates recommendations that simultaneously optimise
two criteria: the reward of the recommendations (as for Q - B A S E ) but also their popularity. Q
P O P P U S H computes the harmonic means of the scores given by Q - B A S E to a POI visit and the
popularity of the POI (visit occurrences in the observed data). To weigh the relative importance
of Q - B A S E original score and the popularity-based score, Q - P O P P U S H uses a parameter  . With
 &gt; 0.5 the popularity of the POI-visit has a higher importance, whereas with  &lt; 0.5 the Q - B A S E
component is weighted more. Equal importance is given with  = 0.5 . We have tested Q-POP
PUSH in an ofline experiment by comparing its performance with Q - B A S E , a popularity baseline
and two Nearest Neighbour next-item RSs: S K N N and s - S K N N [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. Finally, we have assessed the
user perception of the recommendations generated by the best performing (ofline) models in a
user study.
      </p>
      <p>The rest of this paper summarises the ofline experiment, the user study and concludes with
a discussion.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Experimental analysis</title>
      <p>
        The ofline experiment was conducted along the classical Machine Learning evaluation approach,
with train and test sets [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ]. The train set is used to train the models 1 and the test set is used
for recommendations generation and evaluation. In particular, 70% of a test POI-visit sequence
is used to identify the recommendations and the remaining 30% to assess the recommendation
performance. We used a dataset of 1663 POI-visit sequences (over 500 POIs) in the Italian city of
Florence, done by 1163 anonymous users. The metrics used to assess Q - B A S E , Q - P O P P U S H , S K N N ,
s - S K N N and the popularity baseline POP are: reward as in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; popularity intended as a proxy of
novelty; precision; and similarity of a recommendation list to the list generated by S K N N . Since
S K N N is very precise but also afected by a popularity bias, it is interesting to measure how much
Q - B A S E and Q - P O P P U S H deviates from S K N N . All the metrics range in [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ], where values close
to 0 (1) indicate a low (high) performance. For more details about procedures and metrics, we
refer to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
1Nearest neighbour models and POP do not need any training, but use train data at test time.
      </p>
      <p>In Table 1 we report the metric values for Top-5 recommendations, obtained in the ofline
evaluation. S K N N and s - S K N N perform essentially the same on the reward, precision and popularity
* There is a significant diference (  &lt; 0.05 ) between the best performing model and SKNN (two-tailed
paired t-test). The test is run for the metrics Rew, Prec and Pop for 5 repeated train-test split.
metrics. Not surprisingly, s - S K N N produces, on average, recommendations lists that overlap up
to 53% to the list of SKNN (SimKNN ). When comparing SKNN to POP, we note that the overlap
of the recommendations is quite low, but still, POP has a reasonably good precision, considering
the simplicity of the approach. Q - B A S E suggests much less popular items than SKNN and also
with higher reward.</p>
      <p>The analysis of Q - P O P P U S H performance allows to address the research question RQ1. By
introducing a popularity bias to Q - B A S E , as it is done in Q - P O P P U S H , the generated recommendations
become more similar to S K N N . In fact, when  is increased, the popularity of the recommendations
produced by Q - P O P P U S H rises, it reaches and passes that of S K N N . However, it is interesting to
note that with a rather small value of  = 0.009 Q - P O P P U S H obtains a precision of 0.062, which
is very close to the precision of S K N N (0.068) and s - S K N N (0.063). With this setting Q - P O P P U S H has
still a positive reward (0.02), while S K N N and s - S K N N have negative rewards. This means that a
small popularity bias in Q - B A S E could be beneficial to improve the system’s precision. We must
also note that with larger values of  the reward of Q - P O P P U S H is becoming smaller and smaller
and approaching that of the two nearest neighbour methods. Clearly, Q - P O P P U S H can be tuned
to balance two objectives: precision and reward.</p>
      <p>To answer the second research question RQ2 we have designed a live user study aimed at
measuring the user’s perceived novelty and appreciation of the recommendations generated
by Q - B A S E , Q - P O P P U S H , with parameter  = 0.5 (to give equal importance to POIs’ popularity
and reward) and the S K N N baseline. We have developed an online system to assess the quality of
next-POI recommendations ofered to users who have already visited some POIs. The system
initially asks to declare visited POIs, which are used to create a hypothetical itinerary the user is
supposed to have already visited. Then, the system generates next-POI recommendations with
the three tested RS algorithms, combines them in a unique list, and asks the user to evaluate them
based on a description of the recommended items. The user does not know which algorithm
recommends the displayed recommendations. The training data of the online system is the
same we have used in the ofline study. The user evaluates each POI by marking it with one or
more of the following labels: “I already visited it”, “I like it” for a next visit and “I didn’t know
it”. The online user study participants were recruited via social media and mailing lists. Out of
202 users that accessed the application, we identified 158 reliable recommendation sessions.
Table 2 shows the estimated probability that a user marks as “visited”, “novel”, “liked” or
both “liked” and “novel” a POI recommended by a specific RS. Q - B A S E recommends POIs that
are less likely to have been already visited by the user and more likely to be novel than those
suggested by Q - P O P P U S H and S K N N . As in the ofline experiment, Q - P O P P U S H and S K N N performs
similarly. Interestingly, Q - B A S E suggests fewer POIs that are liked when compared to the other
two strategies. Hence, apparently a more precise RS, based on an ofline test, also recommends
online items that the user will like more. Besides, the obtained results falsify our hypothesis
that optimising the reward of a recommendation, as Q - B A S E do, will produce recommendations
that the user will like more. However, Q - B A S E suggests more novel POIs and, interestingly, more
recommendations that are both liked and novel (last column in Table 2).</p>
      <p>In a successive analysis we computed the probability that a user will like a recommendation
given the following three conditions: she knows the POI but has not yet visited it; she has
already seen it; the item is novel (unknown). We have derived the following conclusions. The
users liked more the novel POI-visit suggestions generated by S K N N and Q - P O P P U S H than those
produced by Q - B A S E . This is due to the tendency of Q - B A S E to suggest POI visits that even if
they have the properties typically liked by the user, e.g., they are of the same type of the POIs
liked by the user, they are also not popular POIs. Hence, these suggested POIs are hard to be
appreciated. The conclusion we derived from the user study is that users tend to like more the
items they are familiar with, e.g., previously visited items or items that are not novel.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Discussion</title>
      <p>The results of our experiments seem to confirm that RSs that precisely predict the user choices
(ofline) are also liked most by real users. In our case, this means that Q - P O P P U S H and S K N N are
better RSs than Q - B A S E . Our explanation of this result is that both high ofline precision and
large probability of liked recommendations (online) are influenced by the popularity of the
recommended items. In fact, these popular items are often in the users’ test sets, and users are
likely to be familiar with them.</p>
      <p>Despite its lower ofline precision performance and a lower extent of liked recommendations
in the user study, Q - B A S E is the RS that may better accomplish the true goal of a tourism RS: it
suggests more next-POIs that are both liked and novel. So, by optimising the reward Q - B A S E
is capable of discovering novel items that are also appreciated (when users are able to assess
them).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Massimo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ricci</surname>
          </string-name>
          ,
          <article-title>Popularity, novelty and relevance in point of interest recommendation: an experimental analysis</article-title>
          ,
          <source>J. Inf. Technol. Tour</source>
          .
          <volume>23</volume>
          (
          <year>2021</year>
          )
          <fpage>473</fpage>
          -
          <lpage>508</lpage>
          . URL: https://doi.org/10. 1007/s40558-021
          <source>-00214-5. doi:1 0 . 1 0 0 7 / s 4 0</source>
          <volume>5 5 8 - 0 2 1 - 0 0 2 1 4 - 5</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Massimo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ricci</surname>
          </string-name>
          ,
          <article-title>Harnessing a generalised user behaviour model for next-poi recommendation</article-title>
          , in: S. Pera,
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ekstrand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Amatriain</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. O'Donovan</surname>
          </string-name>
          (Eds.),
          <source>Proceedings of the 12th ACM Conference on Recommender Systems, RecSys</source>
          <year>2018</year>
          , Vancouver, BC, Canada, October 2-
          <issue>7</issue>
          ,
          <year>2018</year>
          , ACM,
          <year>2018</year>
          , pp.
          <fpage>402</fpage>
          -
          <lpage>406</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kamehkhosh</surname>
          </string-name>
          , L. Lerche,
          <article-title>Leveraging multi-dimensional user models for personalized next-track music recommendation</article-title>
          , in: A.
          <string-name>
            <surname>Sefah</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Penzenstadler</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Alves</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          Peng (Eds.),
          <source>Proceedings of the Symposium on Applied Computing, SAC</source>
          <year>2017</year>
          , Marrakech, Morocco, April 3-
          <issue>7</issue>
          ,
          <year>2017</year>
          , ACM,
          <year>2017</year>
          , pp.
          <fpage>1635</fpage>
          -
          <lpage>1642</lpage>
          . URL: https://doi.org/10.1145/3019612. 3019756.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 0 1 9 6 1 2 . 3 0 1 9 7 5 6 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ludewig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <article-title>Evaluation of session-based recommendation algorithms, User Model</article-title>
          .
          <source>User-Adapt. Interact</source>
          .
          <volume>28</volume>
          (
          <year>2018</year>
          )
          <fpage>331</fpage>
          -
          <lpage>390</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bellogín</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Said</surname>
          </string-name>
          ,
          <source>Recommender Systems Evaluation</source>
          , Springer New York, New York, NY,
          <year>2018</year>
          , pp.
          <fpage>2095</fpage>
          -
          <lpage>2112</lpage>
          . URL: https://doi.org/10.1007/978-1-
          <fpage>4939</fpage>
          -7131-2_
          <fpage>110162</fpage>
          .
          <source>doi:1 0 . 1 0</source>
          <volume>0 7 / 9 7 8 - 1 - 4 9 3 9 - 7 1 3 1 - 2 _ 1 1 0 1 6 2 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Terveen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <article-title>Evaluating collaborative filtering recommender systems</article-title>
          ,
          <source>ACM Transactions on Information Systems (TOIS) 22</source>
          (
          <year>2004</year>
          )
          <fpage>5</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Koren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Turrin</surname>
          </string-name>
          ,
          <article-title>Performance of recommender algorithms on top-n recommendation tasks</article-title>
          ,
          <source>in: Proceedings of the Fourth ACM Conference on Recommender Systems</source>
          , RecSys '10,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2010</year>
          , p.
          <fpage>39</fpage>
          -
          <lpage>46</lpage>
          . URL: https://doi.org/10.1145/1864708.1864721.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 1 8 6 4 7 0 8 . 1 8 6 4 7 2 1 .</volume>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>