<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Item Familiarity Effects in User-Centric Evaluations of Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lukas Lerche TU Dortmund</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germany lukas.lerche@tu- dortmund.de</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Recommender Systems; User-centric Evaluation; Bias</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dietmar Jannach TU Dortmund</institution>
          ,
          <addr-line>Germany dortmund.de</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Michael Jugovac TU Dortmund</institution>
          ,
          <addr-line>Germany dortmund.de</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Laboratory studies are a common way of comparing recommendation approaches with respect to di erent quality dimensions that might be relevant for real users. One typical experimental setup is to rst present the participants with recommendation lists that were created with di erent algorithms and then ask the participants to assess these recommendations individually or to compare two item lists. The cognitive e ort required by the participants for the evaluation of item recommendations in such settings depends on whether or not they already know the (features of the) recommended items. Furthermore, lists containing popular and broadly known items are correspondingly easier to evaluate. In this paper we report the results of a user study in which participants recruited on a crowdsourcing platform assessed system-provided recommendations in a between-subjects experimental design. The results surprisingly showed that users found non-personalized recommendations of popular items the best match for their preferences. An analysis revealed a measurable correlation between item familiarity and user acceptance. Overall, the observations indicate that item familiarity can be a potential confounding factor in such studies and should be considered in experimental designs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Studies with users in a controlled environment are a
powerful means to assess qualities of a recommendation system
which can often not be evaluated in o ine experimental
designs. A common setup in the research literature is that
the participants of an experiment use a software tool that
implements two or more variations of a certain
recommendation functionality. After interacting with the system, the
participants are asked to explicitly evaluate certain aspects
of the system, including, e.g., the suitability or the perceived
diversity of the recommendations or other aspects like the
value of system-provided explanations [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref5">1, 2, 3, 5</xref>
        ].
      </p>
      <p>
        In the recent studies presented in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the
subjects were asked to assess the presented movie
recommendations in dimensions such as diversity, novelty or perceived
accuracy and the participants had to either evaluate lists of
recommended movies individually or make side-by-side
comparisons. One typical problem in such setups is that the
recommendation lists contain both movies that the users already
know and movies unknown to the participants. In the second
case, additional information about the movies is often
provided [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ] and users have to make their assessment based
on plot summaries or movie trailers. This situation may in
turn lead to two possible e ects. First, in case unknown
movies are displayed, the cognitive load for the participants
to assess, e.g., the suitability of the recommended movies,
is higher, which can result in a reduced overall satisfaction
with the system. Second, an assessment like \Would I enjoy
this movie?" based only on the meta-information or a trailer
could be an unreliable predictor of the assessment of a movie
after a participant has actually watched it. To our
knowledge, how item familiarity can impact the users' perception
of a recommendation system in di erent dimensions has not
been discussed explicitly in the literature before. The study
in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] does not consider item familiarity as a factor; the
authors of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] cover item familiarity in their \novelty" construct
but base it on the self-reported familiarity with the
recommendation list as a whole and do not explicitly ask users to
indicate if they know the individual movies.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>EXPERIMENT</title>
      <p>
        We conducted a user study1 in the style of [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
The participants were rst asked to rate a set of movies
known to them using a speci cally designed web application
based on MovieLens data. In the second step, they were
presented with movie recommendations created with ve di
erent algorithms including Matrix Factorization (Funk-SVD),
Bayesian Personalized Ranking (BPR), SlopeOne, a
contentbased technique (CB), and a non-personalized
popularitybased baseline (PopRank). The participants had to rate the
presented movies individually (based on meta-information)
and furthermore assessed the lists as a whole regarding
factors like diversity, transparency, or surprise. For each
presented movie, the users had to state if they already knew
the movie or not. The participants were recruited via
Mechanical Turk. From the 175 \Turkers" we ltered
unreliable ones through di erent automated and comparably strict
measures. At the end 96 participants (about 20 per
treatment) were considered as being reliable.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>OBSERVATIONS</title>
      <p>
        Accuracy. Fig. 1(a) shows how the participants
answered the question how well the presented list of movies
1Details are described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
(a)
(b)
as a whole matched their preferences; Fig. 1(b) displays the
average rating assigned to the recommended movies.
      </p>
      <p>To some surprise, the popularity-based method PopRank is
perceived by the users as the most accurate method, followed
by the BPR technique, which has a comparably strong bias
to recommend popular items to everyone. Movie
recommendations that contained only blockbusters { about 94% of the
recommendations made by these two methods were known
to the users { were considered the best preference match for
the participants (signi cant at p 0:05).</p>
      <p>In comparison, recommendation lists that were actually
personalized and contained various niche items2 received
lower scores. The preference match is inversely related to
the number of items known to the user. Funk-SVD users
knew about 50% of the items and SlopeOne users even less.</p>
      <p>We contrasted these ndings with an o ine accuracy
analysis of the underlying MovieLens dataset, which led to the
expected superiority of the Funk-SVD method in terms of
precision and the RMSE. We then computed the accuracy of
the di erent algorithms for those recommended items that
were rated by the participants. We took individual
measurements for the set of known and unknown movies for the
algorithms which did not only contain popular movies (Fig. 1).
The \o ine" accuracy measurement { except for SlopeOne {
shows to be a comparably good predictor for the movies that
the users already knew (\seen"). When applied to the
unseen movies, the predictions made by the algorithms largely
deviate from the user's ratings3.</p>
      <p>
        Diversity, Surprise, Transparency. Fig. 2 shows the
averaged questionnaire answers regarding perceived
diversity, surprise and transparency. Again, we see unexpected
results, in particular that the content-based (CB)
recommendations were perceived to be diverse. When measuring
the inverse Intra-List-Similarity (ILS) of the
recommendations using TF-IDF vectors of the movie descriptions, the
CB method as expected led to the lowest diversity, which
raises the question if the ILS measure is a suitable proxy for
perceived diversity. The surprise factor for the
popularitybiased methods was low, as expected. Finally, users felt
that they could understand the logic of the recommendations
2In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], unpopular items recommended by Funk-SVD were ltered.
3The absolute RMSE values are comparably high as the participants
only had to rate 15 items in the rst phase. The observations for
precision are comparable; RMSE values for PopRank and BPR are
missing as these methods generate no rating predictions.
1.0
0.8
0.6
0.4
0.2
0.0
1.0
0.8
0.6
0.4
0.2
0.0
      </p>
      <p>Perceived diversity</p>
      <p>Surprise</p>
      <p>Transparency
(\transparency") in particular when popular items were
presented (both in a non-personalized and personalized way).</p>
      <p>User Acceptance. Figure 3 nally reports the average
answers regarding user acceptance in terms of ease-of-use,
intention-to-reuse, and intention to recommend the system
to a friend. \Ease of use" is generally high, but users had
more trouble using the system (assigning ratings) when
unfamiliar movies were presented. The other two satisfaction
indicators in Fig. 3 are correlated with the assessment of
the preference match shown in Fig. 1.</p>
      <p>Ease of use</p>
      <p>Intention to reuse</p>
      <p>Recommendation to a friend
4.</p>
    </sec>
    <sec id="sec-4">
      <title>DISCUSSION</title>
      <p>Recommending popular and familiar items has shown to
be a well-suited strategy in this user study to achieve high
satisfaction with the system and the presented
recommendations, even though in practice recommending only popular
items is typically of limited value.</p>
      <p>Our preliminary study { experiments with more
participants, non-Turkers, and a more speci c questionnaire
focusing on item familiarity are still required { suggests that item
familiarity can be a possible confounding factor in user
studies. Speci cally, lab experiments in which users are asked to
assess items unknown to them might have limited predictive
power with respect to the true usability of the tested system.
5.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Garzotto</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Turrin</surname>
          </string-name>
          .
          <article-title>User-centric vs. system-centric evaluation of recommender systems</article-title>
          .
          <source>In Proc. INTERACT</source>
          <year>2013</year>
          , pages
          <fpage>334</fpage>
          {
          <fpage>351</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ekstrand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Harper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. A. Konstan.</surname>
          </string-name>
          <article-title>User perception of di erences in recommender algorithms</article-title>
          .
          <source>In Proc. RecSys '14</source>
          , pages
          <fpage>161</fpage>
          {
          <fpage>168</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gedikli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ge</surname>
          </string-name>
          .
          <article-title>How should I explain? A comparison of di erent explanation types for recommender systems</article-title>
          .
          <source>Int. J. Hum.-Comput</source>
          . Stud.,
          <volume>72</volume>
          (
          <issue>4</issue>
          ):
          <volume>367</volume>
          {
          <fpage>382</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lerche</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Jugovac</surname>
          </string-name>
          .
          <article-title>Item familiarity as a possible confounding factor in user-centric recommender systems evaluation. i-com Journal for Interactive Media</article-title>
          ,
          <volume>14</volume>
          (
          <issue>1</issue>
          ):
          <volume>29</volume>
          {
          <fpage>40</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Said</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fields</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Jain</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Albayrak</surname>
          </string-name>
          .
          <article-title>User-centric evaluation of a k-furthest neighbor collaborative ltering recommender algorithm</article-title>
          .
          <source>In Proc. CSCW '13</source>
          , pages
          <fpage>1399</fpage>
          {
          <fpage>1408</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>