<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Measuring the Impact of Recommender Systems - A Position Paper on Item Consumption in User Studies</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benedikt Loepp</string-name>
          <email>benedikt.loepp@uni-due.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jürgen Ziegler</string-name>
          <email>juergen.ziegler@uni-due.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Duisburg-Essen</institution>
          ,
          <addr-line>Duisburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>19</volume>
      <issue>2019</issue>
      <abstract>
        <p>While participants of recommender systems user studies usually cannot experience recommended items, it is common practice that researchers ask them to fill in questionnaires regarding the quality of systems and recommendations. While this has been shown to work well under certain circumstances, it sometimes seems not possible to assess user experience without enabling users to consume items, raising the question of whether the impact of recommender systems has always been measured adequately in past user studies. In this position paper, we aim at exploring this question by means of a literature review and at identifying aspects that need to be further investigated in terms of their influence on assessments in users studies, for instance, the diference between consumption of products or only of related information as well as the efect of domain, domain knowledge and other possibly confounding factors.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Information systems → Recommender systems.</p>
    </sec>
    <sec id="sec-2">
      <title>THE PROBLEM WITH USER STUDIES</title>
      <p>
        Questionnaires for assessing quality of recommendations and user
experience of recommender systems (RS) have been, for instance,
proposed in [
        <xref ref-type="bibr" rid="ref4 ref5 ref8">4, 5, 8</xref>
        ]. These established instruments are often
employed in academic user studies, where participants usually have to
ifrst use a RS and are subsequently asked to fill in a questionnaire.
However, recommended items in these scenarios are almost
always represented through “proxy presentations”, i.e. items are only
shown to users by means of images, descriptive texts, metadata,
etc. The actual consumption of items is in contrast to real-world
situations rarely possible. There, it is mostly required to have, for
instance, bought a product, visited a hotel, or watched a movie,
before being even able to provide an opinion.
      </p>
      <p>
        Previously, we have investigated whether the consumption of
items during user studies has an impact on the succeeding
assessment of recommendations by means of questionnaires [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In other
studies, e.g. on explanations [
        <xref ref-type="bibr" rid="ref1 ref9">1, 9</xref>
        ], the impact of consumption has
never been directly addressed. However, we found, among others,
that it strongly depends on domain as well as type and amount
of presented information whether it is possible for participants
to adequately assess recommendation quality and aspects related
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>LITERATURE REVIEW</title>
      <p>
        First, for putting our findings from [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] more into context, we have
performed a literature review. We analyzed all 46 papers accepted
to the five editions of the Joint Workshop on Interfaces and Human
Decision Making for Recommender Systems (IntRS)1 which were
held from 2014 to 2018 in conjunction with the ACM Conference
on Recommender Systems (RecSys). In 66 % of these papers, a user
study had been reported (there were a few more, which however
did not focus on recommendation issues but e.g. “only” on
comparing diferent interfaces). In some of the papers without a user
study, applying such an evaluation method would not have been
appropriate for investigating the respective research question (or
even impossible). Accordingly, this number seems actually quite
high, especially considering that user studies are still rarely used
in broader recommender research [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Taking a closer look at the procedure of the user studies however
sheds a bit diferent light: As far as we were able to grasp the details
from the papers, it was possible only in 44 % of the reported user
studies to actually consume products (i.e. in 30 % of all papers; see
Figure 1). Admittedly, in some papers, this would have made no
sense or consumption would have been unrealistic (e.g. hotel or
date recommendations). Sometimes, it simply was not necessary
for answering the underlying research question. We decided not
to count in consumption of movie trailers [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or song excerpts
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], but included cases where, for instance, recommended research
papers were only accessible via a link to an external website [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
making it less likely that many participants took that chance. In
summary, with the smaller number of user studies presented at
1Website of this year’s edition: https://intrs19.wordpress.com/
11
2
5
      </p>
      <p>Accepted papers ...</p>
      <p>without user studies
with user studies
with consumption
0 20314 20315 20316 201317 20418
Figure 1: Results from our literature review showing how
many papers were accepted to past IntRS workshop editions,
how many of these papers contained user studies, and in
how many studies item consumption was possible.
less user-centric venues in mind that likely allowed consuming
items in even fewer cases, the question arises whether evaluation
results would have been overall the same if item consumption had
been possible. While our literature review is indeed limited, the
impact of RS has most likely not always been measured accurately
since participants might not have had everything they needed to
adequately assess recommendation quality and user experience.
3</p>
    </sec>
    <sec id="sec-4">
      <title>ASPECTS TO INVESTIGATE</title>
      <p>With the importance of item consumption in mind, the literature
review points out possible omissions in past research, emphasizing
the need to take this aspect more into account when designing
future user experiments. For doing so, a number of research
questions still need to be answered. This may help to decide in a more
structured manner, for example, whether it is necessary to provide
participants with the possibility to consume items at all, or which
substitutes may be used otherwise. It may also indicate which
factors that might confound the assessment and thus lead to a distorted
impression of the recommender’s impact need to be considered—for
planning the study, analyzing results and drawing conclusions.</p>
      <p>The following (non-exhaustive) list contains aspects we think
are generally important and possibly mediate the efects of item
consumption. Concretely, we suggest investigating the influence:
• of item consumption also in other domains, depending on
domain knowledge of participants as well as product type
and attributes (e.g. search vs. experience products),
• of presenting diferent kinds of information (subjective vs.
objective item descriptions) as possible substitutes for item
consumption at varying level of detail (only metadata or
additional content descriptions, other item-related information
such as user reviews, system-generated explanations, paper
abstracts, song excerpts, movie trailers, etc.),
• of user characteristics such as personality or decision-making
style (making decisions in either a rational or intuitive way
might afect the need for actual item consumption),
• and of the point in time assessments take place (since the
efect of item consumption might diminish over time).</p>
      <p>
        Beyond that, there are certainly many other aspects that may
influence study results when trying to quantify the impact of RS.
For instance, the improvements made regarding user experience in
the past couple of years led to higher perceived recommendation
quality without any changes to recommendations [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However,
apart from these attempts intended to positively afect the impact of
RS, some aspects may unintentionally cause diferences due to the
specific characteristics of user experiments. First, the experimental
situation itself (e.g. presence of supervisor, lab study), with systems
specifically designed for the purpose of the study, thus also limited
to this purpose, might afect ecological validity: The assessment
might be diferent compared to when a recommendation set is
integrated into a real-world e-commerce platform. Among others,
economic reasons (real money needs to be spent) or diferent
relationships between students and researchers vs. customers and
commercial system providers might afect e.g. reported purchase
intention or perceived trustworthiness. Also, questionnaires might
interfere with internal validity as item formulations can be
ambiguous, e.g. regarding whether recommended products are novel
(recently released vs. only new to the participant) or the set appears
well-chosen (because products fit together or actually represent the
participant’s taste). More generally, using such instruments at all
might be an issue as they might provoke a more conscious
assessment, possibly afecting decision-making (i.e. participants could
settle for diferent items if not confronted with a questionnaire).
4
      </p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS AND PERSPECTIVES</title>
      <p>
        We have now positioned our work on the efects of item
consumption [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] in context of the broader question of how the impact of RS
can adequately be measured by means of user studies in academia.
We identified a number of aspects that still need to be investigated
in order to pursue the superordinate goal of deriving a set of
guidelines for promoting validity of future experiments and fostering
reproducibility. Currently, we are planning a study to investigate
the influence of the aspects listed in the previous section. In
addition, we would like to address the questions beyond and encourage
others to do so as well, possibly also by employing unprecedented
means for assessing the impact of RS. For instance, developing
methods that use eye-tracking to determine which items are of
most interest for participants might help to avoid interventions
and make them switch decision-making styles. Overall, the insights
that may be gained could also have broader impact, for example,
by finding solutions for algorithms to adequately deal both with
ratings provided in real-world systems without previously
experiencing the products (e.g. “This recipe sounds awesome” → 5-star
rating) and ratings resulting from actual consumption.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bilgic</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Mooney</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Explaining recommendations: Satisfaction vs. promotion</article-title>
          .
          <source>In Proc. Beyond Personalization Workshop.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Graus</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Can trailers help to alleviate popularity bias in choice-based preference elicitation?</article-title>
          .
          <source>In Proc. IntRS</source>
          '
          <volume>16</volume>
          .
          <fpage>22</fpage>
          -
          <lpage>27</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          .
          <year>2015</year>
          . Recommender Systems Handbook. Springer US,
          <article-title>Chapter Evaluating recommender systems with user experiments</article-title>
          ,
          <fpage>309</fpage>
          -
          <lpage>352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gantner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Soncu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Newell</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Explaining the user experience of recommender systems</article-title>
          .
          <source>User Model. User-Adap</source>
          .
          <volume>22</volume>
          ,
          <issue>4</issue>
          -
          <fpage>5</fpage>
          (
          <year>2012</year>
          ),
          <fpage>441</fpage>
          -
          <lpage>504</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Kobsa</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A pragmatic procedure to support the user-centric evaluation of recommender systems</article-title>
          .
          <source>In Proc. RecSys '11. ACM</source>
          , New York, NY, USA,
          <fpage>321</fpage>
          -
          <lpage>324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Loepp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Donkers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kleemann</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Impact of item consumption on assessment of recommendations in user studies</article-title>
          .
          <source>In Proc. RecSys '18. ACM</source>
          , New York, NY, USA,
          <fpage>49</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Lu</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Tintarev</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A diversity adjusting strategy with personality for music recommendation</article-title>
          .
          <source>In Proc. IntRS '18</source>
          .
          <fpage>7</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Hu</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>A user-centric evaluation framework for recommender systems</article-title>
          .
          <source>In Proc. RecSys '11. ACM</source>
          , New York, NY, USA,
          <fpage>157</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Tintarev</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Masthof</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Evaluating the efectiveness of explanations for recommender systems</article-title>
          .
          <source>User Model. User-Adap</source>
          .
          <volume>22</volume>
          ,
          <issue>4</issue>
          -
          <fpage>5</fpage>
          (
          <year>2012</year>
          ),
          <fpage>399</fpage>
          -
          <lpage>439</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Verbert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Parra</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Brusilovsky</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The efect of diferent set-based visualizations on user exploration of recommendations</article-title>
          .
          <source>In Proc. IntRS</source>
          '
          <volume>14</volume>
          .
          <fpage>37</fpage>
          -
          <lpage>44</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>