<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recommender Systems Alone Are Not Everything: Towards a Broader Perspective in the Evaluation of Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benedikt Loepp</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Duisburg-Essen</institution>
          ,
          <addr-line>Duisburg</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Thus far, in most of the user experiments conducted in the area of recommender systems, the respective system is considered as an isolated component, i.e., participants can only interact with the recommender that is under investigation. This fails to recognize the situation of users in real-world settings, where the recommender usually represents only one part of a greater system, with many other options for users to ifnd suitable items than using the mechanisms that are part of the recommender, e.g., liking, rating, or critiquing. For example, in current web applications, users can often choose from a wide range of decision aids, from text-based search over faceted filtering to intelligent conversational agents. This variety of methods, which may equally support users in their decision making, raises the question of whether the current practice in recommender evaluation is suficient to fully capture the user experience. In this position paper, we discuss the need to take a broader perspective in future evaluations of recommender systems, and raise awareness for evaluation methods which we think may help to achieve this goal, but have not yet gained the attention they deserve.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>are supported in a holistic fashion: Even if multiple decision aids are available, they are often
isolated, i.e., the input provided to one component hardly afects the results generated by
another. For instance, when a user constrains the set of items by selecting several filter criteria,
a recommender that is part of the same system will not necessarily reflect this selection in
the generated recommendations, but only take the user’s implicit or explicit item feedback
into account (and vice versa). Similarly, the interaction with a chatbot will usually start from
nothing, ignoring any other interests or needs that have been expressed before.</p>
      <p>This separation is dificult to understand in light of the fact that it is well known that diferent
decision aids contribute diferently to the user’s progress in accomplishing typical choice
and decision-making tasks [8, 9]. Recent studies on systems that provide multiple support
components have confirmed that users use diferent mechanisms before settling on an item,
and that this is (partly) due to personal and situational characteristics [10, 11, 12, 13]. For
these reasons, there have, of course, been several calls over the years to bring the methods and
their fields closer together [ 14, 15, 16, 17]. These calls, however, have largely been focused on
methodological aspects, whereas it should not be overseen, that this narrow perspective is also
reflected in the evaluation of the methods—regardless of whether they are tightly integrated
with other decision aids, or appear in a decoupled fashion, as it is currently common.</p>
      <p>The work presented in [18] is one of the few exceptions from information retrieval research
that argues for a more holistic evaluation approach, in particular, “to think outside the (search)
box.” In general, however, this aspect has, to the best of our knowledge, not yet received
much attention, neither in the area of information retrieval, nor conversational user interfaces,
nor recommender systems [cf. 19, 20, 1, 21]. This means that whenever user experience of
recommender systems is studied in empirical experiments, participants can typically only
interact with the recommendation component, i.e., the subject of the evaluation, which is
often designed specifically for the purpose of the study, especially in academia. Switching to
a diferent method, which participants may find more appealing depending on their personal
preferences and progress in the decision-making process, is, however, not possible. Hence, there
is a lack of data to indicate which method works best for which user. As a consequence, we
argue that it is necessary for all kinds of decision aids “to think outside the box.”</p>
      <p>In the following, as another motivating example, we present qualitative feedback from
a user experiment on multiple decision aids that we recently conducted. Afterwards, we
provide an overview of methods that may help to apply a broader perspective when evaluating
recommender systems, and thus, to obtain a more accurate picture of real-world scenarios, in
which these systems usually represent only one out of many available decision aids.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Motivating example and potential future directions</title>
      <p>Because of the reasons described above, we argue that a broader perspective is required when
evaluating recommender systems. Otherwise, there is the danger of continuing to follow certain
paths in the development of interactive and conversational recommendation approaches (cf. the
surveys in [22, 23]), without knowing whether these approaches are really what users want.
To highlight that this is an issue already, we refer to a recent experiment, in which we asked
participants ( = 100, 47 females, 2 non-binary, age: M = 35.36, SD = 12.60) to use diferent
decision aids. For this purpose, we confronted them with one of two tasks, either related to
a goal-driven or an explorative scenario. In both scenarios, participants had to find at least
two suitable laptops. To accomplish the respective task, they were allowed to choose (and to
switch) between: a faceted filtering component, a content-based recommender with the option
to like and dislike items, a product advisor with a dialog showing a limited number of guiding
questions, and a natural language chatbot implemented using Google Dialogflow .</p>
      <sec id="sec-2-1">
        <title>2.1. Qualitative insights from an experiment with multiple decision aids</title>
        <p>While the study is described in more detail in [13], we here want to provide additional comments
made by participants when they were asked why they did not use one of the components. For
example, this was the case because they were reluctant to use chatbots. One participant stated:
“I generally do not like interacting with chatbots. I feel whatever input I give to select a product,
I might as well use a filter component.” In contrast to the increasing popularity of conversational
agents for recommendation purposes, another participant indicated that he or she uses chatbots
only when there is “a problem with the product or a payment issue, [but that a] chatbot is
not required [for] browsing.” Others were even more direct in their criticism, writing: “I hate
chatbots. I feel I spend a lot of time typing, and I hate the fake cheeriness of them,” or that they
turn “something really simple like searching for a new laptop into a dark comedy.”</p>
        <p>Other participants cited personal characteristics as reasons for their concerns, such as domain
knowledge (“I assume the chatbot would be more useful for someone who does not know what
to look for.”) or need for control (“I prefer doing things myself, I can chat to the bot when I cannot
seem to find what I am looking for using other [. . . ] options.”) For the other options, however,
we obtained similar feedback. For example, with respect to the advisor, one participant stated
that he or she “should have tried this component, but [likes] to shop for these products using
iflters.” Although it was clearly an academic study, another participant refrained from using
the recommender because he or she suspected “that this component uses sponsored companies
which pay money to have their products included in the recommendation section.” Also in this
case, others were generally reluctant, writing: “I never use recommendations as they are never
in line with what I need. It may be the bestseller for the store in question, but not for my needs.”
Domain knowledge again played a role, notably in both directions, with participants stating to
be “knowledgeable enough [that they] do not need recommendations,” but also that they “do
not know enough about computers to give a thumbs up or thumbs down.” A lack of knowledge
was also a major reason to stay away from faceted filtering , of which participants “thought you
need technical understanding of laptops,” or mentioned that they “often feel overwhelmed using
iflters, and generally do not know where to start in regard to buying laptops.”</p>
        <p>All these comments suggest that users often have an idea of which decision aid to use.
However, this is not necessarily the one that is ofered by a system—or is the subject of the
experiment they participate in. For a holistic user-centered evaluation, this means, that the
perspective is too narrow, as only the user experience with the specific recommender for a
given task is addressed, ignoring the interdependencies with other methods that may exist, and
may be more suitable depending on the current circumstances. Some participants explicitly
indicated that they would like to “specify certain filters and let the recommendations pick only
from the result set of the filters,” or to “combine components [so that] the advisor/chatbot works
within the result set created by the must-have filters.” However, this is exactly what participants
usually cannot do, since most experiments are strongly focused on individual decision aids. To
evaluate recommender systems from a broader perspective, we therefore suggest to improve
the current practice by applying the following evaluation methods, which, to date, are used
only for single methods, or not at all.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Ofline experiments and simulation studies: Richer data, entire systems</title>
        <p>
          Ofline experiments are well established in recommender research [24]. However, they are
increasingly criticized as they do not allow obtaining insights into the quality dimensions
that are relevant from a user perspective [
          <xref ref-type="bibr" rid="ref1">1, 25, 26</xref>
          ]. Nevertheless, with the large datasets that
are available today (e.g., MovieLens, Netflix , Amazon), they remain essential to make objective
decisions whether or not to use a specific recommendation method. However, most datasets
are limited to implicit or explicit user-item feedback. Even though they represent diferent
domains, contain a varying amount of side information, and are nowadays often available
in sequential form, this limits what can be concluded from the corresponding experiments,
i.e., which algorithm generates better item recommendations. Therefore, we argue that future
ofline experiments should be conducted based on richer datasets, which include data from
all components of a system, such as both user-item feedback and search queries issued. Thus,
provided adequate metrics are found, one could also determine which decision aids work best
for individual users and keep them longer engaged. More generally, one could examine system
support above the item level, e.g., with respect to the objective quality of recommendations for
item features, or even of recommendations for switching to other decision aids.
        </p>
        <p>The same applies to simulation studies, which have only recently gained more attention in
recommender research [27]. By simulating typical user behavior, e.g., with respect to critiquing
mechanisms [28] or interaction with items over time [29], this type of experiment has shown
strong potential in economic terms. With multiple decision aids, and thus, a larger design
space, this will become even more important, also on a global level, e.g., to study long-term user
behavior with respect to the question of when a recommender is used, or another component is
perceived as more suitable and may contribute more to the user’s progress. In addition, domain
and other factors such as product type (search vs. experience), product category (cheap streaming
content vs. expensive high-risk products), and the given task (goal-oriented or explorative) may
afect which method works best. Thus, simulation studies may be the only way to investigate
how preferences evolve over time when using diferent decision aids, and to understand which
interaction efects can occur between a recommender and other components—something that
would never be possible with actual users.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. User-centered evaluation: Multiple decision aids, insightful methods</title>
        <p>
          Well-known qualitative methods from human-computer interaction research, which are
frequently used at the beginning of user-centered design processes, e.g., focus groups, interviews,
or contextual inquiries [cf. 30], are rarely used in the area of recommender systems. We argue,
however, that these techniques could be useful for obtaining insights into users’ actual needs
with respect to the interaction with such a system, in particular, when it comes to the relation to
other decision aids. In this context, it is worth noting that it has been found only recently, that
users’ mental models do not necessarily correspond to the implementations of recommender
systems, and are subject to large inter-individual diferences [ 31]. However, identifying the
understanding users have of the system behavior is considered highly important for evaluating
the impact of a recommender and improving it [32]. Accordingly, to better inform the design
of applications that embed multiple support components, it will be inevitable to explore these
models in more depth, in particular, with a focus on the users’ comprehension of possible
interactions between the components—by means of both qualitative methods such as grounded
theory [33] and quantitative approaches such as proposed in [
          <xref ref-type="bibr" rid="ref2">34</xref>
          ].
        </p>
        <p>
          Once a (prototypical) recommender is implemented, questionnaire-based assessment is the
most common way of measuring the diferent qualities related to user experience. For this
purpose, well-established frameworks and questionnaires exist [e.g., 35, 36, 37], which, however,
are strongly focused on dimensions that are specific to recommender systems. On the other hand,
general usability questionnaires, e.g., SUS [
          <xref ref-type="bibr" rid="ref6">38</xref>
          ] and UEQ [
          <xref ref-type="bibr" rid="ref7">39</xref>
          ], are too broad to draw conclusions
about the suitability of this or other decision aids for preferential choice and decision-making
tasks. Therefore, existing instruments need to be extended in order to allow for a more global,
subjective assessment of recommendation components—in the context of the applications in
which they are embedded, and thus, of the interplay with other methods. Otherwise, it will be
hardly possible to gain insights into why users prefer a specific method to perform a certain task,
and, more generally, whether they would use more strongly connected decision aids, or dislike
this idea, e.g., because they expect higher complexity or have increasing privacy concerns.
        </p>
        <p>
          Finally, we would like to highlight two further aspects that do not receive much attention
in current recommender research: in-situ and long-term evaluation. The former is important,
among others, because questionnaires sufer from the problem that they typically require self
reflection disconnected from actual system usage, and, worse, often from consumption or
experience of the items, which has shown to influence the assessment of recommendations [
          <xref ref-type="bibr" rid="ref8">40</xref>
          ].
Therefore, we deem it necessary to develop methods for a quantitative in-situ assessment of the
users’ motivation to use diferent components, e.g., via questionnaires directly embedded into
the respective applications, as it has been done to study the reasons to switch between search
engines [
          <xref ref-type="bibr" rid="ref9">41</xref>
          ] or selected decision aids [13]. This, however, may need to be complemented by
qualitative methods, such as the think aloud procedure, systematic user observation, or shadowing—
techniques that are generally popular but have also rarely been applied in recommender research.
Moreover, eye tracking may be considered as a useful alternative. Being less disruptive, it
has become more popular in recent years, e.g., in studies on recommender interfaces and
recommendation presentation [
          <xref ref-type="bibr" rid="ref10 ref11">42, 43</xref>
          ], critiquing [
          <xref ref-type="bibr" rid="ref12 ref13">44, 45</xref>
          ], and efects of personal characteristics
[
          <xref ref-type="bibr" rid="ref14">46</xref>
          ]. However, also in these cases, participants’ behavior was observed only in relation to the
recommendation component, largely ignoring its surroundings. Either way, following these
directions will not be suficient to understand how users interact with these surroundings over
a longer period of time. Therefore, while repeated study designs, longitudinal studies, and field
studies are still rare in recommender research, with only very few exceptions [e.g., 47, 48],
these methods appear to be of particular importance when taking a broader perspective: Goals
and tasks may vary over time, which can have a substantial impact on the usage of diferent
components. Thus, experimental data that represent long-term user behavior only with respect
to a single decision aid may distort the picture, since other components may be perceived as
more appropriate in other contexts or stages of the decision-making process.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Conclusions</title>
      <p>Overall, it seems important for future research in the recommender area, but also for other
communities, to face the challenge of evaluating the respective methods in a context that is
more similar to real-world settings, where decision aids rarely stand on their own. In this
position paper, we explained why we think this way, and outlined how this challenge may be
addressed. By this means, we hope to raise awareness that contemporary decision aids not only
need to be brought together from a methodological perspective, but that a broader perspective
is required when evaluating the methods. Of course, this may open up new issues. For instance,
the more holistic the evaluation, the higher efort and costs for running an experiment, which
are factors that already limit many studies in academia. Moreover, the dificulty of designing an
experiment and analyzing its results, but also, the number of application-specific parameters
and possible confounding factors, increase with the consideration of more than one decision aid.
For these reasons, it will remain important to keep in mind the specific circumstances in which
most experiments take place, and not to think that a broader perspective in the evaluation of
the methods automatically allows more general conclusions to be drawn.</p>
      <p>Nevertheless, we are certain that “thinking outside the (recommendations) box” will help
gain a better understanding not only of the degree to which a recommender can satisfy a
user in a given situation, but, in particular, of how the interplay with other decision aids can
afect the assessment of the system. In the end, this may shift the focus away from further
improving decision aids that are less efective or users do not want to use for specific tasks,
to those that exhibit the greatest potential for providing support at the respective stage of the
decision-making process. For now, however, we hope to encourage at least a discussion about
using the mentioned evaluation methods more extensively to gain more in-depth insights into
the users’ understanding of and preference for recommendation components in relation to other
decision aids. Of course, other methods may equally well be used, but we leave it for future
work to provide more concrete suggestions on which methods to use and in which order.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>Thanks to Timm Kleemann, who implemented the system for the study that is mentioned here,
and contributed to this study to the same extent as the author of the present paper.
The study was partially supported by the Eurostars project ACODA (grant no. 01QE1946C).
[2] C. A. Gomez-Uribe, N. Hunt, The Netflix recommender system: Algorithms, business
value, and innovation, ACM Transactions on Management Information Systems 6 (2015)
13:1–13:19.
[3] B. Smith, G. Linden, Two decades of recommender systems at Amazon.com, IEEE Internet</p>
      <p>Computing 21 (2017) 12–18.
[4] R. Baeza-Yates, B. Ribeiro-Neto, Modern Information Retrieval, ACM, New York, NY, USA,
1999.
[5] M. A. Hearst, Search User Interfaces, Cambridge University Press, Cambridge, UK, 2009.
[6] K. Ramesh, S. Ravishankaran, A. Joshi, K. Chandrasekaran, A survey of design techniques
for conversational agents, in: ICICCT ’17: Proceedings of the 2nd International
Conference on Information, Communication and Computing Technology, Springer Singapore,
Singapore, 2017, pp. 336–350.
[7] A. Anand, L. Cavedon, H. Joho, M. Sanderson, B. Stein, Conversational search (Dagstuhl</p>
      <p>Seminar 19461), Dagstuhl Reports 9 (2020) 34–83.
[8] G. Häubl, V. Trifts, Consumer decision making in online shopping environments: The
efects of interactive decision aids, Marketing Science 19 (2000) 4–21.
[9] S. Castagnos, N. Jones, P. Pu, Recommenders’ influence on buyers’ decision process, in:
RecSys ’09: Proceedings of the 3rd ACM Conference on Recommender Systems, ACM,
New York, NY, USA, 2009, pp. 361–364.
[10] J. Schafer, J. Humann, J. O’Donovan, T. Höllerer, Quantitative modeling of dynamic
human-agent cognition, in: M. D. McNeese, E. Salas, M. R. Endsley (Eds.), Contemporary
Research: Models, Methodologies, and Measures in Distributed Team Cognition, CRC
Press, Boca Raton, FL, USA, 2020, pp. 137–186.
[11] P. Virdi, A. D. Kalro, D. Sharma, Online decision aids: The role of decision-making styles
and decision-making stages, International Journal of Retail &amp; Distribution Management
48 (2020) 555–574.
[12] T. Kleemann, M. Wagner, B. Loepp, J. Ziegler, Modeling user interaction at the convergence
of filtering mechanisms, recommender algorithms and advisory components, in: Mensch
&amp; Computer 2021 – Tagungsband, ACM, New York, NY, USA, 2021, pp. 531–543.
[13] T. Kleemann, B. Loepp, J. Ziegler, Towards multi-method support for product search and
recommending, in: UMAP ’22: Adjunct Proceedings of the 30th ACM Conference on User
Modeling, Adaptation and Personalization, ACM, New York, NY, USA, 2022, pp. 74–79.
[14] H. Garcia-Molina, G. Koutrika, A. Parameswaran, Information seeking: Convergence
of search, recommendations, and advertising, Communications of the ACM 54 (2011)
121–130.
[15] E. H. Chi, Blurring of the boundary between interactive search and recommendation, in:
IUI ’15: Proceedings of the 20th International Conference on Intelligent User Interfaces,
ACM, New York, NY, USA, 2015, p. 2.
[16] B. Loepp, On the convergence of intelligent decision aids, in: UCAI ’21: Proceedings of
the 2nd Workshop on User-Centered Artificial Intelligence, 2021.
[17] A. D. Starke, M. Lee, Unifying recommender systems and conversational user interfaces,
in: CUI ’22: Proceedings of the 4th International Conference on Conversational User
Interfaces, ACM, New York, NY, USA, 2022.
[18] P. Clough, Evaluation: Thinking outside the (search) box, in: FIRE ’14: Proceedings of the</p>
      <p>Forum for Information Retrieval Evaluation, ACM, New York, NY, USA, 2015, pp. 1–9.
[19] D. Kelly, Methods for evaluating interactive information retrieval systems with users,</p>
      <p>Foundations and Trends in Information Retrieval 3 (2009) 1–224.
[20] C. Mulwa, S. Lawless, M. Sharp, V. Wade, The evaluation of adaptive and personalised
information retrieval systems: A review, International Journal of Knowledge and Web
Intelligence 2 (2011) 138–156.
[21] A. B. Kocaballi, L. Laranjo, E. Coiera, Understanding and measuring user experience in
conversational interfaces, Interacting with Computers 31 (2019) 192–207.
[22] M. Jugovac, D. Jannach, Interacting with recommenders – Overview and research
directions, ACM Transactions on Interactive Intelligent Systems 7 (2017) 10:1–10:46.
[23] D. Jannach, A. Manzoor, W. Cai, L. Chen, A survey on conversational recommender
systems, ACM Computing Surveys 54 (2022) 105:1–105:36.
[24] A. Gunawardana, G. Shani, S. Yogev, Evaluating recommender systems, in: F. Ricci,
L. Rokach, B. Shapira (Eds.), Recommender Systems Handbook, Springer, New York, NY,
USA, 2022, pp. 547–601.
[25] M. Rossetti, F. Stella, M. Zanker, Contrasting ofline and online results when evaluating
recommendation algorithms, in: RecSys ’16: Proceedings of the 10th ACM Conference on
Recommender Systems, ACM, New York, NY, USA, 2016, pp. 31–34.
[26] T. Rehorek, O. Biza, R. Bartyzal, P. Kordik, I. Povalyev, O. Podstavek, Comparing ofline
and online evaluation results of recommender systems, in: REVEAL ’18: Proceedings of
the Workshop on Ofline Evaluation for Recommender Systems, 2018.
[27] N. Hazrati, F. Ricci, Simulating users’ interactions with recommender systems, in: Adjunct
Proceedings of the 30th ACM Conference on User Modeling, Adaptation and
Personalization, ACM, New York, NY, USA, 2022, pp. 95–98.
[28] H. Xie, D. D. Wang, Y. Rao, T.-L. Wong, L. Y. K. Raymond, L. Chen, F. L. Wang, Incorporating
user experience into critiquing-based recommender systems: A collaborative approach
based on compound critiquing, International Journal of Machine Learning and Cybernetics
9 (2018) 837–852.
[29] J. McInerney, E. Elahi, J. Basilico, Y. Raimond, T. Jebara, Accordion: A trainable simulator
for long-term interactive systems, in: RecSys ’21: Proceedings of the 15th ACM Conference
on Recommender Systems, ACM, New York, NY, USA, 2021, pp. 102–113.
[30] G. Cockton, Usability evaluation, The Encyclopedia of Human-Computer Interaction (2nd</p>
      <p>Edition) (2013).
[31] T. Ngo, J. Kunkel, J. Ziegler, Exploring mental models for transparent and controllable
recommender systems: A qualitative study, in: UMAP ’20: Proceedings of the 28th ACM
Conference on User Modeling, Adaptation and Personalization, ACM, New York, NY, USA,
2020, pp. 183–191.
[32] M. M. Ghori, A. Dehpanah, J. Gemmell, H. Qahri-Saremi, B. Mobasher, Does the user have
a theory of the recommender? A grounded theory study, in: Adjunct Proceedings of the
30th ACM Conference on User Modeling, Adaptation and Personalization, ACM, New
York, NY, USA, 2022, pp. 167–174.
[33] J. Corbin, A. Strauss, Basics of Qualitative Research: Techniques and Procedures for
Developing Grounded Theory, 3 ed., Sage Publications, Inc., Thousand Oaks, CA, USA,
2008.
USA, 2020, pp. 173–182.
[47] Y. Zhong, T. L. S. Menezes, V. Kumar, Q. Zhao, F. M. Harper, A field study of related video
recommendations: Newest, most similar, or most relevant?, in: Proceedings of the 12th
ACM Conference on Recommender Systems, ACM, New York, NY, USA, 2018, pp. 274–278.
[48] Y. Liang, M. C. Willemsen, Exploring the longitudinal efects of nudging on users’ music
genre exploration behavior and listening preferences, in: RecSys ’22: Proceedings of the
16th ACM Conference on Recommender Systems, ACM, New York, NY, USA, to appear.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          ,
          <article-title>Evaluating recommender systems with user experiments</article-title>
          , in: F.
          <string-name>
            <surname>Ricci</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Rokach</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          Shapira (Eds.),
          <source>Recommender Systems Handbook</source>
          , Springer US, Boston, MA, USA,
          <year>2015</year>
          , pp.
          <fpage>309</fpage>
          -
          <lpage>352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kunkel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ngo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Krämer</surname>
          </string-name>
          ,
          <article-title>Identifying group-specific mental models of recommender systems: A novel quantitative approach</article-title>
          , in: C. Ardito,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lanzilotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Malizia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Petrie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Piccinno</surname>
          </string-name>
          , G. Desolda,
          <string-name>
            <given-names>K.</given-names>
            Inkpen (Eds.),
            <surname>Human-Computer</surname>
          </string-name>
          <string-name>
            <surname>Interaction - INTERACT</surname>
          </string-name>
          <year>2021</year>
          , volume
          <volume>12935</volume>
          of Lecture Notes in Computer Science, Springer, Berlin, Germany,
          <year>2021</year>
          , pp.
          <fpage>383</fpage>
          -
          <lpage>404</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kobsa</surname>
          </string-name>
          ,
          <article-title>A pragmatic procedure to support the user-centric evaluation of recommender systems</article-title>
          ,
          <source>in: RecSys '11: Proceedings of the 5th ACM Conference on Recommender Systems</source>
          , ACM, New York, NY, USA,
          <year>2011</year>
          , pp.
          <fpage>321</fpage>
          -
          <lpage>324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>P.</given-names>
            <surname>Pu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <article-title>A user-centric evaluation framework for recommender systems</article-title>
          ,
          <source>in: RecSys '11: Proceedings of the 5th ACM Conference on Recommender Systems</source>
          , ACM, New York, NY, USA,
          <year>2011</year>
          , pp.
          <fpage>157</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Knijnenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Willemsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gantner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Soncu</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Newell, Explaining the user experience of recommender systems, User Modeling and User-Adapted Interaction 22 (</article-title>
          <year>2012</year>
          )
          <fpage>441</fpage>
          -
          <lpage>504</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Brooke</surname>
          </string-name>
          , SUS
          <article-title>- A quick and dirty usability scale</article-title>
          , in: Usability Evaluation in Industry, Taylor &amp; Francis, London, UK,
          <year>1996</year>
          , pp.
          <fpage>189</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>B.</given-names>
            <surname>Laugwitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Held</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schrepp</surname>
          </string-name>
          ,
          <article-title>Construction and evaluation of a user experience questionnaire</article-title>
          , in: A.
          <string-name>
            <surname>Holzinger</surname>
          </string-name>
          (Ed.),
          <source>HCI and Usability for Education and Work</source>
          , volume
          <volume>5298</volume>
          of Lecture Notes in Computer Science, Springer, Berlin, Germany,
          <year>2008</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>B.</given-names>
            <surname>Loepp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Donkers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kleemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <article-title>Impact of item consumption on assessment of recommendations in user studies</article-title>
          ,
          <source>in: RecSys '18: Proceedings of the 12th ACM Conference on Recommender Systems</source>
          , ACM, New York, NY, USA,
          <year>2018</year>
          , pp.
          <fpage>49</fpage>
          -
          <lpage>53</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Dumais</surname>
          </string-name>
          ,
          <article-title>Why searchers switch: Understanding and predicting engine switching rationales</article-title>
          ,
          <source>in: SIGIR '11: Proceedings of the 34th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2011</year>
          , pp.
          <fpage>335</fpage>
          -
          <lpage>344</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. M.</given-names>
            <surname>Harper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <article-title>Gaze prediction for recommender systems</article-title>
          ,
          <source>in: RecSys '16: Proceedings of the 10th ACM Conference on Recommender Systems</source>
          , ACM, New York, NY, USA,
          <year>2016</year>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>P.</given-names>
            <surname>Gaspar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kompan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Simko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bielikova</surname>
          </string-name>
          ,
          <article-title>Analysis of user behavior in interfaces with recommended items - An eye-tracking study</article-title>
          ,
          <source>in: IntRS '18: Proceedings of the 5th Joint Workshop on Interfaces and Human Decision Making for Recommender Systems</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>32</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>An eye-tracking study: Implication to implicit critiquing feedback elicitation in recommender systems</article-title>
          ,
          <source>in: UMAP '16: Proceedings of the 24th ACM Conference on User Modeling, Adaptation and Personalization</source>
          , ACM, New York, NY, USA,
          <year>2016</year>
          , pp.
          <fpage>163</fpage>
          -
          <lpage>167</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          , W. Wu,
          <article-title>Inferring users' critiquing feedback on recommendations from eye movements</article-title>
          , in: A.
          <string-name>
            <surname>Goel</surname>
          </string-name>
          , M. B.
          <string-name>
            <surname>Díaz-Agudo</surname>
          </string-name>
          , T. Roth-Berghofer (Eds.),
          <source>Case-Based Reasoning Research and Development</source>
          , volume
          <volume>9969</volume>
          of Lecture Notes in Computer Science, Springer, Berlin, Germany,
          <year>2016</year>
          , pp.
          <fpage>62</fpage>
          -
          <lpage>76</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>M.</given-names>
            <surname>Millecamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. N.</given-names>
            <surname>Htun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Conati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Verbert</surname>
          </string-name>
          ,
          <article-title>What's in a user? Towards personalising transparency for music recommender interfaces</article-title>
          ,
          <source>in: UMAP '20: Proceedings of the 28th ACM Conference on User Modeling, Adaptation and Personalization</source>
          , ACM, New York, NY,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>