<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>F. Diaz)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>On Evaluating Session-Based Recommendation with Implicit Feedback</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fernando Diaz</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Google</institution>
          ,
          <addr-line>Montréal</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>Session-based recommendation systems are used in environments where system recommendation actions are interleaved with user choice reactions. Domains include radio-style song recommendation, session-aware related-items in a shopping context, and next video recommendation. In many situations, interactions logged from a production policy can be used to train and evaluate such session-based recommendation systems. This paper presents several concerns with interpreting logged interactions as reflecting user preferences and provides possible mitigation to those concerns.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;session-based recommendation systems</kwd>
        <kwd>evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>interface constraints, user cognitive biases, and uncontrolled sequential dependencies. As
a result, we believe that evaluation data gathered under current practices may be distorted
and susceptible to misidentification of quality recommendation systems, even when existing
debiasing methods are employed.</p>
      <p>We propose addressing these concerns in several ways. First, we suggest that current
rewarddefinition practices should be revisited and user behavior studied further. As part of this, we
believe that there is an opportunity to adjust logging and interaction modeling practices to
control for problematic behavior. Second, we suggest that system designers develop mechanisms
and interface tools to gather evaluation data not prone to the biases discussed in Section 3.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Session-based Recommender Systems</title>
      <p>Let  be the set of users and  the set of items. We define a session prefix  = (︀ 1, 2,  1
− ⌋︀
as a sequence of items engaged with by a user  ∈  . Given a session prefix, each item
has an associated reward  reflecting the quality of the item, if we were to present it to user

immediately after  . The next-item recommendation task is to, given a session prefix, produce
we can define  as,
a ranking  of</p>
      <p>
        such that high quality items are above lesser-quality items. In practice, this
ranking is truncated for display or eficiency reasons. Given a ranking  and  , we can evaluate
performance using an information retrieval metric  , which models user browsing behavior [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>We are interested in how  is defined. One way to do this is to use an oracle to label the
quality of all items given a prefix. This oracle would have access to the entire inventory, the
internal state of the user, and suficient time to assemble the ideal reward. In practice, we rarely
have access to such an oracle, so we use logged interactions to infer . Let  be an observed
length-ℓ session, from which we can extract ℓ − 1 prefixes for evaluation. For a prefix
 =  (︀ 1,),
Since, in our logs, we only observe one selected item for a prefix, this underestimates the reward
for similar or substitutable items. To address this, we can introduce an assumption of ‘in-session
substitutability’ which considers all selected items in the session substitutes,


︀]
︀]
 = ⌋︀︀⌉︀⌉ 0
︀⌉︀)︀⌉ 1  =</p>
      <p>otherwise
 = ⌋︀︀⌉︀⌉ 0
︀⌉︀)︀⌉ 1  ∈  (︀ ,ℓ︀⌋</p>
      <p>otherwise
We can imagine more elaborate definitions inspired by reinforcement learning, but all are based
on immediate implicit user feedback.</p>
      <p>For the purpose of our discussion, we assume that we have access to the oracle  . This
allows us to compute, for each session, the set of optimal sequences. Such an oracle has access
to the full catalog of items and understands any potential interactions between items that may
afect their utility.</p>
      <p>In addition, we will consider the protocol for gathering session data described in Algorithm
1. This covers a large class of production recommendation system workflows. Radio-style
(1)
(2)
 ←GetFeedback(,  )
 ←GetRanking(,  ) ▷ get slate from logging policy based on the current prefix
▷ gather session  from a specific user</p>
      <p>▷ initialize session sequence
▷ observe item selected by user
▷ append selected item to prefix
▷ terminate sequence if the user abandons
▷ return the session sequence
recommendation is a special case where ⋃︀  ⋃︀ = 1 and the user has an additional SKIP reaction,
which does not get appended to  . Furthermore, we are agnostic about the logging policy and
include those that incorporate randomization.</p>
      <p>Algorithm 1 Session Data Collection</p>
      <p>1: function SessionCollect()</p>
    </sec>
    <sec id="sec-3">
      <title>3. Problems with Current SBRS Evaluation</title>
      <p>We now turn the potential issues with this method of collecting labeled data to evaluate SBRS.
Our claim is not that all of these concerns are present in all SBRSs, although we suspect that
many are.</p>
      <p>
        First, consider the impact of the user selecting items under incomplete information. Due to
system constraints, the user is only ever presented with a ranking of a subset of the catalog.
Moreover, because users scan rankings from top to bottom, with an increasing probability
of abandoning the scan (i.e. position bias), a choice will often be made amongst the
topranked items. As a result, sequential choices are made with severely limited options and
information. This is referred to as choice bracketing [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and we depict it in Figure 1a. There
are several implications of choice bracketing. First, limited options can result in selecting a
suboptimal item in the sequence, since a user may not see superior options. Moreover, these
unexamined, superior options disappear in future rankings when a recommendation system
removes previously-recommended items to reduce the perception of redundancy. Second, choice
bracketing potentially narrows and distorts a user’s decision context (i.e. inspected relevant and
non-relevant items) leading to priming and potentially inaccurate choices [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. A new system
that presents diferent rankings will bracket choices diferently and potentially result in diferent
optimal choices.
      </p>
      <p>Second, we turn to the problem of label sparsity in SBRS. For a given prefix,  will be
incomplete, especially if we only consider the next observed item   as relevant (Equation 1).
Even if adopt the ‘in-session substitutability’ assumption (Equation 2), a user will rarely exhaust
the set of relevant items in a session. We refer to this as the problem of incomplete substitutes
and depict it in Figure 1b. Here, because we are only selecting one item from the ranking at a
time, potential substitutes are unobserved and considered nonrelevant. Such sparsity issues have
resulted in distorted evaluation in both recommendation system [8] and information retrieval
contexts [9]. And, although this can be mitigated by algorithms based on equal exposure [10],

from an evaluation perspective, collecting a large set of equally efective trajectories is unlikely
for tail interests.</p>
      <p>Third, in many sequential recommendation tasks, there are sequential dependencies between
items. For example, users may not want to hear two songs from the same musician one after the
other, even though they find the two songs enjoyable otherwise. Unfortunately, the ‘in-session
substitutability’ assumption can overestimate the value of a substituted item from ′ &gt;  if, for
example, the item degrades in value when recommended immediately after  −1. Similarly,
an unselected item at time ′ &gt;  may be underestimated in value if it is substitutable with  
but was not selected at time ′ because of its recommendation immediately after  ′−1. The
contextual utility of items, especially for entertainment goods, could be due to satiation [11] or
other order efects [ 12]. We refer to this as the problem of inter-item dependency and depict it
in Figure 1c, where the oracle preferences at time  depend on the choice made at time  − 1.</p>
      <p>Fourth, many session-based recommendation systems include a default choice, usually the
topranked item. In streaming media platforms, this often means automatically playing the default
choice after some short period of time for the user to select an alternative. This functionality
results in reinforcing the default option (or, more generally, the system ranking), even if the
user compares it to alternatives [13]. We call this the problem of default preferences and depict
it in Figure 1d. As with choice bracketing and incomplete substitutes, this can result in missing
labels and, in cases where the default option is not relevant, the implicit feedback is incorrect.</p>
      <p>Fifth, consider user interfaces that include a thumbnail summary for each recommended item.
The visual attributes of this summary can vary across items and can cause users to inspect and
select items that are more visually salient [14, 15]. Salient items can disrupt inspection by rank
order, resulting in selection of inferior items. We refer to this as the problem of presentation
bias and depict it Figure 1e. Given the similarity to choice bracketing, it results in the same
problems.</p>
      <p>Finally, in some cases, oracle preferences may be inconsistent with the observed preferences
because preferences change at the moment of choice. Consider the example of choosing what
to eat. Prior work has demonstrated that people often choose healthy options if there is some
temporal delay between the choice of what to eat and when they actually eat their choice; those
preferences reverse in favor of the less healthy option if the choice is made immediately before
consumption [16]. Similarly, experiments have shown that people will select ‘highbrow’ movies
if asked days before watching the film; their preferences will shift to ‘lowbrow’ movies if asked
on the day of the watching the film [ 17]. In the context of SBRS, this means that observed
choices made instantaneously and sequentially may be reversals of ‘healthier’ preferences
expressed with foresight. We call this the problem of immediacy efects and depict it in Figure
1f. By construction, we defer to the oracle preferences when they disagree with immediate
preferences, which may be susceptible to impulsive behavior.</p>
      <p>In addition to these concerns with implicit labels from session data, how session data is
gathered and segmented into evaluation prefixes can distort performance. Simple issues like
over-representing sessions from active users are familiar to recommendation systems researchers.
Sessions can introduce additional issues. Prefixes are often selected to include all but the final
item in the session. This can hide under-performance at earlier points in the session, which is a
problem when user preferences or behavior changes over the length of the session. Evaluation
preparation that considers all prefixes in a session for evaluation can, for situations where
session lengths are not fixed, over-emphasize performance at the beginning of the session.
Separately, evaluation trajectories themselves are biased by the data-gathering policy and may
not be representative of the prefixes encountered when the evaluated SBRS is deployed [18].</p>
      <p>The magnitude of these problems depends on the domain. For example, in radio-style music
recommendation, users have an aversion to silence [19]; this may result in urgency and, as
a result, amplify position bias and narrow choice brackets. In some shopping settings, order
efects may be less pronounced. Text-only interfaces will be less susceptible to presentation
bias. Furthermore, these problems can interact and compound efects. For example, visually
salient items can increase impulsivity and potentially lead to immediacy efects [20].</p>
      <p>
        The implications of label unreliability depends largely on the domain. In the context of search,
cognitive biases in labels have been demonstrated to impact performance of learning-to-rank
systems [21]. In the context of traditional recommendation system evaluation, failing to consider
biased labeling can result in system under-performance in practice [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Addressing the Problems with Current SBRS Evaluation</title>
      <p>Since this is a position paper, we would like to conclude with possible next steps for the
community to consider, given the potential issues with current SBRS evaluation. Specifically,
a research program on SBRS preference elicitation could be built around the following two
themes: (i) recognizing that a domain is susceptible to problems in Section 3, and (ii) mitigation
strategies for those problems.</p>
      <p>Some of these problems can be recognized by looking at out-of-session information indicating
higher-level user satisfaction with the SBRS. One way to do this is to understand the relationship
between in-session behavior and longer-term user retention, surveys, and other tools [22, Part
4]. Alternatively, smaller scale, controlled, laboratory experiments and qualitative research can
also provide an indication of these problems and is especially efective when combined with
larger scale log data [23, 24].</p>
      <p>
        These problems, when detected, can be addressed in a variety of ways. In the context of search,
there is a small body of work focused on extracting relevance information in the presence of click
feedback under position bias [25]. Randomization and other of-policy evaluation techniques
can be used to address some, although not all, of the concerns in SBRS [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Explicit models of
‘unhealthy’ items can also be used to guide recommendations toward healthier options [26].
      </p>
      <p>A diferent way to approach these problems is to change interface elements to support
decision-making. For example, widgets like shortlists can help expand brackets [27]. In the
context of exploratory search, assistive tools like note-taking devices can also improve long-term
goals like task completion [28, 29].</p>
      <p>A final option, in domains where the space of information needs is small, we can consider
directly modeling the oracle. In the context of music recommendation, users often spend time
manually-curating playlists for future consumption in specific situations [ 30]. As such,
manuallycurated playlists provide a rich source of ‘gold standard’ data for SBRS [31, 32]. Therefore,
developing exploratory search and other tools to support curation, in music or other domains,
can provide one way to crowd-source oracle data [33]. Similar methods have been used in the
context of query autocompletion, which can be considered a character-level recommendation
task [34]. That said, there are some cases where asking a user to provide oracle decisions is
unsuccessful because people can fail to consider important contextual information necessary for
understanding the appropriate choice [35]. For example, one might select meals for a week but
not consider the time pressures or exhaustion that may make the efort to prepare the healthiest
meal not worth it. This tension between immediacy efects and inaccurate forecasting, then,
comes to the fore.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this position paper, we have argued that several of the current practices for gathering label
and reward data from implicit feedback is susceptible to error and may impact evaluation of
SBRSs. While several of these problems have been discussed in the context of algorithm design,
we believe that moving the investigation to our practice of evaluation will add nuance to our
understanding of how users interact with recommendation systems and, as a result, improve
the design of these systems.
information retrieval, in: Proceedings of the 2021 Conference on Human Information
Interaction and Retrieval, CHIIR ’21, Association for Computing Machinery, New York,
NY, USA, 2021, pp. 27–37. URL: https://doi.org/10.1145/3406522.3446023. doi:10.1145/
3406522.3446023.
[8] P. Kouki, I. Fountalis, N. Vasiloglou, X. Cui, E. Liberty, K. Al Jadda, From the lab to
production: A case study of session-based recommendations in the home-improvement
domain, in: Fourteenth ACM Conference on Recommender Systems, RecSys ’20,
Association for Computing Machinery, New York, NY, USA, 2020, pp. 140–149. URL:
https://doi.org/10.1145/3383313.3412235. doi:10.1145/3383313.3412235.
[9] N. Arabzadeh, A. Vtyurina, X. Yan, C. L. A. Clarke, Shallow pooling for sparse labels, 2021.
[10] F. Diaz, B. Mitra, M. D. Ekstrand, A. J. Biega, B. Carterette, Evaluating stochastic rankings
with expected exposure, arXiv e-prints (2020) arXiv:2004.13157.
[11] J. Galak, J. P. Redden, The properties and antecedents of hedonic decline, Annual Review
of Psychology 69 (2018) 1–25. URL: https://doi.org/10.1146/annurev-psych-122216-011542.
doi:10.1146/annurev-psych-122216-011542, pMID: 28854001.
[12] M. Eisenberg, C. Barry, Order efects: A study of the possible influence of presentation
order on user judgments of document relevance, Journal of the American Society for
Information Science 39 (1988) 293–300.
[13] W. Samuelson, R. Zeckhauser, Status quo bias in decision making, Journal of Risk
and Uncertainty 1 (1988) 7–59. URL: https://doi.org/10.1007/BF00055564. doi:10.1007/
BF00055564.
[14] Y. Yue, R. Patel, H. Roehrig, Beyond position bias: examining result attractiveness as a
source of presentation bias in clickthrough data, in: Proceedings of the 19th international
conference on World wide web, WWW ’10, ACM, New York, NY, USA, 2010, pp. 1011–1018.
URL: http://doi.acm.org/10.1145/1772690.1772793. doi:http://doi.acm.org/10.1145/
1772690.1772793.
[15] F. Diaz, R. W. White, G. Buscher, D. Liebling, Robust models of mouse movement on
dynamic web search results pages, in: Proceedings of the 22nd ACM conference on
Information and knowledge management (CIKM 2013), Association for Computing Machinery,
New York, NY, USA, 2013, pp. 1451–1460. URL: https://doi.org/10.1145/2505515.2505717.
[16] D. Read, B. van Leeuwen, Predicting hunger: The efects of appetite and delay on choice,
Organizational Behavior and Human Decision Processes 76 (1998) 189–205. URL: https:
//www.sciencedirect.com/science/article/pii/S0749597898928035. doi:https://doi.org/
10.1006/obhd.1998.2803.
[17] D. Read, G. Loewenstein, S. Kalyanaraman, Mixing virtue and vice: combining the
immediacy efect and the diversification heuristic, Journal of Behavioral Decision Making
12 (1999) 257–273.
[18] K.-W. Chang, A. Krishnamurthy, A. Agarwal, H. Daume, J. Langford, Learning to search
better than your teacher, in: Proceedings of The 32nd International Conference on Machine
Learning, 2015, pp. 2058–2066.
[19] A. J. Lonsdale, A. C. North, Why do we listen to music? a uses and gratifications
analysis, British Journal of Psychology 102 (2011) 108–134. URL: https://bpspsychub.
onlinelibrary.wiley.com/doi/abs/10.1348/000712610X506831. doi:https://doi.org/10.
1348/000712610X506831.
[20] B. V. den Bergh, S. Dewitte, L. Warlop, J. D. served as editor, B. S. served as associate editor
for this article., Bikinis instigate generalized impatience in intertemporal choice, Journal
of Consumer Research 35 (2008) 85–97. URL: http://www.jstor.org/stable/10.1086/525505.
[21] C. Eickhof, Cognitive biases in crowdsourcing, in: Proceedings of the Eleventh ACM
International Conference on Web Search and Data Mining, WSDM ’18, Association for
Computing Machinery, New York, NY, USA, 2018, pp. 162–170. URL: https://doi.org/10.
1145/3159652.3159654. doi:10.1145/3159652.3159654.
[22] P. Chandar, F. Diaz, B. St. Thomas, Beyond accuracy: Grounding evaluation metrics for
human-machine learning systems, in: Advances in Neural Information Processing Systems,
2020.
[23] P. Chandar, F. Diaz, C. Hosey, B. St. Thomas, Mixed method development of evaluation
metrics, in: KDD ’21: Proceedings of the 27th ACM SIGKDD International Conference on
Knowledge Discovery &amp; Data Mining, 2021. URL: https://kdd2021-mixedmethods.github.
io/.
[24] Q. Zhao, M. C. Willemsen, G. Adomavicius, F. M. Harper, J. A. Konstan, Interpreting
user inaction in recommender systems, in: Proceedings of the 12th ACM Conference on
Recommender Systems, RecSys ’18, Association for Computing Machinery, New York,
NY, USA, 2018, pp. 40–48. URL: https://doi.org/10.1145/3240323.3240366. doi:10.1145/
3240323.3240366.
[25] A. Chuklin, I. Markov, M. d. Rijke, Click models for web search, Synthesis Lectures on
Information Concepts, Retrieval, and Services 7 (2015) 1–115. URL: https://doi.org/10.2200/
S00654ED1V01Y201507ICR043. doi:10.2200/S00654ED1V01Y201507ICR043.
[26] A. Singh, Y. Halpern, N. Thain, K. Christakopoulou, E. H. Chi, J. Chen, A. Beutel, Building
healthy recommendation sequences for everyone: A safe reinforcement learning approach,
in: FAccTRec Workshop, 2020.
[27] T. Schnabel, P. N. Bennett, S. T. Dumais, T. Joachims, Using shortlists to support decision
making and improve recommender system performance, in: Proceedings of the 25th
International Conference on World Wide Web, WWW ’16, International World Wide Web
Conferences Steering Committee, Republic and Canton of Geneva, CHE, 2016, pp. 987–997.</p>
      <p>URL: https://doi.org/10.1145/2872427.2883012. doi:10.1145/2872427.2883012.
[28] D. Donato, F. Bonchi, T. Chi, Y. Maarek, Do you want to take notes? identifying research
missions in yahoo! search pad, in: Proceedings of the 19th International Conference on
World Wide Web, WWW ’10, Association for Computing Machinery, New York, NY, USA,
2010, pp. 321–330. URL: https://doi.org/10.1145/1772690.1772724. doi:10.1145/1772690.
1772724.
[29] A. Crescenzi, Y. Li, Y. Zhang, R. Capra, Towards better support for exploratory search
through an investigation of notes-to-self and notes-to-share, in: Proceedings of the 42nd
International ACM SIGIR Conference on Research and Development in Information Retrieval,
SIGIR’19, Association for Computing Machinery, New York, NY, USA, 2019, pp. 1093–1096.</p>
      <p>URL: https://doi.org/10.1145/3331184.3331309. doi:10.1145/3331184.3331309.
[30] A. N. Hagen, The playlist experience: Personal playlists in music streaming services,
Popular Music and Society 38 (2015) 625–645. URL: https://doi.org/10.1080/03007766.2015.
1021174. doi:10.1080/03007766.2015.1021174.
[31] I. Kamehkhosh, D. Jannach, User perception of next-track music recommendations, in:
Proceedings of the 25th Conference on User Modeling, Adaptation and Personalization,
UMAP ’17, Association for Computing Machinery, New York, NY, USA, 2017, pp. 113–121.</p>
      <p>URL: https://doi.org/10.1145/3079628.3079668. doi:10.1145/3079628.3079668.
[32] C.-W. Chen, P. Lamere, M. Schedl, H. Zamani, Recsys challenge 2018: Automatic music
playlist continuation, in: Proceedings of the 12th ACM Conference on Recommender
Systems, RecSys ’18, Association for Computing Machinery, New York, NY, USA, 2018, pp. 527–
528. URL: https://doi.org/10.1145/3240323.3240342. doi:10.1145/3240323.3240342.
[33] K. Lukof, U. Lyngs, H. Zade, J. V. Liao, J. Choi, K. Fan, S. A. Munson, A. Hiniker, How the
Design of YouTube Influences User Sense of Agency, Association for Computing Machinery,
New York, NY, USA, 2021. URL: https://doi.org/10.1145/3411764.3445467.
[34] M. Shokouhi, Learning to personalize query auto-completion, in: Proceedings of the 36th
International ACM SIGIR Conference on Research and Development in Information Retrieval,
SIGIR ’13, Association for Computing Machinery, New York, NY, USA, 2013, pp. 103–112.</p>
      <p>URL: https://doi.org/10.1145/2484028.2484076. doi:10.1145/2484028.2484076.
[35] T. D. Wilson, D. T. Gilbert, Afective forecasting, volume 35 of Advances in
Experimental Social Psychology, Academic Press, 2003, pp. 345–411. URL: https://
www.sciencedirect.com/science/article/pii/S0065260103010062. doi:https://doi.org/
10.1016/S0065-2601(03)01006-2.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. Z.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Orgun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lian</surname>
          </string-name>
          ,
          <article-title>A survey on session-based recommender systems</article-title>
          ,
          <source>ACM Comput. Surv</source>
          .
          <volume>54</volume>
          (
          <year>2021</year>
          ). URL: https://doi.org/10.1145/3465401. doi:
          <volume>10</volume>
          .1145/3465401.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Marlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Zemel</surname>
          </string-name>
          ,
          <article-title>Collaborative prediction and ranking with non-random missing data</article-title>
          ,
          <source>in: Proceedings of the Third ACM Conference on Recommender Systems</source>
          , RecSys '09,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2009</year>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>12</lpage>
          . URL: https://doi.org/10.1145/1639714.1639717. doi:
          <volume>10</volume>
          .1145/1639714.1639717.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Charlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McInerney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Blei</surname>
          </string-name>
          ,
          <article-title>Modeling user exposure in recommendation</article-title>
          ,
          <source>in: Proceedings of the 25th International Conference on World Wide Web, WWW '16, International World Wide Web Conferences Steering Committee, Republic and Canton of Geneva, CHE</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>951</fpage>
          -
          <lpage>961</lpage>
          . URL: https://doi.org/10.1145/2872427.2883090. doi:
          <volume>10</volume>
          . 1145/2872427.2883090.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Beutel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Covington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Belletti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          ,
          <article-title>Top-k of-policy correction for a reinforce recommender system</article-title>
          ,
          <source>in: Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining, WSDM '19</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2019</year>
          , pp.
          <fpage>456</fpage>
          -
          <lpage>464</lpage>
          . URL: http://doi.acm.
          <source>org/10</source>
          .1145/3289600.3290999. doi:
          <volume>10</volume>
          .1145/3289600. 3290999.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Carterette</surname>
          </string-name>
          ,
          <article-title>System efectiveness, user models, and user utility: a conceptual framework for investigation</article-title>
          ,
          <source>in: Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval, SIGIR '11</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2011</year>
          , pp.
          <fpage>903</fpage>
          -
          <lpage>912</lpage>
          . URL: http://doi.acm.
          <source>org/10</source>
          .1145/2009916.2010037. doi:
          <volume>10</volume>
          .1145/ 2009916.2010037.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Read</surname>
          </string-name>
          , G. Loewenstein,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rabin</surname>
          </string-name>
          ,
          <article-title>Choice bracketing</article-title>
          ,
          <source>Journal of Risk and Uncertainty</source>
          <volume>19</volume>
          (
          <year>1999</year>
          )
          <fpage>171</fpage>
          -
          <lpage>197</lpage>
          . URL: http://www.jstor.org/stable/41760959.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <article-title>Cognitive biases in search: A review and reflection of cognitive biases in</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>