<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Ranked Bandit Approach for Multi-stakeholder Recommender Systems∗</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>TAHEREH ARABGHALIZI</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>University of Pittsburgh</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>USA ALEXANDROS LABRINIDIS</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>University of Pittsburgh</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>Recommender systems traditionally find the most relevant products or services for users tailored to their needs or interests but they ignore the interests of the other sides of the market (aka stakeholders). In this paper, we propose to use a Ranked Bandit approach for an online multi-stakeholder recommender system that sequentially selects top  items according to the relevance and priority of all the involved stakeholders. We presented three diferent criteria to consider the priority of each stakeholder when evaluating our approach. Our extensive experimental results on a movie dataset showed that the contextual multi-armed bandits with a relevance function make a higher level of satisfaction for all involved stakeholders in the long term.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>1.1
1.2</p>
    </sec>
    <sec id="sec-2">
      <title>Background and Related Work</title>
      <p>
        A traditional user-centric recommender system recommends items based on the interests and
preferences of the user. However, the user is not the sole stakeholder in many real-world applications.
In those cases, the recommendations could benefit other individuals or organizations [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
Moreover, online recommendation tasks, which ingest data one observation at a time, are not well
served by traditional ofline recommendation techniques such as Collaborative Filtering [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] which
rely on historical data. It has recently come to the attention of online recommendation tasks to
study multi-armed bandits (MAB), a classic reinforcement learning problem. As a result of this
approach, a solution can be provided for the dilemma between exploration and exploitation that
maximizes the expected cumulative payof over the long term [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Depending on whether side information (aka context) is taken into account, bandit algorithms
fall into two categories: context-free and contextual. In context-free bandits, the observed reward
depends on the selected arm while it is both the selected arm and its context that determine the
observed reward in contextual bandits [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. While contextual bandits have been used extensively
for online user-centric recommender systems [
        <xref ref-type="bibr" rid="ref11 ref20">11, 20</xref>
        ] and much research has been conducted on
multi-objective multi-armed bandit algorithms [
        <xref ref-type="bibr" rid="ref21 ref23 ref7">7, 21, 23</xref>
        ], multi-stakeholder resommender systems
have received less attention. Most recently, Mehrotra et al. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] proposed a multiple-objective
contextual bandit approach to maximize the long-term payofs for diferent objectives such as
diversity for a multi-stakeholder recommender system on a music streaming platform. Although,
there are several existing works that consider multiple stakeholders in their recommendation
generation [
        <xref ref-type="bibr" rid="ref15 ref19">15, 19</xref>
        ], to the best of our knowledge, there is no work that considers the contextual
information of users, preferences and priority of all stakeholders in an online recommender system.
1.3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Contributions</title>
      <p>In this work, we aim to provide long-term, acceptable levels of satisfaction for all the stakeholders
involved in a recommender system, according to the given context (if available) and priorities. We
propose a bandit-based approach that sequentially selects top candidate items and then accepts the
best candidates that maximize the total payof, given relevance and priority of each stakeholder.
Our contributions are as follows:
• we propose to use a Ranked Bandit approach to recommend multiple items to a user and
provide acceptable levels of satisfaction for all stakeholders in a recommender system over
time.
• we introduce three diferent criteria including Deterministic, Probabilistic and Multi-sided
relevance Function to consider the priority of each stakeholder when evaluating our approach.
• we define and use diferent metrics to evaluate the satisfaction of stakeholders and compare
diferent benchmarks.
• we use real-world and synthetic datasets and multiple sensitivity analyses to experimentally
evaluate the performance of our approach.
2</p>
    </sec>
    <sec id="sec-4">
      <title>PROBLEM STATEMENT</title>
      <p>In this paper, we address the problem of recommending items to users in an online multi-stakeholder
platform, considering the contextual information of users and items (if available), relevance and priority
of involved stakeholder in order to provide them with an acceptable level of satisfaction in the long
term. In the application of coupon recommendation, the recommender system needs to recommend
top  coupons (from the finite set of available coupons at each bus stop) to a bus passenger whose
up-coming bus is going to be full. This system should take the preferences of all stakeholders
such as bus passengers, coupon suppliers and minority-owned businesses (i.e., another stakeholder
group that we want to prioritize) into account and ofer the relevant coupons based on the given
priority (weight) of stakeholders. This recommender system aims to maximize the total payofs
and makes a good balance between the satisfaction of all stakeholders over time.
3</p>
    </sec>
    <sec id="sec-5">
      <title>PROPOSED SOLUTION</title>
      <p>
        To address this problem, we propose to utilize a Ranked Bandit approach that sequentially selects
top  items and accepts the best candidates that maximize the payof, given relevance and priority
of each stakeholder.
In Ranked Bandit algorithm, which was first introduced by [
        <xref ref-type="bibr" rid="ref16 ref18">16, 18</xref>
        ], one multi-armed bandit
algorithm is instantiated for each slot of a ranked list with  slots, to learn the greedy-optimal
solution. If a user clicks on slot , that slot gains a reward of 1, and all other slots  where  &lt; 
receive a reward of 0. In a multi-stakeholder platform, however, the algorithm should consider the
relevance and priority (if applicable) of all involved stakeholders to obtain a reward of 1 for each
slot. To this end, we introduce three diferent criteria including deterministic, probabilistic, and
multi-sided relevance function (see Section 4.1), apply them after displaying the selected items to a
user and receive a reward accordingly.
      </p>
      <p>Our proposed approach is described in Algorithm 1. As you can see, at each round, the Ranked
Bandit algorithm plays a multi-armed bandit instance  for each ranking slot . If the arm 
was already selected at a higher ranking slot, an arbitrary unselected arm is chosen instead. This
process is being sequentially repeated for top  ranking slot (Lines 3-9). Then the algorithm receive
a feedback for each multi-armed bandit instance  through running a criterion. If the arm in
ranking slot  is relevant to the involved stakeholders based on the used criterion and the given
priority (weight) of stakeholders, then bandit  receives a reward of 1 and all higher bandits 
( &lt; ) receive a reward of 0 (Lines 12-17). Each bandit instance updates its reward afterwards.
Algorithm 1 Ranked Bandit for Multi-stakeholder Recommender Systems</p>
    </sec>
    <sec id="sec-6">
      <title>4 EVALUATION</title>
      <p>
        The training process in online algorithms, such as bandits, occurs incrementally over time.
However, ofline evaluation methods can be used to evaluate the performance of such algorithms using
historical datasets. One of the most popular ofline evaluation methods is called Replay. In this
methodology, the bandit algorithm selects an arm for each record in the historical data. If this
arm has been seen by the user before, the replay methodology accepts the arm, otherwise the
arm is denied by the replay method [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This method works well in ofline evaluation of bandits
employed in the user-centric recommender systems which focus on user satisfaction obtained
based on user clicks. In a multi-stakeholder recommender system such as coupon recommendation,
however, the replay evaluation should be able to consider the relevance of a selected arm to multiple
stakeholders given the priority of those stakeholders. To this end, we propose three criteria called
Deterministic, Probabilistic, and Multi-sided Relevance Function to address the relevance of a selected
arm to multiple weighted stakeholders in the replay evaluation.
4.1
      </p>
    </sec>
    <sec id="sec-7">
      <title>Selection/Evaluation Criteria</title>
      <p>Deterministic: based on this criterion, the replay method accepts an arm if it is relevant to the
most prioritized stakeholder which is the one with the highest weight.</p>
      <p>
        Probabilistic: this criterion determines the winner arm by first partitioning the [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] interval to
diferent ranges according to the weights of the diferent stakeholders and then selecting a random
number ∈ [
        <xref ref-type="bibr" rid="ref1">0,1</xref>
        ] which would map to one of the ranges and therefore one of diferent stakeholders.
The replay method would then accept the arm which is relevant to that stakeholder.
Multi-sided Relevance Function: the multi-sided relevance function can be defined as the
weighted sum of the relevance score of each stakeholder to each selected arm a. The relevance
scores can be calculated based on the feedback of users and other stakeholders about the quality of
the recommended items. The relevance function is defined as below:
  (,  ) =  11 + ... + 
(1)
where   ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ] is the estimated multi-sided relevance score, n is the number of stakeholders,
1, 2, ...,  are the relevance scores of arm a for each stakeholder where each score is either
1 (relevant) or 0 (non-relevant),  is the given priority or weight of each stakeholder where
 ∈ [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ], Í=1  = 1. We also define a relevance threshold, , to let the replay evaluation accepts
an arm if the multi-sided relevance score is greater than this threshold [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
4.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>Evaluation Metrics</title>
      <p>
        We use three diferent types of metrics to evaluate the performance of our proposed approach:
• Average Rewards: we accumulate the reward values of the accepted arms in each round and
return the average of the rewards for total rounds. The ultimate goal of a bandit algorithm is
to maximize the rewards in the long term.
• Average NDCG : Normalized Discounted Cumulative Gain (NDCG) measures the usefulness
of an item based on its position in the recommendation list [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. We compute the Average
NDCG at  for each stakeholder over time as the representation of their satisfaction level.
• Balance-Relevance Rate: this metric computes the product of the balance and relevance
rates which are defined as follows:
– Balance Rate: to measure how similar the weights of the stakeholders and their satisfaction
levels (i.e. Average NDCG) are, we compute the cosine similarity between the vector of
weights and their corresponding Average NDCG.
– Relevance Rate: this is a measure that calculates the number of times (out of total rounds)
that the selected arms are relevant to stakeholders (using one of the evaluation criteria)
and are accepted by the replay evaluation.
5
      </p>
    </sec>
    <sec id="sec-9">
      <title>EXPERIMENTAL EVALUATION</title>
      <p>In this section, we conduct several experiments on a movie recommendation data to evaluate the
performance of our proposed approach.
5.1</p>
    </sec>
    <sec id="sec-10">
      <title>Dataset</title>
      <p>
        Since there are currently no public datasets available for a multi-stakeholder platform in the coupon
recommendation domain, we merged two real-world datasets including MovieLens (1m) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] with
IMDB (81k+) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] and also generated a synthetic dataset. In our experiments, we use a scenario
where there are three stakeholders involved in the coupon recommendation system. According to
the data, movies are used as coupons, users as bus passengers (first stakeholder), movie production
companies as local businesses/coupon suppliers (second stakeholder) and movies with a specific
genre as a minority-owned business (third stakeholder). We created the following datasets for
training and ofline evaluation of the bandit algorithms:
Movies: we use top m movies with the highest number of ratings as arms.
      </p>
      <p>Movies’ Features: the genres of movies are used as the context information of the arms.
Users’ Features: users’ demographic features such as age range and gender along with the the
normalized average of the genres of the movies that each user rated, are used as the context
information of the users.</p>
      <p>Users Data: the ratings of the users to the movies (integers between 1 and 5) are used as the
feedback of users in bandit evaluation. If the rating of a selected movie is greater than 3.0, the movie
is likely relevant to the preferences of that user and they will click on it. It should be noted that we
used Singular Value Decomposition (SVD) technique to fill in the missing data in the user ratings.
Suppliers Data: as there aren’t any ratings data for production companies towards the users in the
MovieLens or IMDB datasets, we generated a synthetic dataset. First we clustered all users using
KMeans algorithm and then for each production company, we used a Truncated Normal Distribution
with mean = 3 and standard deviation = 1, to generate random integer numbers between 1 and 5
as the ratings of that production company towards the users in the same cluster. If a production
company has given a rating greater than 3.0 to the current user, that user is likely relevant to that
company’s preferences.</p>
      <p>Minority Data: we consider ‘Sci-Fi’ movies as minority-owned businesses (since only about 7%
of all movies in the MovieLens data are ‘Sci-Fi’). If a selected movie is Sci-Fi, it is assumed to be
relevant to minority-owned businesses.</p>
      <p>Observation Data: we use a stream of users who rated the top movies (arms) in the past to evaluate
the bandits.
5.2</p>
    </sec>
    <sec id="sec-11">
      <title>Base Bandit Algorithms</title>
      <p>
        In this section, we describe the base multi-armed bandit algorithms that are used in our approach:
• Hybrid LinUCB: as mentioned earlier, contextual bandits is a variant of the multi-armed
bandit problems that utilizes contextual information. In many applications, the arms are not
distinct from each other and have shared features. For example, in a coupon recommendation
system where the coupons are arms that are recommended to users, there may be some
similarities between coupons, and thus they share similar features. Li et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] introduced
LinUCB with Hybrid Linear Model where the payof of each arm is a linear function of shared
and non-shared components. We use the contextual information of users along with the
arms features for this algorithm. LinUCB has a hyper-parameter, , which should be tuned
in advance.
• Thompson Sampling: it is a Bayesian multi-armed bandit algorithm where a posterior
distribution is used to summarize reward values inferred from past data. Under this distribution,
the algorithm selects the arm proportionately to its likelihood of being optimal [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
• Epsilon-greedy: in each round, this algorithm selects a random arm with probability ,
and chooses the arm with the highest empirical mean with probability 1 −  [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].  is a
hyper-parameter and needs to be tuned properly.
• UCB1: an upper confidence bound has to be calculated for each arm for the algorithm to be
able to choose an arm in each round [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>• Random: it randomly selects an arm to pull at each round.
5.3</p>
    </sec>
    <sec id="sec-12">
      <title>Experimental Results</title>
      <p>In this section we present and discuss the results of our experiments.</p>
      <p>Results with Diferent Criteria : we conducted multiple experiments for all base bandit algorithm
with default settings (see Table 1) to compare the impact of our proposed evaluation criteria in
terms of Balance-Relevance Rate. As you can see in Table 2, where we show the results for the
Hybrid LinUCB algorithm, the multi-sided relevance function outperforms the other two criteria
(i.e. Deterministic and Probabilistic) in almost all experiments with diferent weight combinations.
This shows that given the priority of stakeholders, the multi-sided relevance function can provide
a better trade-of between the satisfaction of all stakeholders in the long term. For this reason, we
carried out the next experiments using the the multi-sided relevance function as the evaluation
criterion.
Results with Multi-sided Relevance Function: in this section we describe and illustrate the
resultsof several experiments with the default settings (see Table 1) and compare the base bandit
algorithms in terms of average rewards and average NDCG while prioritizing diferent stakeholders.
As it is shown in Figure 1, Hybrid LinUCB has the highest average rewards compared to the other
context-free bandit algorithms when we prioritize users (Figure 1a) and suppliers (Figure 1b). This
indicates that utilizing the contextual features of users and items can make higher average rewards
for all stakeholders when applying the Multi-sided Relevance Function as the evaluation criterion.
It should be mentioned that our results show that the Hybrid LinUCB algorithm outperforms other
base bandits when we also prioritize minority (as the third stakeholder).</p>
      <p>Furthermore, Figure 2 illustrates the Average NDCG for diferent stakeholders when running
diferent bandit algorithms with multi-sided relevance function. As one can see, Hybrid LinUCB
provides higher Average NDCG for the prioritized stakeholder and makes a good balance rate
(a) weights: (user: 0.6, supplier: 0.3, minority: 0.1)</p>
      <p>(b) weights: (user: 0.3, supplier: 0.6, minority: 0.1)
between the Average NDCG and the weights of stakeholders, in all scenarios when we prioritize
users (a), suppliers (b) or minority (c).</p>
      <p>Sensitivity Analysis: we performed multiple sensitivity analyses to see how Balance-Relevance
Rate changes with diferent set of stakeholders’ weights, diferent number of arms and diferent
number of ranking slots:
• Stakeholders’ Weights: we set the weights of stakeholders to diferent values (e.g., 0.5, 0.25,
0.25), where sum is equal to 1. The results of these experiments showed a similar pattern to
the the results of experiments with the default weights.
• Number of Arms: we set the number of arms to 100 (default), 200 and 300 for diferent bandits
while prioritizing diferent stakeholders. As the figures in Table 3 show, Hybrid LinUCB with
with multi-sided relevance function outperforms the other bandit algorithms in terms of
balance-relevance rate even when we change the number of arms.
• Number of Slots: we change the number of ranking slots (ofered items) to 1, 3 (default) and 5
while prioritizing one stakeholder at a time for 100 arms. As the results of these experiments
in Table 4 show, Hybrid LinUCB with multi-sided relevance function works better than other
algorithms in terms of balance-relevance rate for various number of slots. There are only
Stakeholders’ Prioritized Number of Hybrid</p>
      <p>Weights Stakeholder Arms LinUCB</p>
      <p>100 0.851
(0.6, 0.3, 0.1) User 200 0.835
300 0.834
100 0.967
(0.3, 0.6, 0.1) Supplier 200 0.930
300 0.934
100 0.948
(0.1, 0.3, 0.6) Minority 200 0.966
300 0.953</p>
      <p>TS</p>
      <p>EG</p>
      <p>UCB1 RAND
two scenario where Thompson Sampling works a little better than Hybrid LinUCB but the
diference between their balance-relevance rates is negligible (&lt;0.006). It is worth noting that
balance-relevance rates grow when the number of ofered items (slots) increases which is
reasonable because when the recommender system ofers higher number of items to users,
there is a higher probability that the ofered items are relevant to stakeholders.</p>
    </sec>
    <sec id="sec-13">
      <title>6 CONCLUSIONS</title>
      <p>In this paper, we addressed the problem of recommending coupons in a multi-stakeholder platform
where stakeholders can be prioritized. We proposed to use a Ranked Bandit approach that
sequentially selects top  items and accepts the best candidates that maximize the payof, given relevance
and priority of each stakeholder. We introduced three diferent criteria including Deterministic,
Probabilistic and Multi-sided Relevance Function to consider the priority of each stakeholder when
evaluating our approach. We also defined diferent metrics to evaluate the satisfaction of
stakeholders and compare diferent benchmarks. Our experimental results on real-world and synthetic
datasets showed that contextual multi-armed bandits (i.e. Hybrid LinUCB) with a multi-sided
relevance function outperforms context-free bandits with any evaluation criterion.</p>
    </sec>
    <sec id="sec-14">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work is part of the PittSmartLiving project which is supported by NSF award CNS-1739413.
0.429
0.348
0.293
0.424
0.444
0.434
0.262
0.250
0.201</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Himan</given-names>
            <surname>Abdollahpouri</surname>
          </string-name>
          , Gediminas Adomavicius, Robin Burke, Ido Guy, Dietmar Jannach, Toshihiro Kamishima, Jan Krasnodebski, and
          <string-name>
            <given-names>Luiz</given-names>
            <surname>Pizzato</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Multistakeholder recommendation: Survey and research directions</article-title>
          .
          <source>User Modeling and User-Adapted Interaction 30</source>
          ,
          <issue>1</issue>
          (
          <year>2020</year>
          ),
          <fpage>127</fpage>
          -
          <lpage>158</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Himan</given-names>
            <surname>Abdollahpouri</surname>
          </string-name>
          and
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          .
          <year>2022</year>
          .
          <article-title>Multistakeholder recommender systems</article-title>
          .
          <source>In Recommender systems handbook</source>
          . Springer,
          <fpage>647</fpage>
          -
          <lpage>677</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Tahereh</given-names>
            <surname>Arabghalizi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexandros</given-names>
            <surname>Labrinidis</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Data-driven bus crowding prediction models using contextspecific features</article-title>
          .
          <source>ACM Transactions on Data Science 1</source>
          ,
          <issue>3</issue>
          (
          <year>2020</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Tahereh</given-names>
            <surname>Arabghalizi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alexandros</given-names>
            <surname>Labrinidis</surname>
          </string-name>
          .
          <year>2022</year>
          .
          <article-title>Context-aware Multi-stakeholder Recommender Systems</article-title>
          .
          <source>In The International FLAIRS Conference Proceedings</source>
          , Vol.
          <volume>35</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Auer</surname>
          </string-name>
          , Nicolo Cesa-Bianchi, and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Fischer</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Finite-time analysis of the multiarmed bandit problem</article-title>
          .
          <source>Machine learning 47, 2</source>
          (
          <year>2002</year>
          ),
          <fpage>235</fpage>
          -
          <lpage>256</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Nicolo</given-names>
            <surname>Cesa-Bianchi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Fischer</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Finite-Time Regret Bounds for the Multiarmed Bandit Problem.</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>ICML</given-names>
          </string-name>
          , Vol.
          <volume>98</volume>
          .
          <fpage>100</fpage>
          -
          <lpage>108</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Madalina</surname>
            <given-names>M Drugan</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Ann</given-names>
            <surname>Nowe</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Designing multi-objective multi-armed bandits algorithms: A study</article-title>
          .
          <source>In The 2013 International Joint Conference on Neural Networks (IJCNN)</source>
          .
          <source>IEEE</source>
          , 1-
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>F Maxwell</given-names>
            <surname>Harper and Joseph A Konstan</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The movielens datasets: History and context</article-title>
          .
          <source>Acm transactions on interactive intelligent systems (tiis) 5</source>
          ,
          <issue>4</issue>
          (
          <year>2015</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Kalervo</given-names>
            <surname>Järvelin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jaana</given-names>
            <surname>Kekäläinen</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Cumulated gain-based evaluation of IR techniques</article-title>
          .
          <source>ACM Transactions on Information Systems (TOIS) 20</source>
          ,
          <issue>4</issue>
          (
          <year>2002</year>
          ),
          <fpage>422</fpage>
          -
          <lpage>446</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Stefano</given-names>
            <surname>Leone</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>IMDb movies extensive dataset</article-title>
          . https://www.kaggle.com/stefanoleone992/imdb-extensive-dataset
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Lihong</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wei Chu</surname>
          </string-name>
          , John Langford, and
          <string-name>
            <given-names>Robert E</given-names>
            <surname>Schapire</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>A contextual-bandit approach to personalized news article recommendation</article-title>
          .
          <source>In Proceedings of the 19th international conference on World wide web. 661-670.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Lihong</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wei Chu</surname>
          </string-name>
          , John Langford, and
          <string-name>
            <given-names>Xuanhui</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Unbiased ofline evaluation of contextual-bandit-based news article recommendation algorithms</article-title>
          .
          <source>In Proc of the 4th ACM intl conference on Web search and data mining</source>
          .
          <volume>297</volume>
          -
          <fpage>306</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Tyler</surname>
            <given-names>Lu</given-names>
          </string-name>
          , Dávid Pál, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Pál</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Contextual multi-armed bandits</article-title>
          .
          <source>In Proceedings of the Thirteenth international conference on Artificial Intelligence and Statistics . JMLR Workshop and Conference Proceedings</source>
          ,
          <fpage>485</fpage>
          -
          <lpage>492</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Rishabh</surname>
            <given-names>Mehrotra</given-names>
          </string-name>
          , Niannan Xue, and
          <string-name>
            <given-names>Mounia</given-names>
            <surname>Lalmas</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Bandit based Optimization of Multiple Objectives on a Music Streaming Platform</article-title>
          .
          <source>In Proc of the 26th ACM SIGKDD Intl Conference on Knowledge Discovery &amp; Data Mining</source>
          .
          <fpage>3224</fpage>
          -
          <lpage>3233</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Phong</surname>
            <given-names>Nguyen</given-names>
          </string-name>
          , John Dines, and
          <string-name>
            <given-names>Jan</given-names>
            <surname>Krasnodebski</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A multi-objective learning to re-rank approach to optimize online marketplaces for multiple stakeholders</article-title>
          .
          <source>arXiv preprint arXiv:1708.00651</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Filip</surname>
            <given-names>Radlinski</given-names>
          </string-name>
          , Robert Kleinberg, and
          <string-name>
            <given-names>Thorsten</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Learning diverse rankings with multi-armed bandits</article-title>
          .
          <source>In Proceedings of the 25th international conference on Machine learning</source>
          .
          <fpage>784</fpage>
          -
          <lpage>791</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Francesco</surname>
            <given-names>Ricci</given-names>
          </string-name>
          , Lior Rokach, and
          <string-name>
            <given-names>Bracha</given-names>
            <surname>Shapira</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Introduction to recommender systems handbook</article-title>
          .
          <source>In Recommender systems handbook. Springer</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Matthew</surname>
            <given-names>Streeter</given-names>
          </string-name>
          , Daniel Golovin, and
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Krause</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Online learning of assignments</article-title>
          .
          <source>Advances in neural information processing systems</source>
          <volume>22</volume>
          (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Özge</surname>
            <given-names>Sürer</given-names>
          </string-name>
          , Robin Burke, and Edward C Malthouse.
          <year>2018</year>
          .
          <article-title>Multistakeholder recommendation with provider constraints</article-title>
          .
          <source>In Proceedings of the 12th ACM Conference on Recommender Systems</source>
          .
          <volume>54</volume>
          -
          <fpage>62</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Liang</surname>
            <given-names>Tang</given-names>
          </string-name>
          , Yexi Jiang,
          <string-name>
            <given-names>Lei</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Tao</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Ensemble contextual bandits for personalized recommendation</article-title>
          .
          <source>In Proc. of the 8th ACM Conf. on Recommender Systems</source>
          .
          <volume>73</volume>
          -
          <fpage>80</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Cem</given-names>
            <surname>Tekin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Eralp</given-names>
            <surname>Turğay</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Multi-objective contextual multi-armed bandit with a dominant objective</article-title>
          .
          <source>IEEE Transactions on Signal Processing</source>
          <volume>66</volume>
          ,
          <issue>14</issue>
          (
          <year>2018</year>
          ),
          <fpage>3799</fpage>
          -
          <lpage>3813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>William</surname>
            <given-names>R</given-names>
          </string-name>
          <string-name>
            <surname>Thompson</surname>
          </string-name>
          .
          <year>1933</year>
          .
          <article-title>On the likelihood that one unknown probability exceeds another in view of the evidence of two samples</article-title>
          .
          <source>Biometrika</source>
          <volume>25</volume>
          ,
          <fpage>3</fpage>
          -
          <lpage>4</lpage>
          (
          <year>1933</year>
          ),
          <fpage>285</fpage>
          -
          <lpage>294</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Saba</surname>
            <given-names>Q Yahyaa</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madalina M Drugan</surname>
            ,
            <given-names>and Bernard</given-names>
          </string-name>
          <string-name>
            <surname>Manderick</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The scalarized multi-objective multi-armed bandit problem: An empirical study of its exploration vs. exploitation tradeof</article-title>
          .
          <source>In 2014 International Joint Conference on Neural Networks (IJCNN)</source>
          . IEEE,
          <fpage>2290</fpage>
          -
          <lpage>2297</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>