<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Application of Causal Bandit to Content Optimization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sameer Kanase</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yan Zhao</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shenghe Xu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mitchell Goodman</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manohar Mandalapu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benjamyn Ward</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chan Jeon</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shreya Kamath</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ben Cohen</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yujia Liu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hengjia Zhang</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yannick Kimmel</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saad Khan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brent Payne</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patricia Grao</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amazon</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Amazon encompasses a large number of discrete businesses such as Retail, Advertising, Fresh, Business (B2B e-commerce), and Prime Video, most of which maintain a presence across its e-commerce website. They produce content for our customers that belong to diverse content types such as merchandising (e.g. product recommendations), product advertisements (e.g. sponsored products and display ads), program adoption banners (e.g. Amazon Fresh), and consumption (e.g. Prime Video). When customers visit a web page on the website, it triggers a content allocation process where we determine the specific content to show in regions of customer shopping experience on that web page. Content produced by the aforementioned businesses then needs to be arbitrated during this process. We present a causal bandit based framework to address the problem of content optimization in this context. The framework is responsible for fairly balancing the difering objectives and methods of these businesses, and selecting the right content to display to the customers at the right time. It does so with the goal of improving the overall site-wide customer shopping experience. In this paper, we present our content optimization framework, describe its components, demonstrate the framework's efectiveness through online randomized experiments, and share learnings from deploying and testing the framework in production.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Personalization</kwd>
        <kwd>Recommender system</kwd>
        <kwd>Content optimization</kwd>
        <kwd>Content ranking</kwd>
        <kwd>Content diversity</kwd>
        <kwd>Causal bandit</kwd>
        <kwd>Contextual bandit</kwd>
        <kwd>View-through attribution</kwd>
        <kwd>Holistic optimization</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>When customers visit a web page on the website, it</title>
        <p>triggers a content allocation process where we
deterAmazon encompasses a large number of discrete busi- mine the specific content to show in the widget groups
nesses such as Retail, Advertising, Fresh, Business (B2B on that web page. Content produced by the
aforemene-commerce), and Prime Video, most of which maintain a tioned discrete businesses then needs to be arbitrated
presence across its e-commerce (or retail) website. These during this process. As the common integration point,
discrete businesses produce content for our customers Amazon’s content optimization framework is
responsithat belong to diverse content types such as merchandis- ble for this content arbitration. It accomplishes this by
ing (e.g. product recommendations), product advertise- fairly balancing the difering objectives and methods of
ment (e.g. sponsored products and display ads), program these businesses through optimization capabilities, and
adoption banner (e.g. Amazon Fresh), and consumption by taking into account customer, content, and shopping
(e.g. Prime Video). Each such content is rendered in the context. This results in the right content being shown
form of a widget within independent ‘regions of customer to the customers at the right time thereby providing a
shopping experience’ on the website, also known as wid- consistent and personalized shopping experience. The
get groups. For instance, widgets such as ‘customers who content optimization framework is an ecosystem which
viewed this also viewed’ and ‘customers who bought this enables businesses to interoperate independently by
enalso bought’ are displayed on product detail pages of the abling content creators, customer shopping experience
website alongside other organic and advertising content. providers, and web page owners to eficiently construct
The region of customer shopping experience on the web- and serve content for the retail website.
site where the collection of such widgets are displayed is In this paper, we present a causal bandit based
framean example of a widget group. We illustrate the concept work to address the problem of content optimization
of a product (or an item), widget, and widget group in with the objective of improving the overall customer
(figure 1). shopping experience on Amazon’s retail website. Our
contributions include:
• application of a contextual bandit framework to
enable introduction of new content through
online randomized experiments (or A/B tests) and
to learn the value (or benefit) of new content
through exploration,
• an approach to measure the reward for actions</p>
        <p>taken by the contextual bandit framework,
• application of view-through attribution (VTA) to
attribute reward in the context of content ranking
which only requires that content be impressed,
• utilization of an uplift modeling framework to
augment VTA and to optimize for incremental
benefit,
• a methodology to incorporate diversity in content
ranking by using cross-content interactions, and
• learnings from the deployment of a low-latency
learning framework in production that reduces
the delay in feedback and increases the velocity
of our learning loop.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Finally, we also demonstrate the efectiveness of our framework through online A/B tests, and share results and insights gathered through the same.</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>formally define the problem we address as determining
ranked set  from  given contexts  and  so as to
Application of exploration strategies in the context of rec- maximize the expected reward ℛ. Here, reward ℛ is a
ommender systems is an active area of research. In recent measure of improved customer shopping experience on
years, multiple exploration strategies have emerged and the retail website. We denote the metric for measuring
shown promising results [1, 2]. They include epsilon- reward ℛ by  , short for ‘metric of interest’. In our
greedy [3, 4], upper confidence bound (UCB) [ 5, 6], setting,   takes into account the short-term as well
adding random noise to parameters [7, 8, 9], and boot- as long-term impact to the customer’s shopping
experistrap sampling [10, 11]. We adopt the Thompson sam- ence, and helps us to fairly balance multiple and difering
pling algorithm [12] to balance exploration with exploita- objectives of various stakeholders. It is computed
ustion under the contextual bandit setting. Originally in- ing actions taken by the customer after interacting with
troduced in 1933 [13], Thompson sampling has been content such as impressions, clicks, purchases and other
widely adopted in the context of bandit problems recently high-value actions. Note that our problem is diferent
[14, 15, 16]. It has been shown to achieve state-of-the-art from that of ranking products (or items) within a single
results on some real-world use cases and be robust to widget for a particular recommender system.
delay [17, 18]. A key challenge we face in predicting   using</p>
      <p>Uplift modeling is a widely used approach to mea- and  is that of the estimate being biased due to the
coldsure incremental efect [ 19, 20, 21, 22]. Our approach to start problem. New content gets continually introduced
estimate incremental efect or benefit is similar to the to be shown on Amazon’s retail website while existing
meta-learning approach presented in [23, 24]. In [25], the content can be sunsetted at any point of time. Empirically,
authors presented an application of a causal bandit in tar- we observe a propensity in customers to interact more
geting campaigns. They estimated incremental efect to with content displayed higher up in the widget group and
optimize for clicks in email marketing campaigns and ad- on the web page. Furthermore, we only observe reward
vertisement campaigns on Amazon’s mobile homepage. for content that was shown to customers before, but we
In this work, we explore an application of causal bandit only show content to customers for which we predict
for content optimization by estimating and optimizing there will be suficient reward. Consequently, content
for heterogeneous treatment efect [26]. with few or no prior observations is unlikely to be ranked
higher or chosen to be shown to the customers even if
3. Problem Description it could generate a high-reward in the counterfactual
event where a customer were to interact with it. Here,
we could use aggregate-level features to partially address
the cold-start problem but cannot fully solve it. Moreover,
we observe that customer preferences and their
interactions with content change over time. To address these
challenges, we use a contextual bandit based framework
to create a learning loop for new content that has never
been shown before to the customers and to dynamically
adapt to changing customer preferences.</p>
      <p>Let  be the set of all web pages and  be the set of
all widget groups on Amazon’s retail website. Here, a
widget group refers to real estate or region of customer
shopping experience on the website which can be
populated with content  in the form of a widget. Content
 can belong to diverse types of content such as
product advertisements (e.g. sponsored products and display
ads), merchandising (e.g. product recommendations),
program adoption banners (e.g. Amazon Fresh), and
consumption (e.g. Prime Video). Each widget group  ∈ 
in turn can render (or display) a set of ranked content
 = { |  ∈    ∈ {1, . . . , }} where  is
the rank of content rendered in widget group ,  is
the total number of content that can be rendered in ,
and  is the set of all possible candidate content that is
eligible to be rendered in . Here, the cardinality of set
 &gt;&gt; . Eligibility for rendering content  in widget
group  is typically determined by business rules and
content creators.</p>
      <p>When a customer visits web page  ∈  on Amazon’s
retail website, a request is generated with customer and
shopping context  to optimize and display content for
widget group  on page . Context for candidate content
 ∈  can be constructed and is denoted by . We now</p>
    </sec>
    <sec id="sec-3">
      <title>4. Methodology</title>
      <sec id="sec-3-1">
        <title>In this section, we present a causal bandit based framework to address the problem of content optimization.</title>
        <sec id="sec-3-1-1">
          <title>4.1. Features</title>
          <p>When a customer visits web page  ∈  on Amazon’s
retail website, we receive customer and shopping
context . Context  corresponding to each candidate
content can also be generated separately. We then
combine contexts  and  non-linearly to form a single
d-dimensional vector  ∈ R. We also include
secondand third-order interaction terms between the
explanatory variables observed in the context. For reference, we
include a few examples of context below: we observe that results from such an approach can be
mixed in that it may improve customer shopping
expe• Shopping context: region, web page type, widget rience on some web pages of the website but not all of
group id, page item, metadata of page item, and them.</p>
          <p>
            search query To address these challenges, we propose optimizing
• Customer context: recent interaction events, cus- directly for overall down-session value generated after
tomer signed-in status, and prime membership customer has interacted with content. In this approach,
status once content has been ranked and rendered, we record
• Content context: widget id, widget meta informa- customer’s interaction events with it such as impressions,
tion, and content attributes clicks, purchases and other high-value actions.
Thereafter, we measure the aggregate value generated from
4.2. Ranking Model these events over a subsequent time horizon to compute
our metric of interest  . The measured value is
atWe formulate the problem of content optimization as that tributed to content as reward, if it meets a predefined
of learning to rank the set of eligible content . Our aim criteria. Content ranking models then learn to predict for
is to determine the rank of each eligible candidate content this down-session value of showing content to customers
 ∈  and return the  −  ranked content  so given a context, and make ranking decisions based on
as to render them in widget group . To do so, we need the predicted value. This approach enables us to measure
a utility function using which we can evaluate eligible and attribute site-wide impact across all devices, apps,
content and rank them. We propose using reward ℛ to widget groups, and web pages from the moment a
cusbe generated over a subsequent time horizon in the event tomer has interacted with content. We call this approach
content were to be shown to a customer,  ∈ {0, 1}, to define reward and rank diverse type of content using
as our utility function. We model it using a generalized aggregate down-session value as holistic optimization.
linear model,
(|, ) = (⊤ )
(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )
4.4. Attribution
where, g is the link function. Since reward ℛ takes The predefined criteria used to attribute aggregate
downcontinuous values in our problem setting, we choose session value as reward also defines the form of
attrian identity link function. We use the set of past obser- bution such as click-through attribution (CTA) or
viewvations  made up of triplets of context, action and through attribution (VTA). The distinction between these
reward {( ,  ,  ),  = 1, . . . ,  − 1} to train the two forms of attributions is the customer interaction
ranking model and estimate regression parameters using event that triggers the measurement of reward. In CTA,
a Bayesian framework. reward is measured after a click event with content
occurs while in VTA reward is measured after a view event
with content occurs. Note that both VTA and CTA are
4.3. Reward a form of equal credit attribution model. Likewise, the
A fundamental challenge in our problem setting is that time horizon over which the reward is measured is called
of defining and measuring reward so as to evaluate di- an attribution window. The window is triggered after a
verse types of content together on an equal footing [27]. customer interaction event with content occurs. We
deWhen content optimization systems seek to maximize termine attribution windows by performing exploratory
the attributed value (or reward) to individual content, data analysis of the length of customer shopping sessions
we observe that it leads to development and launch of and use multiple windows in practice to cater to varied
bespoke recommender systems that optimize for individ- use cases. We illustrate the concept of VTA and CTA
ual objectives and cater to page specific use cases. For with an attribution window using the example in (figure
instance, recommender systems displayed on diferent 2).
web pages can optimize for increasing customer inter- A key drawback of CTA is that it cannot attribute
reactions with themselves through view, clicks, purchases ward to content that cannot be clicked or where clicking
and other high-value actions without being complemen- on content does not necessarily indicate a positive
custary (or incremental) to the customer’s current shopping tomer shopping experience. We observe that CTA also
intent. This often results in a poor customer shopping leads ranking models to favor content that has a high
experience which in turn leads to a negative impact to click propensity. Consequently, such models promote
business metrics such as revenue. An alternative here is content which at times is not relevant to customer’s
onto attribute value to individual content only if interaction going shopping mission. This distracts the customer from
with it is in addition to purchase of the page item (or their mission which in turn results in a negative impact
product) wherever applicable. In empirical evaluation, to their shopping experience. VTA on the other hand
allows us to capture both the positive and negative im- observed down-session value as reward without
accountpact of presenting content to customers. It enables us ing for the counterfactual outcome could lead to models
to capture the value of showing content which inspires overestimating the predicted benefit at inference time.
customer shopping missions including scenarios where To address these challenges, we use an uplift modeling
customers can compare selection without requiring di- framework. It estimates Conditional Average Treatment
rect interaction with content. Furthermore, it is closer in Efect (CATE) [ 28] between exposure and non-exposure
alignment with how an experimentation framework for of content to customers using observational data. We
conducting online A/B tests may measure and attribute assume conditional unconfoundedness in our problem
aggregate downstream impact after a customer has been setting [29, 30]. Then,
exposed to a new shopping experience (or treatment).
          </p>
          <p>
            Thus, it can enable parity in the methodology used to   ≡ [(
            <xref ref-type="bibr" rid="ref1">1</xref>
            ) − (0)| = ] (
            <xref ref-type="bibr" rid="ref2">2</xref>
            )
attribute reward in the content optimization and online = [(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )| = ] − [(0)| = ]
experimentation systems.
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>4.5. Uplift Modeling Framework</title>
          <p>Both VTA and CTA assume a causal relationship between
customers interacting with shown content through views
and clicks (cause), and observed reward (efect). In the
case of CTA, there is a strong connection between the
cause and efect as often times a click is an intentional
action on the part of a customer. However, with VTA
we cannot establish this direct connection between a
customer viewing content and the observed reward. As
such, we assume a causal relationship which introduces
noise in our observations. Models incrementally trained
using such observations are likely to have a high variance
in the predictions.</p>
          <p>
            In addition, the attribution model described in the
previous section does not capture the incremental value of
showing content. Customers can have an underlying
propensity to shop products or consume content based
on prior exposure or afinity. As a result, attributing
where,  is the d-dimensional feature vector. Here,
[(
            <xref ref-type="bibr" rid="ref1">1</xref>
            )| = ] is the mean of the treated group
where content  is shown in the shopping session, and
[(0)| = ] is the mean of the untreated group where
content  is not shown in the shopping session. We have
explored two approaches to estimate the latter: i) using
the mean of untreated group calculated from our ranking
logs as a biased estimate for [(0)| = ], and ii) by
estimating [(0)| = ] from randomized controlled
trials.
          </p>
          <p>The uplift modeling framework is then defined using
a two-part model. First, a baseline model estimates the
expected counterfactual reward when content  is ranked
but not shown in the shopping session. We illustrate the
underlying theory using a linear regression model:
 0 = ⊤ +</p>
          <p>
            (
            <xref ref-type="bibr" rid="ref3">3</xref>
            )
Treatment or incremental efect for each observation in
the treated group where content  is ranked and shown
in the shopping session is then estimated as: parameters by a small fraction of their existing value
to account for changes in the environment [31]. This
(1) =  − ˆ0() (
            <xref ref-type="bibr" rid="ref4">4</xref>
            ) completes the feedback loop which allows us to
continwhere,  is the observed down-session reward for ob- uously explore actions and expand our knowledge for
servation , and (1) is the imputed incremental efect making better decisions in the future. This in turn
enfor observation  in the treated group. In the second ables us to support running online A/B tests using which
part, pseudo-efect  is used as the target variable in content creators can introduce new content across
Amaour ranking model, described in (eqn. 1), to predict the zon’s retail website with the goal of improving customer
incremental benefit of showing content  to a customer shopping experience and measure its benefit while doing
in widget group . so.
          </p>
        </sec>
        <sec id="sec-3-1-3">
          <title>4.6. Exploration Strategy</title>
        </sec>
        <sec id="sec-3-1-4">
          <title>4.7. Incorporating Diversity in Ranking</title>
          <p>The exploration component of our content optimization Showing high-relevance content without taking content
framework explores content with few observations from diversity into account leads to monotony and tends to
the past. To do so, it aims at solving a contextual bandit make the holistic shopping experience less meaningful
problem. Here, we use Thompson sampling, an algo- for the customers. Optimizing for the whole widget group
rithm widely used to balance exploration and exploita- involves balancing relevance and diversity of the content
tion. It suggests to randomly play each arm according to therein, where the whole-widget group efect is
repreits probability of being optimal. In our problem setting, it sented using the amount of similar content displayed
means choosing content proportional to the probability in it. One approach followed here is to model this as a
of it being optimal. This implies we won’t be necessarily submodular optimization problem [32]. In [33, 34], the
choosing content with the highest expected incremental authors propose using submodular functions which have
benefit at each time step. It is a trade-of we make to ex- a diminishing returns property. In their approach, the
plore content with few observations from the past which total score for a content is derived from its relevance
have high uncertainty but ultimately may drive a higher while also accounting for the decreasing utility of
showreward. In practice, we apply the Thompson sampling ing multiple content of the same type. As a result, the
algorithm by sampling model parameters ˆ from their value of selecting content from a given category or type
posterior distributions followed by choosing content that decreases as a function of the number of content
belongmaximizes the reward. ing to that type already selected. A key shortcoming of
this approach is that it lacks a feedback loop and
paramAlgorithm 1 Thompson Sampling Algorithm for Con- eters of the diversity scoring function aren’t learned to
tent Optimization optimize for the same objective as the relevance scoring
1: for  = 1, . . . ,  do function.
2: for all  = 1, . . . ,  do ◁  is the value Instead we propose a two-stage model for
incorporatof  corresponding to widget group  ing diversity into content. We first rank all the eligible
3: Receive context  content  ∈  to be shown in widget group  using our
4: Sample ˆ from the posterior distribution underlying ranking model (eqn. 1). Thereafter, we
iter</p>
          <p>Pr( |) atively re-rank content at each position  in the widget
5: Select  = argmax ⊤ˆ  group by taking into account the content that is already
6: end for ranked in the previous  − 1 positions. This is
accom7: Choose  −  arms and observe reward  plished by using a second ranking model which includes
8:  = − 1 ∪ cross-content interaction features. To capture these
in{( ,  ,  ),  = 1, . . . , } teractions, we categorize each content  in to one of 
9: end for distinct categories or types. The goal is to then select
an optimal number of highly relevant widgets in each
category. For a given widget group  ∈ , the overall
value of widget group  is represented as:</p>
          <p>
            Candidates that are ranked and chosen to be displayed
by the ranking model are then logged along with their
observed reward in the form of triplets (, , ). There- 
after, we estimate the incremental efect for each obser-  (|, ) = ∑︁  (|(− 1) , ,  ) (
            <xref ref-type="bibr" rid="ref5">5</xref>
            )
vation in the logged feedback using our uplift modeling =1
framework. The incremental efects are subsequently  (|(− 1) , ,  ) = (⊤ ) (
            <xref ref-type="bibr" rid="ref6">6</xref>
            )
used as target variables to incrementally train our
ranking model using a batch update under the Thompson Sam- where,  is the content at rank  in widget group ,
pling framework. While doing so, we decay the model (− 1) is the set of content allocated in the top  − 1
positions,  is the request and customer context received than before as we retrained the models at a faster
caas before,  are the model parameters, and  is a gener- dence. This impacts the model’s learning process in two
alized linear model. ways: i) outliers in the dataset can cause the model to
incorrectly associate higher potential reward for some
content despite winsorization techniques, and ii)
insuf5. Low Latency Learning ifcient data can limit the model’s ability to learn about
Framework content’s reward distribution especially in regions with
small amount of trafic. Empirically, we observe that both
In production, we observed that our content optimiza- of these scenarios lead to over-exploration of content.
tion framework sufered from a delay in feedback as it We address this challenge using Bayesian
Regularizaneeded multiple days on average to complete the learning tion. The use of Gaussian priors has been established in
loop. This involves logging the feedback after customers [35] as a form of L2 regularization wherein the following
are shown with content on the website, measuring and equivalence is explored:
attributing reward in our data pipelines, incrementally
training the models at a daily cadence, and deploying ( +  )− 1  = E[︁  (| ) ( ) ]︁ (
            <xref ref-type="bibr" rid="ref7">7</xref>
            )
the retrained models in production. A delay in feedback  ()
has the following consequences in production for our
content optimization framework:
Here, instead of initiating the feature weights from a
static prior (i.e. mean 0 and variance 1), we derive a
• When new content is introduced into the ecosys- prior distribution from previously learned weight
distritem, the optimization framework is not able to butions in order to allow for a more pessimistic
exploefectively estimate its potential benefit, and the ration regime. For instance, by using a prior representing
content is subject to exploration as expected. Due 20th percentile mean of all features and 75th percentile
to the delay in feedback, this can result in new variance of all features, we can reduce the chances of
content being explored at a higher show-rate over-exposure for new content during the learning
pefor the duration of the delay during the initial riod. This improvement in turn has enabled us to launch
learning period before suficient observations are the L3 framework in production and reduced the delay
logged for the model to learn from its own feed- in feedback by 90%.
back. This in turn can result in sub-optimal
decision making and introduction of poor customer
shopping experience during the learning period. 6. Experiments
• While the cost of exploration may be amortized We first evaluate our content optimization framework
over long running content campaigns, a longer using both traditional ofline evaluation and of-policy
feedback loop induces limitations in realizing ben- evaluation methodologies [36, 37]. This allows us to
evalefit during high-value events such as Cyber Mon- uate and prune alternative treatment policies before
introday where new content may be introduced for a ducing them in online randomized experiments [38, 39].
short period of time. In such cases, some content Here, we use regression and ranking metrics to
evalupromoting sales or other events will be turned of ate the framework quantitatively, and content’s
share-ofeven before their benefit is efectively learned by voice and ranking distributions to evaluate it qualitatively
the model. using domain knowledge. Subsequently, we demonstrate
• A longer feedback loop also decreases the velocity the efectiveness of our framework through five online
of running online A/B tests where new content A/B tests.
and improvements to optimization framework
may be introduced with the aim of improving
customer shopping experience. 6.1. Online Experiment Setting
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>To address these challenges, we developed a Low Latency Learning (L3) framework which has reduced the learning loop for our content optimization framework by 90% from multiple days to a couple of hours.</title>
        <sec id="sec-3-2-1">
          <title>5.1. Bayesian Regularization</title>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>A key challenge we encountered in the development of</title>
        <p>L3 framework was that of the number of samples
available for incremental training being significantly lower
In our online experimentation setting, observational
units (or shopping sessions) are randomly exposed to
either the baseline control policy or the alternative
treatment policies. Here, we track the impact to our metric
of interest  , which is a measure of improved
sitewide customer shopping experience. In the results, we
include the causal efect w.r.t percentage improvement
in this metric at Amazon’s scale. The experiments are
conducted across all of Amazon’s world-wide
marketplaces and product categories. Level of significance  for
these experiments was determined by Amazon’s business
objectives and was set to 0.10. Duration for these
experiments was estimated from statistical power analysis. We
allocated equal trafic to both the control and treatment
groups. During the course of the experiment, the models
were incrementally trained using their own set of logged
feedback.</p>
        <sec id="sec-3-3-1">
          <title>6.2. Experiment 1: Application of the</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>Holistic Optimization Framework</title>
          <p>We first test the efectiveness of our holistic optimization
framework to rank content in a widget group on product
detail pages of Amazon’e retail website. This is a region
of customer shopping experience on the website where
we usually see organic content such as ‘customers who
viewed this also viewed’ and ‘customers who bought this
also bought’ widgets being displayed alongside
advertising content. A key challenge in dynamically ranking
content in this setting was that of attribution of reward
to diverse type of content which were generated by
content creators who optimized for difering business
objectives. As such, our framework needed to arbitrate
content during the content allocation process and fairly
balance the difering objectives. Since holistic
optimization framework measures reward using the aggregate
down-session value after customer has interacted with
content, we wanted to test its efectiveness in addressing
this problem. In the control group, content was statically
ranked by a rule-based system while in the treatment
group, our framework dynamically ranked content using
the holistic optimization framework. In the results
(EXP1), we observe a practically and statistically significant
improvement in the   metric which is a measure of
site-wide improvement in customer shopping experience.</p>
        </sec>
        <sec id="sec-3-3-3">
          <title>6.3. Experiment 2: Application of</title>
        </sec>
        <sec id="sec-3-3-4">
          <title>View-through Attribution</title>
          <p>In this experiment, we applied our content optimization
framework to the image size selection problem. Usually,
product display images on Amazon’s detail page exist in
three sizes – small, medium and large. Here, the size of a
rendered image can influence the customer’s
understanding of the product. Hence, we want to select and render
an optimal size of the same product image so as to help
the customers evaluate products better especially for high
consideration purchases. This is a use case where
clickthrough attribution cannot be used as clicking on the
content does not necessarily indicate a positive customer
shopping experience. We formulate the task of optimal
image size selection as a learning to rank problem, and
use view-through attribution to measure and attribute
reward to the rendered image size. To demonstrate the
efectiveness of this approach, we ran two experiments –
one each for desktop and mobile surfaces. In the control
group, image size was selected by a rule-based system
while in the treatment group, our framework ranked the
image size variations and chose the top ranked variation
to render. In both the experiments (EXP-2A and EXP-2B),
we observe an improvement in the   metric which
is practically significant at Amazon’s scale.</p>
        </sec>
        <sec id="sec-3-3-5">
          <title>6.4. Experiment 3: Application of the</title>
        </sec>
        <sec id="sec-3-3-6">
          <title>Causal Bandit Framework</title>
          <p>After demonstrating the efectiveness of VTA, we tested
the utility of the uplift modeling framework. The
framework allows us to measure and optimize for incremental
value generated by content, and reduces the
observational bias in data. We conducted an experiment in a
widget group which is located at the bottom of product
detail pages on the desktop retail website where
personalized content that is usually generated by taking recent
browsing history into account is shown. This in turn
allowed us to test our hypothesis that customers can
have an underlying propensity to shop products or
consume content based on prior exposure or afinity, and
optimizing for incremental benefit can result in a
positive customer shopping experience. In the control group,
content was ranked by a linear bandit without using the
uplift modeling framework while rewards were measured
and attributed using CTA. In the treatment group,
content was ranked using a linear causal bandit with VTA.
In the results (EXP-3), we observe that the linear causal
bandit using VTA performed better than the linear
bandit which did not use uplift modeling framework. The
improvement in   metric was both practically and
statistically significant.</p>
        </sec>
        <sec id="sec-3-3-7">
          <title>6.5. Experiment 4: Application of</title>
        </sec>
        <sec id="sec-3-3-8">
          <title>Incorporating Diversity in Ranking</title>
          <p>Subsequently, we ran an experiment (EXP-4) on the
desktop homepage of Amazon’s retail website to test the
impact of incorporating diversity in content. In the control
group, content was ranked using just the single baseline
ranking model, while in the treatment group, content learnings from the deployment of a low-latency learning
was ranked using the two-stage ranking model – first framework in production that has reduced the delay in
using the baseline model followed by a re-ranking model feedback and shortened the learning loop by 90%. Here,
which incorporates diversity using cross-content interac- we described our application of Gaussian prior as a form
tion features. Here, we observe a practically significant of L2 regularization which in turn enabled the launch
improvement in the   metric. Based on the results, of the L3 framework. We then demonstrated the
efecwe infer that incorporating diversity into content ranking tiveness of our methodology through multiple online
can lead to a better customer shopping experience. experiments, and shared results and insights gathered
through the same. Finally, we believe our methodology
6.6. Experiment 5: Application of the and learnings are generic and can be extended to
content optimization problems in other domains. It can also
Low Latency Learning Framework be extended to rank items (or products) within a single
widget for a product recommendation system.</p>
          <p>L3 pipeline has shortened the delay in feedback for our
contextual bandit based framework by 90%. As a result,
we expect the bandit retrained at hourly cadence to
converge sooner and perform better than the one retrained
at a daily cadence. To test the benefit and measure the
impact of low latency learning, we ran an experiment
on the mobile homepage of Amazon’s retail website. In
the control group, content was ranked by a linear bandit
incrementally trained at a slower cadence with a learning
loop of multiple days, while in the treatment group,
content was ranked by a linear bandit incrementally trained
at a faster cadence with a learning loop of a few hours.</p>
          <p>In the results (EXP-5), we observe that the bandit with
a shorter delay in feedback performed better w.r.t our
metric of interest   where the improvement was
both practically and statistically significant. Based on the
results, we infer that reducing the delay in feedback and
increasing the velocity of learning loop has a positive
impact on customer shopping experience.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>7. Conclusion</title>
      <sec id="sec-4-1">
        <title>In this paper, we presented a causal bandit framework</title>
        <p>to address the problem of content optimization with the
objective of improving the overall customer shopping
experience on Amazon’s e-commerce (or retail) website.</p>
        <p>Therein, we introduced a holistic optimization
framework that enables us to define reward and rank diverse
types of content using aggregate down-session value;
presented the concept of view-through attribution;
discussed how it addresses some of the shortcomings of
click-through attribution; and presented applications of
VTA in ranking content belonging to diverse type. To
address the shortcomings of view-through attribution, we
used an Uplift modeling framework which has enabled
us to rank content using incremental or causal
beneift instead of overall value. Subsequently, we proposed
a two-stage model to incorporate diversity in content
ranking by using cross-content interaction features. It
helps us to balance relevance with diversity in content
shown on Amazon’s retail website and provide a
meaningful experience to our customers. Thereafter, we shared
national Conference on Learning Representations, ternational Conference on Predictive Applications
2018. and APIs, PMLR, 2017, pp. 1–13.
[9] M. Fortunato, M. G. Azar, B. Piot, J. Menick, M. Hes- [23] S. R. Künzel, J. S. Sekhon, P. J. Bickel, B. Yu,
Metsel, I. Osband, A. Graves, V. Mnih, R. Munos, D. Has- alearners for estimating heterogeneous treatment
sabis, O. Pietquin, C. Blundell, S. Legg, Noisy net- efects using machine learning, Proceedings of the
works for exploration, in: International Conference national academy of sciences 116 (2019) 4156–4165.
on Learning Representations, 2018. [24] Z. Zhao, T. Harinen, Uplift modeling for multiple
[10] I. Osband, B. Van Roy, Bootstrapped thompson sam- treatments with cost optimization, in: 2019 IEEE
pling and deep exploration, 2015. doi:10.48550/A International Conference on Data Science and
AdRXIV.1507.00300. vanced Analytics (DSAA), IEEE, 2019, pp. 422–431.
[11] I. Osband, C. Blundell, A. Pritzel, B. Van Roy, Deep [25] N. Sawant, C. B. Namballa, N. Sadagopan, H.
Nasexploration via bootstrapped dqn, in: Advances in sif, Contextual multi-armed bandits for causal
Neural Information Processing Systems 29, Curran marketing, in: ICML 2018, 2018. URL: h t t p s :
Associates, Inc., 2016, pp. 4026–4034. //www.amazon.science/publications/contextu
[12] D. Russo, B. V. Roy, A. Kazerouni, I. Osband, A tuto- al-multi-armed-bandits-f or-causal-marketing.
rial on thompson sampling, CoRR abs/1707.02038 [26] Y. Zhao, M. Goodman, S. Kanase, S. Xu, Y. Kimmel,
(2017). arXiv:1707.02038. B. Payne, S. Khan, P. Grao, Mitigating targeting
[13] W. R. Thompson, On the likelihood that one un- bias in content recommendation with causal
banknown probability exceeds another in view of the dits, in: Proceedings of the 2nd Workshop on
Multievidence of two samples, Biometrika 25 (1933) 285– Objective Recommender Systems (MORS 2022), in
294. conjunction with the 16th ACM Conference on
[14] M. Strens, A bayesian framework for reinforcement Recommender Systems (RecSys 2022), Seattle, WA,
learning, in: ICML, volume 2000, 2000, pp. 943–950. USA, 2022.
[15] S. L. Scott, A modern bayesian look at the multi- [27] S. Xu, Y. Zhao, S. Kanase, M. Goodman, S. Khan,
armed bandit, Applied Stochastic Models in Busi- B. Payne, P. Grao, Machine learning attribution:
ness and Industry 26 (2010) 639–658. Inferring item-level impact from slate
recommen[16] L. Li, W. Chu, J. Langford, R. E. Schapire, A dation in e-commerce, in: KDD 2022 Workshop
contextual-bandit approach to personalized news on First Content Understanding and Generation for
article recommendation, in: Proceedings of the e-Commerce, 2022. URL: https://www.amazon.sci
19th International Conference on World Wide Web, ence/publications/machine-learning-attribution-i
WWW ’10, Association for Computing Machinery, nf erring-item-level-impact-f rom-slate-recomm
New York, NY, USA, 2010, p. 661–670. doi:10.114 endation-in-e-commerce.</p>
        <p>5/1772690.1772758. [28] S. Athey, G. Imbens, Recursive partitioning for
[17] O. Chapelle, L. Li, An empirical evaluation of heterogeneous causal efects, Proceedings of the
thompson sampling, in: Proceedings of the 24th National Academy of Sciences 113 (2016) 7353–7360.
International Conference on Neural Information doi:10.1073/pnas.1510489113.
Processing Systems, NIPS’11, Curran Associates [29] D. Rubin, Estimating causal efects of treatments in
Inc., Red Hook, NY, USA, 2011, p. 2249–2257. randomized and nonrandomized studies., Journal
[18] S. Agrawal, N. Goyal, Thompson sampling for con- of Educational Psychology 66 (1974) 688–701.
textual bandits with linear payofs, in: International [30] G. W. Imbens, D. B. Rubin, Causal Inference for
Conference on Machine Learning, 2013, pp. 127– Statistics, Social, and Biomedical Sciences: An
135. Introduction, Cambridge University Press, 2015.
[19] B. Hansotia, B. Rukstales, Incremental value mod- doi:10.1017/CBO9781139025751.
eling, Journal of Interactive Marketing 16 (2002) [31] T. Graepel, J. Q. n. Candela, T. Borchert, R. Herbrich,
35. Web-scale bayesian click-through rate prediction
[20] V. S. Y. Lo, The true lift model: A novel data mining for sponsored search advertising in microsoft’s bing
approach to response modeling in database mar- search engine, in: Proceedings of the 27th
Internaketing, SIGKDD Explor. Newsl. 4 (2002) 78–86. tional Conference on International Conference on
doi:10.1145/772862.772872. Machine Learning, ICML’10, Omnipress, Madison,
[21] N. Radclife, Using control groups to target on WI, USA, 2010, p. 13–20.</p>
        <p>predicted lift: Building and assessing uplift model, [32] Y. Yue, C. Guestrin, Linear submodular bandits and
Direct Marketing Analytics Journal (2007) 14–21. their application to diversified retrieval, in:
Ad[22] P. Gutierrez, J.-Y. Gérardy, Causal inference and vances in Neural Information Processing Systems,
uplift modelling: A review of the literature, in: In- volume 24, Curran Associates, Inc., 2011.
[33] C. H. Teo, H. Nassif, D. Hill, S. Srinivasan, M. Good- D. Coey, M. Curtis, A. Deng, W. Duan, P. Forbes,
man, V. Mohan, S. Vishwanathan, Adaptive, person- B. Frasca, T. Guy, G. W. Imbens, G. Saint Jacques,
alized diversity for visual discovery, in: Proceed- P. Kantawala, I. Katsev, M. Katzwer, M.
Konutings of the 10th ACM Conference on Recommender gan, E. Kunakova, M. Lee, M. Lee, J. Liu, J.
McSystems, RecSys ’16, Association for Computing Queen, A. Najmi, B. Smith, V. Trehan, L. Vermeer,
Machinery, New York, NY, USA, 2016, p. 35–38. T. Walker, J. Wong, I. Yashkov, Top challenges from
[34] H. Nassif, K. O. Cansizlar, M. Goodman, S. V. N. the first practical online controlled experiments
Vishwanathan, Diversifying music recommenda- summit, SIGKDD Explor. Newsl. 21 (2019) 20–35.
tions, in: ICML 2016, 2016. doi:10.1145/3331651.3331655.
[35] M. A. Figueiredo, Adaptive sparseness for super- [39] T. S. Richardson, Y. Liu, J. Mcqueen, D. Hains, A
vised learning, IEEE transactions on pattern analy- bayesian model for online activity sample sizes, in:
sis and machine intelligence 25 (2003) 1150–1159. Proceedings of The 25th International Conference
[36] A. Swaminathan, T. Joachims, The self-normalized on Artificial Intelligence and Statistics, volume 151
estimator for counterfactual learning, in: Advances of Proceedings of Machine Learning Research, PMLR,
in Neural Information Processing Systems, vol- 2022, pp. 1775–1785.</p>
        <p>ume 28, Curran Associates, Inc., 2015.
[37] T. Schnabel, A. Swaminathan, A. Singh, N.
Chandak, T. Joachims, Recommendations as treatments: APPENDIX
Debiasing learning and evaluation, in: Proceedings
of the 33rd International Conference on
International Conference on Machine Learning - Volume A. An Illustration of Diverse Type of
48, ICML’16, JMLR.org, 2016, p. 1670–1679. Content on Amazon’s Homepage
[38] S. Gupta, R. Kohavi, D. Tang, Y. Xu, R.
Andersen, E. Bakshy, N. Cardin, S. Chandran, N. Chen,</p>
      </sec>
      <sec id="sec-4-2">
        <title>Below, (figure 3) illustrates diverse type of content being shown on the homepage of Amazon’s retail website.</title>
        <p>Figure 3: Homepage of Amazon’s Retail Website.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Riquelme</surname>
          </string-name>
          , G. Tucker,
          <string-name>
            <given-names>J.</given-names>
            <surname>Snoek</surname>
          </string-name>
          ,
          <article-title>Deep bayesian bandits showdown: An empirical comparison of bayesian deep networks for thompson sampling</article-title>
          ,
          <source>in: International Conference on Learning Representations</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bietti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Langford</surname>
          </string-name>
          ,
          <article-title>A contextual bandit bake-of, arXiv preprint</article-title>
          arXiv:
          <year>1802</year>
          .
          <volume>04064</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Mnih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Rusu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Veness</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Bellemare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Riedmiller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fidjeland</surname>
          </string-name>
          , G. Ostrovski,
          <string-name>
            <given-names>S.</given-names>
            <surname>Petersen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Beattie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sadik</surname>
          </string-name>
          , I. Antonoglou,
          <string-name>
            <given-names>H.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kumaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wierstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Legg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hassabis</surname>
          </string-name>
          ,
          <article-title>Human-level control through deep reinforcement learning</article-title>
          ,
          <source>Nature</source>
          <volume>518</volume>
          (
          <year>2015</year>
          )
          <fpage>529</fpage>
          -
          <lpage>533</lpage>
          . doi:
          <volume>10</volume>
          .1038/nature14236.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Schaul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Quan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Antonoglou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <article-title>Prioritized experience replay</article-title>
          ,
          <source>in: 4th International Conference on Learning Representations, ICLR</source>
          <year>2016</year>
          , San Juan, Puerto Rico, May 2-
          <issue>4</issue>
          ,
          <year>2016</year>
          , Conference Track Proceedings,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Robbins</surname>
          </string-name>
          ,
          <article-title>Asymptotically eficient adaptive allocation rules</article-title>
          ,
          <source>Advances in Applied Mathematics</source>
          <volume>6</volume>
          (
          <year>1985</year>
          )
          <fpage>4</fpage>
          -
          <lpage>22</lpage>
          . doi:https://doi.org/10.1016/
          <fpage>0196</fpage>
          -
          <lpage>8858</lpage>
          (
          <issue>85</issue>
          )
          <fpage>90002</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Auer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Cesa-Bianchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <article-title>Finite-time analysis of the multiarmed bandit problem</article-title>
          ,
          <source>Mach. Learn</source>
          .
          <volume>47</volume>
          (
          <year>2002</year>
          )
          <fpage>235</fpage>
          -
          <lpage>256</lpage>
          . doi:
          <volume>10</volume>
          .1023/A:
          <fpage>101368</fpage>
          <lpage>9704352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <article-title>Dropout as a bayesian approximation: Representing model uncertainty in deep learning</article-title>
          ,
          <source>in: Proceedings of the 33rd International Conference on International Conference on Machine Learning - Volume 48, ICML'16</source>
          , JMLR.org,
          <year>2016</year>
          , p.
          <fpage>1050</fpage>
          -
          <lpage>1059</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Plappert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Houthooft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sidor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Asfour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Abbeel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Andrychowicz</surname>
          </string-name>
          ,
          <article-title>Parameter space noise for exploration</article-title>
          , in: Inter-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>