<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Delayed Rewards in the context of Reinforcement Learning based Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Debmalya Biswas</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>We present a Reinforcement Learning (RL) based approach to implement Recommender systems. The results are based on a real-life Wellness app that is able to provide personalized health related content to users in an interactive fashion. Unfortunately, current recommender systems are unable to adapt to continuously evolving features, e.g. user sentiment, and scenarios where the RL reward needs to be computed based on multiple and unreliable feedback channels (e.g., sensors, wearables). To overcome this, we propose three constructs: (i) weighted feedback channels, (ii) delayed rewards, and (iii) rewards boosting, which we believe are essential for RL to be used in Recommender Systems. Finally, we also provide some implementation details on how the Wellness App based on Azure Personalizer was extended to accommodate the above RL constructs.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Wellness apps have historically suffered from low adoption rates.
Personalized recommendations have the potential of improving
adoption, by making increasingly relevant and timely
recommendations to users. While recommendation engines (and consequently, the
apps based on them) have grown in maturity, they still suffer from
the ‘cold start’ problem and the fact that it is basically a push-based
mechanism lacking the level of interactivity needed to make such
apps appealing to millennials.</p>
      <p>We present a Wellness app case-study where we applied a
combination of Reinforcement Learning (RL) and Natural Language
Processing (NLP)/Chatbots to provide a highly personalized and
interactive experience to users. We focus on the interactive aspect of the
app, where the app is able to profile and converse with users in
realtime, providing relevant content adapted to the current sentiment and
past preferences of the user.</p>
      <p>
        The core of such chatbots is an intent recognition Natural
Language Understanding (NLU) engine [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], which is trained with
hardcoded examples of question variations. When no intent is matched
with a confidence level above 30%, the chatbot returns a fallback
answer. The user sentiment is computed based on both the
(explicit) user response and (implicit) environmental aspects, e.g.
location (home, office, market, . . . ), temperature, lighting, time of the
day, weather, other family members present in the vicinity, and so
on; to further adapt the chatbot response.
      </p>
      <p>
        RL refers to a branch of Artificial Intelligence (AI), which is able
to achieve complex goals by maximizing a reward function in
realtime. The reward function works similar to incentivizing a child with
candy and spankings, such that the algorithm is penalized when it
takes a wrong decision and rewarded when it takes a right one – this is
reinforcement. The reinforcement aspect also allows it to adapt faster
to real-time changes in user sentiment. For a detailed introduction to
RL frameworks, the interested reader is referred to [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Previous works have explored RL in the context of Recommender
Systems [
        <xref ref-type="bibr" rid="ref10 ref5 ref7">5, 7, 10</xref>
        ], and enterprise adoption also seems to be gaining
momentum with the recent availability of Cloud APIs (e.g. Azure
Personalizer [
        <xref ref-type="bibr" rid="ref2 ref6">2, 6</xref>
        ]) and Google’s RecSim [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However, they still
work like a typical Recommender System. Given a user profile and
categorized recommendations, the system makes a recommendation
based on popularity, interests, demographics, frequency and other
features. The main novelty of these systems is that they are able
to identify the features (or combination of features) of
recommendations getting higher rewards for a specific user; which can then
be customized for that user to provide better recommendations.
Unfortunately, this is still inefficient for real-life systems which need
to adapt to continuously evolving features, e.g. user sentiment, and
where the reward needs to computed based on multiple and
unreliable feedback channels (e.g., sensors, wearables).
      </p>
      <p>
        The rest of the paper is organized as follows: Section 2 outlines the
problem scenario and formulates it as an RL problem. In Section 3,
we propose three RL constructs needed to overcome the above
limitations: (i) weighted feedback channels, (ii) delayed rewards, and (iii)
rewards boosting, which we believe are essential constructs for RL to
be used in Recommender Systems. ‘Delayed Rewards’ in this context
is different from the notion of Delayed RL [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], where rewards in the
distant future are not considered as valuable as immediate rewards. In
our notion of ‘Delayed Rewards’, a received reward is only applied
after its consistency has been validated by a subsequent action. We
provide implementation details in Section 4, on how to extend an RL
powered Wellness App based on Azure Personalizer, to
accommodate the above constructs. Section 5 concludes the paper providing
some directions for future work.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>PROBLEM SCENARIO</title>
      <p>In this section, we set the problem context and formulate it as a
Reinforcement Learning problem.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Wellness App</title>
      <p>The Wellness app supports both push based notifications, where
personalized health, fitness, activity, etc. related recommendations are
pushed to the user; as well as interactive chats where the app reacts
in response to a user query. We assume the existence of a
knowledgebase KB of articles, pictures and videos, with the artifacts ranked
according to their relevance to different user profiles / sentiments.</p>
      <p>The Wellness app architecture is described in Fig. 1, which shows
how the user response and environmental conditions are:
1. gathered using available sensors to compute the ‘current’
feedback, including environmental context (e.g. webcam pic of the
user can be used to infer the user sentiment to a chatbot response
/ notification, the room lighting conditions and other users present
in the vicinity),
2. which is then combined with the user conversation history to
quantify the user sentiment curve and discount any sudden
changes in sentiment due to unrelated factors;
3. leading to the aggregate reward value corresponding to the last
chatbot response / app notification provided to the user.</p>
      <p>
        This reward value is then provided as feedback to the RL agent, to
choose the next optimal chatbot response / app notification from the
knowledgebase. It is worthwhile noting here that capturing the user
sentiment, esp. the environmental aspects, requires a high degree of
knowledge regarding the user context. As such, we need to perform
this in a privacy preserving fashion. We suffice to say here that
appropriate privacy protections are provided by the ‘Privacy’ block in
Fig. 1, and further details are provided in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>RL Formulation</title>
      <p>We formulate the RL Engine for the above scenario as follows
(illustrated in Fig. 2):</p>
      <p>Action (a): An action a in this case corresponds to a KB article
which is delivered to the user either as a push notification, or in
response to a user query, or as part of an ongoing conversation.
Agent (A): is the one performing actions. In this case, the Agent is
the App delivering actions to the users, where an action is selected
based on its Policy.</p>
      <p>Policy ( ): is the strategy that the agent employs to select the next
best action. Given a user profile Up, (current) sentiment Us, and
query Uq; the Policy function computes the product of the article
scores returned by the NLP and Recommendation Engines
respectively, selecting the item with the highest score as the next best
action:
– The NLP Engine (N E) parses the query and outputs a score
for each KB article, based on the “text similarity” of the article
to the user query.
– Similarly, the Recommendation Engine (RE) provides a score
for each article based on the reward associated with each
article, with respect to the user profile and sentiment. The Policy
function can be formalized as follows:
(Up; Us; Uq) = a j max[N E(a; Uq)
a</p>
      <p>RE(a; Up; Us)]
Reward (r): refers to the feedback by which we measure the
success or failure of an agent’s recommended action. The feedback
can e.g. refer to the amount of time that a user spends reading a
recommended article. We consider a 2-step reward function
computation where the feedback fa received with respect to a
recommended action is first mapped to a sentiment score, which is then
mapped to a reward.
(1)
r(a; fa) = s(fa)
(2)
where r and s refer to the reward and sentiment functions,
respectively. Once computed, the KB is updated with the computed
reward / sentiment for the corresponding action.
3</p>
    </sec>
    <sec id="sec-5">
      <title>RL REWARD AND POLICY EXTENSIONS</title>
      <p>In this section, we show how the Reward and Policy functions are
extended to accommodate the real-life challenges posed by our RL
based Wellness App.
3.1</p>
    </sec>
    <sec id="sec-6">
      <title>Weighted (Multiple) Feedback Channels</title>
      <p>As described in Fig. 1, we consider a multi-feedback channel, with
feedback captured from user (edge) devices / sensors, e.g. webcam,
thermostat, smartwatch, or a camera, microphone, accelerometer
embedded within the mobile device hosting the app. For instance, a
webcam frame capturing the facial expression of the user, heart rate
provided by the user smartwatch, can be considered together with
the user provided text response “Thanks for the great suggestion”; in
computing the user sentiment to a recommended action.</p>
      <p>Let ffa1 ; fa2 ; :::fan g denote the feedback received for action a.
Recall that s(f ) denotes the user sentiment computed independently
based on the respective sensory feedback f . The user sentiment
computation can be considered as a classifier outputting a value between
1-10. The reward can then be computed as a weighted average of the
sentiment scores, denoted below:
n
X(wi
i=1
ra(ffa1 ; fa2 ; :::fan g) =
s(fai ))
(3)
where the weights fwa1; wa2; :::wang allow the system to
harmonize the received feedback, as some feedback channels may suffer
from low reliability issues. For instance, if fi corresponds to a user
typed response, fj corresponds to a webcam snapshot; then higher
weightage is given to fi. The reasoning here is that the user might
be ‘smiling’ in the snapshot, however the ‘smile’ is due to his kid
entering the room (also captured in the frame), and not necessarily
in response to the received recommendation / action. At the same
time, if the sentiment computed based on the user text response
indicates that he/she is ‘stressed’, then we give higher weightage to user
explicit (text response) feedback in this case.
3.2</p>
    </sec>
    <sec id="sec-7">
      <title>Delayed Rewards</title>
      <p>A ‘delayed rewards’ strategy is applied in the case of reward
inconsistency, where the current (computed) reward is ‘negative’ for an
action to which the user has been known to react positively in the
past; or vice versa. For instance, let us consider that the user
sentiment is low for a recommendation of category ‘Shopping’, to which
the user has been known to react very positively (to other ‘Shopping’
related recommendations) in the past. Given such inconsistency, the
delayed rewards strategy buffers the computed reward rat for action
at at time t; and provides an indication to the RL Agent-Policy ( )
to try another recommendation of the same type (‘Shopping’) - to
validate the user sentiment - before updating the rewards for both at
and at+1 at time t + 1.</p>
      <p>To accommodate the ‘delayed rewards’ strategy, the rewards
function is extended with a memory buffer that allows the rewards of
last m actions [at+m; at+m 1; :::; at] to be aggregated and applied
retroactively at time (t + m). The delayed rewards function dr is
denoted as follows:
drati 2 fat;at+1;:::;at+mg j (t + m) =
rati )</p>
      <p>(4)
m
X(wi
i=0
where j t + m implies that the reward for the actions
[at+m; at+m 1; :::; at], although computed individually; can only
be applied at time (t + m). As before, the respective weights wi
allow us to harmonize the effect of an inconsistent feedback, where
the reward for an action ati is applied based on the reward computed
for a later action a(t+1)i.</p>
      <p>To effectively enforce the ‘delayed rewards’ strategy, the Policy
is also extended to recommend an action of the same type, as the
previous recommended action; if the delay flag d is set (d = 1): The
”delayed” Policy dt is outlined below:
d = 1 :
d = 0 :</p>
      <p>dt(Up; Us; Uq) =
at j rat rat 1
a j max[N E(a; Uq)
a</p>
      <p>RE(a; Up; Us)]
(5)</p>
      <p>The RL formulation extended with delayed reward / policy is
illustrated in Fig. 3.
Rewards boosting, or rather rewards normalization, applies mainly to
continuous chat interactions. In such cases, if the user sentiment for a
recommended action is ‘negative’; it might not be the fault of the last
action only. It is possible that the conversation sentiment was already
degrading, and the last recommended action is simply following the
downward trend. On the other hand, given a worsening conversation
sentiment, a ‘positive’ sentiment for a recommended action implies
that it had a very positive impact on the user; and hence its
corresponding reward should be boosted. For example, let us consider a
ranking of the user sentiments:</p>
      <p>Disgusted ! Angry ! Sad ! Confused ! Calm ! Happy
Given this, a change from ‘Disgusted’ to ‘Happy’ would lead to
a much higher (positive) boost, than a (negative) change from
‘Confused’ to ‘Sad’.</p>
      <p>The boosted reward rbat for an action at at time t is computed as
follows:</p>
      <p>1
rbat = 2 (rat rbat 1 ) rat (6)</p>
      <p>It is easy to see that a ‘positive’ rat = 7 following a ‘negative’
rbat 1 = 5, will lead to rat getting boosted by a factor of 21 (7
( 5)) = 6. On the same lines, a ‘negative’ rat = 6 following a
‘positive’ rbat 1 = 4, will lead to rat getting further degraded by a
factor of 12 ( 6 4) = 5.</p>
      <p>We leave it as future work to extend the ‘boost’ function to last n
actions (instead of just the last action above). In this extended
scenario, the system maintains a sentiment curve of the last n actions,
and the deviation is computed with respect to a curve, instead of a
discrete value. The expected benefit here is that it should allow the
system to react better to user sentiment trends.
4</p>
    </sec>
    <sec id="sec-8">
      <title>IMPLEMENTATION</title>
      <p>
        In this section, we extended our RL powered Wellness App based
on Azure Personalizer, to accommodate the constructs outlined in
the previous section. Azure Personalizer [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is a Cloud based API
providing an implementation of RL Contextual Bandits [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In short,
Personalizer provides two primary APIs:
      </p>
      <p>Rank API: The mobile app invokes Rank API with a list of
actions and their features, and the user context and features. Given
this, the Rank API returns a list of ranked actions. Internally, the
Personalizer app uses the Explore-Exploit trade-off to rank the
actions:
– Exploit: Ranks the actions based on past data (current inference
model).
– Explore: Select a different action instead of the top action. The
‘explore’ percentage is a configurable parameter, and can be set
along the lines of an epsilon greedy strategy.</p>
      <p>Rewards API: The mobile app presents the ranked content
(returned by Rank API) and computes the reward corresponding to
each action. It then invokes the Reward API to return the
computed rewards to Personalizer. Personalizer correlates the
actionreward, updating its inference model.</p>
      <p>We now provide details of our Wellness Recommender App
(illustrated in Fig. 4). In addition to the usual article / activity
recommendations, the app provides the option to have a video chat as well
to improve the ‘interactive quotient’ of the app. The implementation
details of the RL constructs proposed in this paper are outlined
below:</p>
      <p>Multiple feedback channels: in this case correspond to the live
video feed and user interaction with an article / activity. Given
a user snapshot (captured from the live feed), the sentiment score
is computed using Azure Face API. The article / activity
recommended by the app depends on both the article / activity relevance
score and the ‘current’ user sentiment, e.g. tragic, sad, depressing,
etc. related articles / activities are not shown unless the user is in
a ‘happy’ mood (Fig. 4).</p>
      <p>Personalizer leaves the reward computation on the client side.
We developed a Rewards Computation module (with reference to
Eq. 3) that combines (i) the activity / article related score, i.e. the
activity / app selected and the time spent interacting with it,
together with (ii) the sentiment score computed based on the user
snapshot; with a higher weightage assigned to the latter given its
effectiveness in capturing the user reaction to a recommended
article / activity.</p>
      <p>Rewards boosting: To accommodate this, we consider a ranking
of the user sentiments returned by the Face API:
Disgusted ! Angry ! Sad ! Confused ! Calm ! Happy
The rewards boosting factor (Eq. 6) is assigned proportional to
the change in user sentiment before and after displaying the
recommended activity / article.</p>
      <p>Delayed rewards: is a core RL construct that requires updating
the backend Recommendation Engine and RL Reward and Policy
functions (Eq. 4 and Eq. 5). As such, it is difficult to implement
based on a Cloud API (without having direct access to the
underlying Recommendation and RL Engines). We are in the process of
adapting an Open Source Contextual Bandits implementation to
provide the full ‘delayed rewards’ strategy.</p>
      <p>For now, we only simulated the behavior of the RL Reward
function (Eq. 4) on the client (app) side. We added a memory buffer
to the Rewards Computation module to store rewards computed
for an iteration, until a ‘similar’ reward gets computed for a set
of activities / articles with similar features recommended as part
of another iteration. At this point, the aggregate rewards are
returned to Personalizer (via Reward API) for recommended
activities / articles of both iterations. This ensures that only consistent
rewards are considered while training the Personalizer RL
inference model.
5</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>In this work, we considered the implementation of a RL based
Recommender System, in the context of a real-life Wellness App. RL is a
powerful primitive for such problems as it allows the app to learn and
adapt to user preferences / sentiment in real-time. However, during
the case-study, we realized that current RL frameworks lack certain
constructs needed for them to be applied to such Recommender
Systems.</p>
      <p>To overcome this limitation, we introduced three RL constructs
that we implemented for our Wellness app: (i) weighted feedback
channels, (ii) delayed rewards, and (iii) rewards boosting. The
proposed RL constructs are fundamental in nature as they impact the
interplay between Reward and Policy functions; and we hope that their
addition to existing RL frameworks will lead to increased enterprise
adoption.</p>
    </sec>
    <sec id="sec-10">
      <title>ACKNOWLEDGEMENTS</title>
      <p>I would like to thank Louis Beck and Sami Ben Hassan for their
insights and support in developing the Wellness Recommender App.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Google</surname>
            <given-names>RecSim</given-names>
          </string-name>
          ,
          <year>2020</year>
          (accessed
          <issue>December 9</issue>
          ,
          <year>2020</year>
          ). https://opensource.google/projects/recsim.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Microsoft</given-names>
            <surname>Azure Personalizer</surname>
          </string-name>
          ,
          <year>2020</year>
          (accessed
          <issue>December 9</issue>
          ,
          <year>2020</year>
          ). https://azure.microsoft.com/en-us/services/cognitiveservices/personalizer/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Barto</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <source>Reinforcement Learning: An Introduction</source>
          , MIT Press, Cambridge, MA,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Biswas</surname>
          </string-name>
          , '
          <article-title>Privacy Preserving Chatbot Conversations'</article-title>
          ,
          <source>in Proceedings of the 3rd IEEE Conference on Artificial Intelligence and Knowledge Engineering (AIKE)</source>
          , (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Hwang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ha</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Yoon</surname>
          </string-name>
          .
          <source>Reinforcement Learning based Recommender System using Biclustering Technique</source>
          ,
          <year>2018</year>
          . arXiv:
          <year>1801</year>
          .05532.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Langford</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          , '
          <article-title>A Contextual-Bandit Approach to Personalized News Article Recommendation'</article-title>
          ,
          <source>in Proceedings of the 19th International Conference on World Wide Web (WWW)</source>
          , p.
          <fpage>661</fpage>
          -
          <lpage>670</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Guo</surname>
          </string-name>
          , and
          <string-name>
            <surname>Y. Zhang.</surname>
          </string-name>
          <article-title>Deep Reinforcement Learning based Recommendation with Explicit UserItem Interactions Modeling</article-title>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Nilsson</surname>
          </string-name>
          .
          <string-name>
            <surname>Delayed-Reinforcement</surname>
            <given-names>Learning</given-names>
          </string-name>
          ,
          <year>2020</year>
          (accessed
          <issue>December 9</issue>
          ,
          <year>2020</year>
          ). http://heim.ifi.uio.no/ mes/inf1400/COOL/REF/Standford/ch11.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ricciardelli</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Biswas</surname>
          </string-name>
          , '
          <article-title>Self-improving Chatbots based on Reinforcement Learning'</article-title>
          ,
          <source>in Proceedings of the 4th Multidisciplinary Conference on Reinforcement Learning and Decision Making (RLDM)</source>
          , (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Taghipour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kardan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Ghidary</surname>
          </string-name>
          , '
          <article-title>Usage-Based Web Recommendations: A Reinforcement Learning Approach'</article-title>
          ,
          <source>in Proceedings of the ACM Conference on Recommender Systems (RecSys)</source>
          , p.
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          , (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>