<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Retrieval-Information Filtering.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Collaborative Filtering</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Poster</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ben-Gurion University of the Negev</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Bundle Recommendation</institution>
          ,
          <addr-line>Recommender Systems, E-Commerce</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>Recommender systems (RSs) enhance e-commerce sales by recommending relevant products to their customers. RSs aim at implementing the firm's web-based marketing strategy to increase revenues. Generating bundles is an example of a marketing strategy that aims to satisfy consumer needs and preferences, and at the same time, to increase customers' buying scope and the firm's income. Thus, finding and recommending an optimal and personal bundle becomes very important. In this paper we introduce a novel model of bundle recommendations that integrates collaborative filtering (CF) techniques, personalized demand functions, and price modeling. This model provides a recommendation list by finding pairs of products that maximizes both, the probability of their purchase by the user and the revenue received by selling this bundles. Bundling refers to the practice of selling two or more items together as a package at a price that is below the sum of the independent prices. Optimal bundling would combine items into bundles that best fit the retailer's needs and the user's preferences, and maximize product compliance within the bundle. Thus, a single price &lt;  +   is set for the two products (A, B) if purchased</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>jointly. One challenge is to suggest a price for a bundle that fits both
the customer reservation price i.e., the maximal price buyers are
accepted to pay, and the retailer’s revenue [1]. Very few studies
have combined bundling strategy
with recommender systems
(RSs). The field of frequent item set mining and association rules
deals with finding a basket of items that are frequently bought
together [2]. However, these techniques are not personalized, thus
not applicable for RSs. The recommendation of
bundles were
presented as a tailored solution for the tourism domain using
casebased reasoning where case models representing the travel plan
bundle were matched against the user profile and preferences [3].
The authors of [4] presented a bundle optimization using a genetic
algorithm to maximize the compatibility of the products within a
bundle.</p>
      <p>However,
these
studies
did
not
measure
the
recommendation aspect, i.e., if it is at all feasible and beneficial to
predict bundle purchasing. The study presented in [5] introduces a
held
by
the
1+ ,</p>
      <p>
        1
1 +
1

,
bundle recommendation problem, in which its solution is a set of
items that maximizes some total expected reward. However, the
price aspect was not considered in the model. Our paper maximizes
the expected revenue by considering the item-to-item
cross
dependencies, user-item collaborative filtering techniques and the
demand-price function—resulting in recommendation of the best
bundle and price proposal to the user.
2. BUNDLE RECOMMENDATION MODEL
We maximize the following retailer expected revenue function:
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )  = 
where   (, ,
      </p>
      <p>) is the probability that user i will purchase the
bundle, which is composed of products A and  , at price T. The</p>
      <p>A is the retailer’s cost for product A and 
 is the retailer’s
cost for product B. The proposed bundle and the price T for user i
is set to maximize the expected revenue:
 (, , 
) ∙ ( − 
 − 
 )
(2) (, , 
) = 
∀,,
(, , )
In order to find   (, ,</p>
      <p>
        ) we find the corresponding prices   of
product  and   of product  aggregated to the bundle price  :
(
        <xref ref-type="bibr" rid="ref2">3</xref>
        )   (, , 
) = 
∀  ,  |  +  =   ( ∩  ∩ 
 ∩   )
Thus, we have to find the prices  
and   that maximize the
probability of the user i to buy products  and  while paying those
prices. According to Bayes' law:
(4)   (  ∩ 
  (
 
 
   
   
) =
| 
  ( 
)
According to the Jaccard measure: (5)  = 
) ∙
=
 (∩ )
 (∪ )
Using
      </p>
      <p>combinatorial
principle: (6)  ( ∩ 
mathematics, the</p>
      <p>inclusion–exclusion
) =  ( ) +  ( ) −  ( ∪ 
Using equations (5) + (6): (7)  ( ∩ 
) =
Using Bayes’ law and equation (7):</p>
      <p>(8)   (( ∩   ) ∩ ( ∩   )) =   ( ) ∙   (  | ) +   () ∙   (  |)
  ()
2.1
We assume that the Jaccard measure,  , , which denotes the
products' compatibility, is not affected by the price. The   ( ),
probabilities
are
found
using
the</p>
      <p>CF
technique;
  (  | ),   (  |) is found by the upcoming personal demand.</p>
      <p>Personalized Demand Graph
We would like to assume that each customer has its own demand
graph for each product based on his/her preferences. Thus, we
developed heuristics for estimating the “personalized” demand
graph for user i and item j using very sparse data. Figure 1
demonstrates the demand of a generic customer versus an
enthusiastic one (i.e., one that would pay high prices) as well as
an indifferent one. We assume that the difference between the
demand graphs can be reflected by the following:
(9)   (  | ) = 
( ∙, (  ) ×  , , 100%)
where  ∙, ( A) is the generic demand graph for item  given that
the price  A and  , is the personalized bias factor for user i and
item  . In order to find the personal bias factor,  , , we scan each
customer's previous purchases or his/her highest bid on an item. We
compare his/her price to the median of the generic graph. For
example if customer i purchased item  for price  A* then his/her
bias factor is estimated as: (10)  , =  ∙, 0(.5A∗).
For example (Figure 1), assume that a customer purchased the item
for  A*=1300; according to the generic graph, this price would be
considered only by 35% of the interested population. Thus the
personalized bias for this user is calculated as:  , = 00.3.55 = 1.42.
We can create a bias matrix for all purchases of items by users. The
bias factor of products that have not been purchased by the
customer can be predicted using the SVD method. Given the
complete matrix, we can infer the personalized demand graph of
each user i and item  from the generic demand graph calculated
for item  and multiply it by the predicted alpha.</p>
    </sec>
    <sec id="sec-2">
      <title>3. EVALUATION</title>
      <p>
        Our model was evaluated based on two datasets. (Dataset 1)
consists of transactions from a shopping website that sells
electronics and furniture. (Dataset 2) is a supermarket dataset from
Kaggle
(https://www.kaggle.com/c/acquire-valued-shopperschallenge). We used offline evaluation and compared our model to
SVD and CF as baseline models. We evaluated: (i) The personal
demand function by using a validation set in order to test the
predicted alphas compared to the actual alphas, using the RMSE
measure, and comparing the personal demand graph probability to
0.5 (median probability) of all purchased products in the test
setusing the RMSE measure too; (ii) The product bundling
recommendation by comparing the top 5 bundles to the top 5 items
recommended by CF and SVD algorithms. For this we used
precision, recall, the average quantity that was recommended and
purchased, and the average price paid for the recommended and
purchased products; (iii) The price bundling recommendation by
comparing the recommended price to the actual price the user paid
in the test set, measuring the sum of the absolute difference. The
recommended price was compared to the mean price of the product.
We also compared two strategies: (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) maximizing the bundle
buying probability of the user, and (2) maximizing the expected
revenue. For both datasets we evaluated our model on the top 1,000
customers and top 300 products. The first dataset resulted in 3,425
transactions and the second in 836,846 transactions. A
recommended bundle is considered a hit in the test set if the two
products have been purchased by the user within a week. In dataset
1 for the personal demand graph we received an RMSE of alpha of
0.072 and an RMSE error compared to the median of 0.261. Thus,
the personal graphs are compatible to the users’ preferences. In
dataset 2, for the personal demand graph, we received an RMSE
error of alpha of 1.067 and an RMSE of 0.34, compared to the
median. The results for the product evaluation are presented in table
1 and 3 and the results for the price evaluation are presented in table
2 and 4. Bundle (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) and Bundle (2) represent the two strategies of
maximizing probability and the expected revenue, respectively.
4. CONCLUSIONS AND FUTURE WORK
Our results demonstrate that bundles are predictable and may
increase users' purchase scope. The first dataset is more difficult to
predict, but the bundle model is at least comparable to state of the
art algorithms and is even superior in some cases. The personal
demand graph tends to be very accurate as was observed by the
price recommendation accuracy. The second dataset contains
commodities data, thus a personal demand graph is more difficult
to predict. The recommended price was not as accurate as in the
first dataset. Moreover, for dataset 2 the products are more
predictable and the first bundle strategy yields the best results. For
both datasets maximizing the probability of the user's purchase is
more effective than maximizing the expected revenue. Future work
will aim at improving the personal demand graph of dataset 2,
examining more datasets and providing live user experiments.
[2] Agrawal, R., Imielinski, T., and Swami, A.N. 1993. Mining
association rules between sets of items in large databases.
      </p>
      <p>ACM SIGMOD, volume 22,2 of SIGMOD Record, 207–216.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Guiltinan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>1987</year>
          .
          <article-title>The price bundling of services: a normative framework</article-title>
          .
          <source>The Journal of Marketing</source>
          ,
          <fpage>74</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Ricci</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <year>2002</year>
          .
          <article-title>ITR: a case-based travel advisory system</article-title>
          .
          <source>Advances in Case-Based Reasoning</source>
          . Springer Berlin Heidelberg,
          <fpage>613</fpage>
          -
          <lpage>627</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>