<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Discovering Similar Products in Fashion E-commerce</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Amber Madvariya</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Recommender Systems</institution>
          ,
          <addr-line>Item-Item Collaborative Filtering</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>In recent years, item-item collaborative filtering algorithms have been studied thoroughly in recommender systems. When applied in the context of fashion e-commerce these algorithms can be used to generate similar recommendations, personalize search results and to build a framework for creating clusters of similar products. The efficacy of these algorithms when applied in fashion domain, rely on accurately inferring a user's fashion taste and matching them to a product. Our work hinges around discovery of similar products using two different item-item collaborative filtering algorithms. We identify and address some unique challenges while applying these algorithms in the dynamic fashion e-commerce environment. We study and evaluate their performances through precision/recall measures, live A/B tests and baseline them against a content based approach. We also discuss effects of the transient nature of the industry on a user's fashion taste.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Fashion shopping is a complex intersection of product styles with
users’ taste where products are usually described in terms of
several attributes such as silhouette, line, hem length, color,
fabric, waist length and so forth. These attributes are dynamic
in nature and vary as trends in fashion change. However, these
attributes are not typical representatives of a user’s fashion taste.
Users rely more on look, feel, popularity and other intrinsic
factors to identify their taste.</p>
      <p>The advent and growth of fashion e-commerce, pose even more
challenges towards providing users with relevant products. As the
industry grows, there are thousands of products available in a
catalogue with similar set of attributes. Also users’ needs in fashion
are not specific as compared to hard goods like electronics,
therefore they tend to browse significantly more products on a
fashion portal before clicking on a particular product. In general,
at Myntra we observe the average click depth i.e. number of
products viewed before the first click, to be around 90. Correctly
inferring a user’s intent in a session becomes critical for providing
a better shopping experience. A user’s click on a prod-uct is a
proxy for his/her interest in that product. We can utilize this
information to narrow down his/her intent. Thereafter, serving a
user with similar products to the clicked one, helps them navigate
through the vast catalogue and find relevant products. We also use
similar products to personalize search results as majority of queries
on our platform are broad in nature like ’t-shirts’ and ranking
results based on user’s interest improves the overall search relevance.
Therefore, discovering similar products in fashion e-commerce
becomes an interesting data mining and business problem. In this
paper, we present our work on prescribing similar products using
item-item collaborative filtering algorithms in a fashion context.</p>
      <p>
        Similarity between items is usually computed in terms of content
based similarity, item-item based collaborative filtering [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] or the
hybrid of two. Content based similarity measures rely on a curated
taxonomy to represent a particular content. A widely cited example
of content based system is that of Pandora radio[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which uses
the Music Genome Project to tag their songs and artists from a set
of 450 manually curated attributes. Unlike music, the dynamicity of
trends in fashion, makes its taxonomy ephemeral and new products
continuously require new attributes to be added to the taxonomy.
Maintaining and curating this taxonomy system becomes an
unscalable task over time. Another limitation with this approach is
that attributes are unable to capture user taste completely.
      </p>
      <p>
        Unlike content based similarity measures, item-item
collaborative filtering algorithms use user signals instead of a taxonomy to
compute item-item similarity. In the domain of e-commerce users’
feedback is not explicit [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and needs to be inferred from user
signals. In traditional e-commerce platforms [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], user purchases are
used to compute item similarity. The challenge in using only user
purchase data to compute item similarity in fashion e-commerce
is that it is sparse and erratic. This is due to two primary factors,
the short life cycle of products and the breadth of the catalogue
available. Thus similarity measures relying solely on user purchase
data tend to perform poorly in this setting. So we consider
purchase data along with signals like clicking on a product and adding
a product to a cart to compute item similarity.
      </p>
      <p>The challenges with using these implicit feedback signals are
following:
• Quantifying the chosen signals by assigning relative weights
to each.
• Normalising efect of popular products i.e. products which
have signals from a large number of users.</p>
      <p>In this work, we look at two representations to model item-item
collaborative filtering. We compare these two approaches with a
content based approach, where we use the annotated taxonomy of
a product to get a feature representation for it.</p>
      <p>In the following sections of this work, we describe these three
approaches in detail and compare their performance via A/B tests
and precision-recall test. Finally, we try to conclude by inferring
whether the transient nature of the industry afects user taste or
not.
2</p>
    </sec>
    <sec id="sec-2">
      <title>METHODOLOGY</title>
      <p>For the remaining sections of this paper, we will use the following
terminology. Let P be the set of all products and S be the set of
all user sessions. A session contains all activity by a user within
a 30 minute window from the time he/she logs into the portal
annotated by time-stamp. We record user signals like product list
views, product clicks, addition to carts, orders placed and so on.</p>
      <p>For all the three approaches mentioned below, we split our data
at the article type (e.g Men-Tshirts, Men-Shirts, Women-Dresses)
level, since products similar to a product pi should be from the
same article type. Splitting the data at the article type level has the
added advantage that it reduces the set of products from which we
ifnd similar products for pi .
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Product Attribute Vectors</title>
      <p>
        Each product in our system is annotated with attributes generated
from a manually curated taxonomy. These attributes are broadly
classified into two types, general attributes like brand, color, price
bands etc which are applicable for all article-types and article-type
specific attributes (ATSA) like collar type, sleeve length, neck type
for T-shirts. Each attribute is represented as a key-value pair e.g.
collar type of a t-shirt is the key and diferent collar types like
roundneck, v-neck, polo-neck are values. Each product is represented
as a vector in the real space, where each dimension represents an
attribute. We use binary values to populate a particular dimension
i.e. if the attribute represented by the dimension is a part of the
set of attributes which were annotated to the product, then we
populate the dimension as 1, else we populate it as 0. Along with
product attributes, we extract relevant bi-grams using log likelihood
ratio scores[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] from the product descriptions. We utilize a L2 norm
to normalize these product vectors.In order to generate N similar
products for a product pi , we discover the N closest neighbors of
pi , keeping cosine distance as our distance metric.
      </p>
      <p>We analyse this approach here to benchmark the performance of
item-item collaborative filtering algorithms. The major challenge
with this approach is that the set of attributes representing a
product is not comprehensive and could results in products not being
annotated by some key information.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Item-Item Weighted Graph</title>
      <p>
        In this approach, we use an undirected weighted graph
representation to model product relationships in the system. In this
representation nodes represent products, edges represent associativity
between products and edge weights represent degree of
associativity between products. This approach was first introduced by
YouTube [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], however our formulation to calculate edge weight is
diferent than the one showed in that work.
      </p>
      <sec id="sec-4-1">
        <title>Input Image</title>
        <p>(0.081)
(0.078)
(0.071)
(0.059)
(0.056)</p>
        <p>
          In order to to generate this graph, we use a well-known technique
known as association rule mining[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] or co-visitation counts. We
consider a user session si , where si ⊂ S, and generate a set of
products {p1, p2, . . . pn } clicked in that session. Within this set, we
compute all pair combinations of products (pi ,pj ) which were
cobrowsed together. We count all the occurrences where pi and pj
were co-browsed together, across S and denote the total count of
co-occurrences of (pi ,pj ) as ci j . We assign edge weights as wi j and
calculate it using the following formulation of normalised
pointwise mutual information (NPMI) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]:
        </p>
        <p>wi j = log−p1(i,j) ∗ log pp(i()ip,j()j)
where p(i,j) is the probability of occurrence of the pair (pi ,pj ) and
p(i) and p(j) is the probability of occurrence of pi and pj respectively.
We can use the following formulations for the values of p(i,j), p(i),
p(j):
p(i,j) = ci j p(i) = ci</p>
        <p>C C
where, ci j = count of occurrences of pair (i,j),
p(j) =
ci =</p>
        <p>X cik
k ⊂P
and</p>
        <p>C = X
i, j ⊂P</p>
        <p>Then, to find similar products for a product pi , we locate pi in
the graph, we consider it’s adjacent nodes and among them pick
the top N neighbors, sorted in descending order by edge weight.
Here N is a hyper-parameter which denotes the number of similar
products we want to find for pi .</p>
        <p>The advantage of using this formulation to generate edge weights,
is that it normalises the efect of popular products. If we consider
a pair of products (pi ,pj ) where pj is browsed more across all
sessions as compared to pi , then the value of cj will be high, resulting
in a high value of p(j). The high value of p(j) results in the PMI
of (pi ,pj ) turning out to be low, even though the co-occurrence
counts of (pi ,pj ) might have been higher as compared to other pairs
containing pi .</p>
        <p>One challenge with using this approach is the noise prevalent in
the input data. Some pairs might have a low co-browsing count or
low PMI score to form a meaningful edge. To tackle this challenge,
we put a threshold on both the co-browsing counts and the edge
weights generated using PMI scores. After experimenting, we found
5 to be a suitable threshold for co-browsing count and 0.15 to
be a suitable threshold for PMI score. These thresholds change
depending upon the number of products and user sessions.
cj
C
ci j</p>
        <p>Another source of noise is the session in consideration itself.
This entire approach hinges around the assumption that products
browsed in a session are similar i.e. there is some context to products
browsed by a user in a session. If the number of products browsed in
a session are low, say 2 or 3, or if diferent article types were browsed
randomly in a session, without any intent, then these sessions will
add noise to our system. To tackle this problem, we use the concept
of coherent sessions. We define a session as coherent, if in that
session, the user has browsed at least three products of the same
article type, and not more than two article types.</p>
        <p>As we have considered sessions to construct the graph, it results
in edges being formed only between products which were present
in the system at the same time. Therefore, this graph captures the
ephemeral nature of fashion trends, because products present across
two diferent points in time would never form an edge. So, while
ifnding similar products from this graph for pi , the output products
represent a trend from the time when pi was present.
2.3</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Product-User Vectors</title>
      <p>In this approach, we represent each product as a vector of user
signals in the real space, with each dimension representing a user. We
consider three user signals i.e. product clicks, addition of products
to cart and checkout for populating values in each dimension. We
start by aggregating these user signals across S, for a product-user
combination (pi ,ui ), and generate a chronologically sorted list of
all the identiefid signals that user ui has generated for product pi ,
where pi ∈ P . If we denote a product click by C, addition to cart as
T and checkout as O, one example of such representation could be
(pi , ui ) = {C, C, C, T , O }</p>
      <p>After generating this list, we quantify our user signals, based on
past data. We estimate the relationship between number of product
clicks, add to carts and checkout using a linear classification model
with each session as a data point. In the linear classification model,
we use product clicks and add to cart signals as input and the
checkout event as the target value.</p>
      <p>The weight for checkout event is considered as 1, since it is the
target value in the classification model. So, for the products which
a user has bought, we populate a value of 1 in that user’s dimension
for those products. For products, which the user hasn’t bought but
has clicked or added to cart, we populate it’s vector with weights
that we obtained from the model. For users, who have not generated
any of the above signals for a product, we don’t populate any value
in their dimensions. This results in the vector for each product
being sparse, since the subset of users who have generated signals
for a product is very small as compared to all users.</p>
      <p>After computing the vector for a product pi , we take it’s L2 norm
and find it’s N similar products by finding it’s N nearest neighbors
using the following formulation of cosine similarity as the distance
metric.</p>
      <p>cos (i, j ) =</p>
      <p>p⃗i · p⃗j
|pi | ∗ |pj |
where p⃗i and |pi | represent the vector of pi and it’s magnitude,
"·" represents the dot product between the two vectors.</p>
      <p>
        In order to get accurate cosine similarities, we utilise a
compressed sparse row (CSR) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] matrix representation of the vectors,
instead of approximate nearest neighbor algorithms[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Each row
of this matrix represents a product and each column represents a
user. We multiply this matrix by its transpose, such that the (i,j)th
entry of the resultant matrix gives us the cosine similarity between
pi and pj . Another advantage of using this representation is that
it reduces computation time because of inherent sparse matrix
optimisations.
      </p>
      <p>We observe that products which are globally popular are dense
as compared to products which are less popular. This results in
popular products being close neighbors to many products, if their
vectors are not normalised. Taking a L2 norm reduces the magnitude
along each dimension of popular product and normalizes for global
popularity.</p>
      <p>In this approach, since we have aggregated user signals across
S, it results in the system being agnostic of fashion trends. Unlike
the item graph approach, we find multiple instances of linkage of
products across time. Hence, while finding similar products for a
product pi using this approach, we find that the output products do
not necessarily reflect trends from the same time-frame as when pi
was present in the system.
3
3.1</p>
    </sec>
    <sec id="sec-6">
      <title>Data</title>
    </sec>
    <sec id="sec-7">
      <title>ANALYSIS AND RESULTS</title>
      <p>We split our sessions data into two non overlapping sets, a training
set, which is used to generate the similar products, and a test set,
which is used to evaluate the approaches. On an average, each
session contains 5 products. The training set consists of ∼200M
sessions with ∼12M unique users. Using this training set, we find
similar products for ∼1.3M products. The test set is generated from
∼10M sessions with ∼500K unique users. To generate the test set
of products for pi , we take each session in the test data, where</p>
      <sec id="sec-7-1">
        <title>Input Image</title>
        <p>(0.712)
(0.702)
(0.588)
(0.473)
(0.405)
pi was clicked, and we consider the product which was clicked
immediately after pi in this session. Aggregating across all the
sessions in the test data, we create a set of products which were
clicked immediately after pi . Using the session test data, we were
able to generate test sets for ∼330K products, with the average size
of the set being 164.</p>
        <p>
          We use a series of MapReduce computations [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] to generate the
item graph and to aggregate user signals for each product user
combination. Further, to create the CSR matrix for product user
vectors, we use the CSR sparse matrix routine in scipy [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
3.2
        </p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Analysis</title>
      <p>3.2.1 Browsing behaviour by Gender. To analyse browsing
behaviour by gender, we look at density statistics from both the
item-item graph and product-user vectors.</p>
      <p>For item-item graph, we define, the density of a graph as
e
d =</p>
      <p>n ∗ (n − 1)/2
where e is the number of edges and n is the number of nodes.
Density here represents the ratio of number of edges formed in a graph
to all possible edges in the graph. Below we tabulate density
statistics of graphs for some prominent article types</p>
      <sec id="sec-8-1">
        <title>Article Type</title>
        <sec id="sec-8-1-1">
          <title>Women-Tops</title>
          <p>Men-Tshirts
Women-Jeans
Men-Jeans
Women-Casual Shoes
Men-Casual Shoes
Women-Sports Shoes
Men-Sports Shoes</p>
          <p>As can be observed from the graph, for two comparable
article types like Men-Jeans and Women-Jeans, the women article
types have higher density as compared to men article types, despite
women article types having lower number of nodes. It can be
inferred from the higher density of women article types that women
browse more products than men before making a selection.</p>
          <p>For product user vectors, density of a sparse vector is defined as
the ratio of number of non-zero values in a vector to it’s length. In
the table below, we tabulate density statistics for vectors of diferent
article types.</p>
        </sec>
      </sec>
      <sec id="sec-8-2">
        <title>Article Type</title>
        <sec id="sec-8-2-1">
          <title>Women-Tops</title>
          <p>Men-Tshirts
Women-Jeans
Men-Jeans
Women-Casual Shoes
Men-Casual Shoes
Women-Sports Shoes
Men-Sports Shoes</p>
          <p>Similar to item-item graph, we notice that women article types
are more dense as compared to men products. The higher density
of women product vectors corroborates our earlier inference that
women products are browsed more as compared to men products.</p>
          <p>3.2.2 Diference in Style Catalogued Date . Each product in
our system is tagged with a style catalogued date (SCD), which
represents the date when the product was introduced to the system.
We calculate the average diference between SCD of a product and
it’s similar products and further average it over an article type.</p>
          <p>Fig 8 represents average diference in SCD, between input
products and their similar products averaged over article types for the
two approaches. It can be observed from the plot that average
difference in SCD for product-user vectors is 2-3 times more than that
of item-item graph. Hence, it can be inferred from the plot that
similar products found using item-item graph are from the same
time-frame, while similar products found using product user
vectors link across time. This validates our claim that similar products
found using item-item graph are cohesive of fashion trends, while
those found using product user vectors are agnostic of them.
3.3</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Results</title>
      <p>3.3.1 Precision Scores. Denoting the test set of pi by Ti and
the set of similar products for it as Ri , we use the following
formulation to calculate precision:
precision = |Ri ∩ Ti |
|Ri |</p>
      <p>The size of Ri is a hyper-parameter, since we can set the number
of similar products that we want to find for pi , and we denote it as
N.</p>
      <sec id="sec-9-1">
        <title>Input Image</title>
        <p>(0.151)
(0.112)
(0.108)
(0.107)
(0.100)
Fig 4: Precision versus N for the three approaches
Fig 5: Recall versus N for the three approaches
Fig 6: Precision for diferent article types
Fig 7: Recall for diferent article types</p>
        <p>Fig 4 shows the average precision of the three approaches for
diferent values of N. Fig 6 shows precision by article type for the
three approaches calculated with a value of N as 25.</p>
        <p>3.3.2 Recall Scores. Going with the same notations as above,
we use the following formulation to calculate recall:
recall = |Ri ∩ Ti |</p>
        <p>|Ti |</p>
        <p>Fig 5 shows the average recall of the three approaches for
different values of N. Fig 7 shows recall by article type for the three
approaches, with the same value of N as for Fig 6 i.e. 25.</p>
        <p>It is evident from Fig 4-7 that the two collaborative filtering
approaches significantly out perform the content based approach.
Among the two item-item collaborative filtering approaches we
calculate from the numbers of Fig 4 &amp; 5 that product user vectors
show an average improvement of 7% and 5.1% over item-item graph
for precision and recall, respectively. Also we observe that product
user vectors have better precision recall numbers across majority
of article types as compared to item-item graph.</p>
        <p>
          3.3.3 A/B Tests. For the A/B test, we render the similar
products on the product details page of the input product. We test the
above three approaches by assigning randomly selected 10% of
trafifc to each treatment ( 100k users for each treatment). We recorded
the number of product views and the number of clicks for the three
approaches over a period of three weeks. Figure 9 shows the CTR
(click through rate defined as ratio of clicks to views) recorded
during the test period for the three approaches averaged over each day
of the week. From the test, we observe that product user vectors
show an average CTR improvement of 5% over item-item graph
with a p-value[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] of 0.019. Also, we notice that item-item graph and
product user vectors show an improvement of 50% and 58.4% over
product attribute vectors with a p-value of 1.2*10−5 and 6.6*10−6
respectively.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSION</title>
      <p>In this paper, we have identified and addressed some of the
challenges faced by item-item collaborative filtering approaches which
are unique to fashion e-commerce domain. We formulate a new
method to combine and quantify various user signals, which can
be used as input to these approaches, in e-commerce domain.</p>
      <p>We propose a new method to evaluate similar products in an
ofline setting. We compare and evaluate three approaches to
discovering similar products in both ofline and online tests. We observe
a significant improvement in the performance of item-item
collaborative filtering approaches as compared to the content based
approach. Furthermore, among the two item-item collaborative
ifltering approaches product user vectors perform noticeably better
than item-item graph. We also compare the efects of fashion trends
with respect to the two approaches and based on results of the tests,
infer that a user’s preference to products is independent of rapidly
changing fashion trends.
5</p>
    </sec>
    <sec id="sec-11">
      <title>ACKNOWLEDGEMENTS</title>
      <p>We thank Deepak Warrier, Ankul Batra, Sagar Arora, Ullas
Nambiar, Ashay Tamhane and Kunal Sachdeva for their contributions
Fig 8: Average diference in SCD by article type
Fig 9: Average CTR over a period of 3 weeks
in reviewing this work and their inputs to algorithm design and
evaluation.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Rakesh</given-names>
            <surname>Agrawal</surname>
          </string-name>
          , Tomasz Imieliński, and
          <string-name>
            <given-names>Arun</given-names>
            <surname>Swami</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Mining association rules between sets of items in large databases</article-title>
          .
          <source>In Acm sigmod record</source>
          , Vol.
          <volume>22</volume>
          . ACM,
          <volume>207</volume>
          -
          <fpage>216</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Gerlof</given-names>
            <surname>Bouma</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Normalized (pointwise) mutual information in collocation extraction</article-title>
          .
          <source>Proceedings of GSCL</source>
          (
          <year>2009</year>
          ),
          <fpage>31</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Aydin</given-names>
            <surname>Buluç</surname>
          </string-name>
          , Jeremy T Fineman, Matteo Frigo, John R Gilbert, and
          <string-name>
            <given-names>Charles E</given-names>
            <surname>Leiserson</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Parallel sparse matrix-vector and matrix-transpose-vector multiplication using compressed sparse blocks</article-title>
          .
          <source>In Proceedings of the twenty-first annual symposium on Parallelism in algorithms and architectures. ACM</source>
          ,
          <volume>233</volume>
          -
          <fpage>244</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>James</given-names>
            <surname>Davidson</surname>
          </string-name>
          , Benjamin Liebald, Junning Liu, Palash Nandy, Taylor Van Vleet,
          <string-name>
            <surname>Ullas Gargi</surname>
          </string-name>
          , Sujoy Gupta,
          <string-name>
            <surname>Yu</surname>
            <given-names>He</given-names>
          </string-name>
          , Mike Lambert,
          <string-name>
            <given-names>Blake</given-names>
            <surname>Livingston</surname>
          </string-name>
          , et al.
          <year>2010</year>
          .
          <article-title>The YouTube video recommendation system</article-title>
          .
          <source>In Proceedings of the fourth ACM conference on Recommender systems. ACM</source>
          ,
          <volume>293</volume>
          -
          <fpage>296</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          and
          <string-name>
            <given-names>Sanjay</given-names>
            <surname>Ghemawat</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>MapReduce: simplified data processing on large clusters</article-title>
          .
          <source>Commun. ACM 51</source>
          ,
          <issue>1</issue>
          (
          <year>2008</year>
          ),
          <fpage>107</fpage>
          -
          <lpage>113</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Ted</given-names>
            <surname>Dunning</surname>
          </string-name>
          .
          <year>1993</year>
          .
          <article-title>Accurate methods for the statistics of surprise and coincidence</article-title>
          .
          <source>Computational linguistics 19</source>
          ,
          <issue>1</issue>
          (
          <year>1993</year>
          ),
          <fpage>61</fpage>
          -
          <lpage>74</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Tv</given-names>
            <surname>Genius</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>An Integrated Approach to TV &amp; VOD Recommendations</article-title>
          .
          <article-title>(</article-title>
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Eric</given-names>
            <surname>Jones</surname>
          </string-name>
          , Travis Oliphant,
          <string-name>
            <given-names>Pearu</given-names>
            <surname>Peterson</surname>
          </string-name>
          , et al.
          <fpage>2001</fpage>
          -. SciPy:
          <article-title>Open source scientific tools for Python. (2001-)</article-title>
          . http://www.scipy.org/ [Online; accessed 2016-
          <volume>10</volume>
          -14].
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Greg</given-names>
            <surname>Linden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Brent</given-names>
            <surname>Smith</surname>
          </string-name>
          , and Jeremy York.
          <year>2003</year>
          .
          <article-title>Amazon. com recommendations: Item-to-item collaborative filtering</article-title>
          .
          <source>IEEE Internet computing 7</source>
          ,
          <issue>1</issue>
          (
          <year>2003</year>
          ),
          <fpage>76</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Marius</given-names>
            <surname>Muja</surname>
          </string-name>
          and David G Lowe.
          <year>2009</year>
          .
          <article-title>Fast Approximate Nearest Neighbors with Automatic Algorithm Configuration</article-title>
          .
          <source>VISAPP (1) 2</source>
          ,
          <fpage>331</fpage>
          -
          <lpage>340</lpage>
          (
          <year>2009</year>
          ),
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Pandora</given-names>
            <surname>Radio</surname>
          </string-name>
          . [n. d.].
          <source>Music Genome Project. ([n. d.])</source>
          . https://www.pandora. com/about/mgp
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Badrul</surname>
            <given-names>Sarwar</given-names>
          </string-name>
          , George Karypis, Joseph Konstan,
          <string-name>
            <given-names>and John</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Item-based collaborative filtering recommendation algorithms</article-title>
          .
          <source>In Proceedings of the 10th international conference on World Wide Web. ACM</source>
          ,
          <volume>285</volume>
          -
          <fpage>295</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Bernard</surname>
            <given-names>L</given-names>
          </string-name>
          <string-name>
            <surname>Welch</surname>
          </string-name>
          .
          <year>1947</year>
          .
          <article-title>The generalization ofstudent's' problem when several diferent population variances are involved</article-title>
          .
          <source>Biometrika</source>
          <volume>34</volume>
          ,
          <issue>1</issue>
          /2 (
          <year>1947</year>
          ),
          <fpage>28</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>