<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Finding Similar Products in E-commerce Sites Based on Attributes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Urique Ho mann</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Altigran da Silva</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Moises Carvalho</string-name>
          <email>moisesg@icomp.ufam.edu.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instituto de Computaca~o Universidade Federal do Amazonas Manaus</institution>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present a preliminary study on the problem of nding products similar to a product given as input, based solely on their attributes. We assume that we are given a set of products from a same category of a same on-line store, were each product is described in a catalog by a number of attributes (e.g., general characteristics, technical speci cations, etc.). This problem, which at a rst glance may be seen as straightforward or even mundane, is in fact challenging and intriguing. In fact, any automatic solution for it requires techniques for comparing tens of di erent atributes, whose semantics are often very technical and speci c (e.g., the shutter speed of a camera) and also requires dealing with hundreds of products in the category. To be generic, such a solution must also deal with several distinct product categories. In here, we describe and evaluate a similarity function we have proposed for comparing products based on their attributes. This function uses a number of attribute-speci c similarity functions, which are selected according to a class assigned to the attribute. The assignment of classes to attributes is carried out by a simple classi cation strategy, which we also describe and evaluate. Experiments we carried out to evaluate our proposed similarity function using data from real catalogs in ve distinct popular product categories have shown promising results.</p>
      </abstract>
      <kwd-group>
        <kwd>Similarity Functions</kwd>
        <kwd>E-Commerce</kwd>
        <kwd>Recommender Systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Recommendation Systems are used by most e-commerce sites to suggest
products to their users and provide additional information to help customers to decide
which products to acquire [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Products can be recommended based on several
di erent types of information such as top overall sellers on a site, customer's
demographics, customer's past buying behaviour, or product attributes, e.g.,
technical speci cations, general characteristics, brand, etc. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Recommendations
based on this last type of information are called content-based or
knowledgebased recommendations [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>A simple way of enabling content-based recommendation is, given a product,
presenting to the user other products that are similar to it with respect to their
attributes. This is useful, for instance, when costumers explicitly want to nd
products with certain characteristics or when the seller wants to present to a
customer products similar to a product she is interest in, e.g., for the sake of
comparison or to provide alternatives to out-of-stock itens.</p>
      <p>However, in typical e-commerce sites, looking for similar products may
require the user to browse manually through a large number of pages and
products. For instance, suppose a user is interested in a speci c camera, say, \Nikon
S3500". Currently, if this user wants to nd alternative cameras that are
similar to this model (i.e., having similar features), for the sake of comparing their
prices, it is likely that she would have to browse over hundreds of other cameras
in the catalog to nd them. On the other hand, if this camera is not in stock,
it would be interesting to provide the user with similar alternative cameras in
stock, without having her to look over the whole catalog.</p>
      <p>Another interesting aspect of this kind of recommendation is that it enables
suggesting products to the customers without relying on historical data. It means
that the system can recommend products and provide buying options even if a
costumer is new to the system or if the item is new to the catalog.</p>
      <p>To nd whether two products are similar it is necessary to compare them.
Products on e-commerce sites are often described by their attributes. It means
that, to make a comparison between two products, it is necessary to compare
their attributes. This can be unfeasible to be carried out manually by casual
users on the Web.</p>
      <p>For instance, in a certain e-commerce site, to verify whether the \Nikon
S3500" camera is similar to another camera, say the \Sony W830", a user has
the option of comparing the 26 atributes provided for the rst camera with the
corresponding attributes of the second cameras. The lists of attributes available
for each camera in this site are presented in Figure 1. Notice that the second
camera has only 18 attributes. Also, notice that many attributes are di cult to
be compared, unless the user is an expert in the eld.</p>
      <p>In general, the same situation occurs in many other categories, that is,
comparing products requires comparing tens of attributes, some of them with very
speci c semantics.</p>
      <p>In this paper we present a preliminary study on the problem of nding
products similar to a given product. We assume that we are given a set of products
from a same category of a same on-line store, along with their attributes. For
instance, one of the datasets used in our experiments comprises a set of 489
camera models under the Cameras category of a real on-line store.</p>
      <p>Speci cally, we describe and evaluate a generic similarity function we have
proposed for comparing products based on their attributes. This function uses a
number of attribute-speci c similarity functions, which are selected according to
a class assigned to the attribute. The assignment of classes to attributes is carried
out by a simple but e ective classi cation strategy, which we also describe and
evaluate here.</p>
      <p>An experimental evaluation we carried out and reported here has shown
promising results. Our proposed similarity function showed to be accurate in
nding similar products, achieving average F-1 values above 0.75 in 5
representative product categories we have tested. Also, our strategy for attribute
classication has correctly classi ed most of the attributes from these categories.</p>
      <p>Attribute Nikon S3500</p>
      <p>Brand Nikon
Type of Camera Compact
Monitor/Display 2,7" LCD / TFT 230.000</p>
      <p>Resolution 20,1
Internal Memory 25MB</p>
      <p>Memory Cards Yes
Compatible Memory Cards SD, SDHC and SDXC</p>
      <p>Sensor
Optical Zoom
Digital Zoom</p>
      <p>Lenses
Shutter Speed
Focus range</p>
      <p>Opening
Flash Modes
Flash range
Battery Type
Video Features
Scene modes</p>
      <p>File Formats
Built-in microphone</p>
      <p>Tripod mount
Menu Languages</p>
      <p>Color
Dimensions (HxWxD)</p>
      <p>Weight</p>
      <p>Sony W830
Although important and challenging, e ective methods for nding similar
products are scarce both in industry and in the academy.</p>
      <p>
        Kagie et. al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] proposed a content-based graphical shopping interface based
on product attributes to recommend similar products. To use this interface,
the user must rst de ne an ideal product by providing desired values to its
attributes. The interface then shows products considered as similar to this ideal
product in a 2D Map. By interacting with this map, the user chooses, from the
products plotted, the most similar to the ideal. The interface then recalculates
the similarity between the ideal product to all other products in the dataset.
This process continues until the interface shows a product the user considers as
the most similar. In this work the authors consider only two of attribute classes:
categorical and numeric.
      </p>
      <p>Our approach di ers from this in many aspects. First, in our approach the
user does not need to specify an ideal product. In fact, this is avoided, since we
consider that casual users in e-commerce sites are not willing to specify desired
values for several attributes. Instead, we only require the user to select one
product to be used for comparison. Second, besides categorical and numerical
attributes, we consider two additional classes of atributes: multi-categorial and
dimensional. We adopted these two additional attributes classes because they
are very common in e-commerce products. Third, in our case there is no need to
ask the user to provide the class of each attribute involved in the comparison.
Fourth, while Kagie's work seems to focused on a single category, our approach
was conceived to deal with many categories typically found in e-commerce sites.
Fifth, we instead of using a 2D map with several products, our approach can
produce, as output, a ranking of products in order of similarity.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Attribute Classi cation</title>
      <p>Prior to the application of our similarity function, it is necessary to take each
attribute found in the products of a given category we are interested in and
assign each one to a single class of a simple attribute taxonomy comprising four
classes, namely: Numerical, Categorical, Multicategorical and Dimensional.</p>
      <p>
        This taxonomy was created based on previous work by Kagie et. al. [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ] and
in our own experience in dealing with e-commerce catalogs. The original
taxonomy by Kagie et. al. in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] included only Numerical and Categorical attributes. It
was extended in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to include the Multicategorical class. We further expanded it
with the Dimensional class to handle the common case of atributes that describe
the dimensions of products, displays, etc.
      </p>
      <p>
        Although a number of di erent approaches could have been used for this
task, we opted for using a simple strategy in which the values expected for the
attributes in a given class are described by a regular expression we call domain
descriptors. Domain descriptors are similar to the Data Frames used by Embley
et. al. in several methods (e.g., in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]) and provide a description on how values
of attributes of the four classes above are written.
      </p>
      <p>The classi cation of a attribute is carried out as follows. Let Ai be an
attribute that occurs for products p1;: : :;pm in a given category. For instance,
attribute Scene Modes occurs in the description of many products in the
Compact Cameras category. First, for all products pj (1 j p), we take the value vi;j
for Ai occurring in pj .</p>
      <p>Next, we perform several cleaning and standardization operations over set of
values vi;j of Ai taken from products. These operations include duplicate values
removal, white space and case normalization, among others. The result is a set of
values a1;: : :;am which we call the occurrences of Ai. Notice that by doing so we
assume that all values of Ai have the same semantics in all pj . For instance, we
assume that the attribute Scene Modes has the same semantics in the description
of all products in the Compact Cameras category.</p>
      <p>Finally, we test each occurrence a1;: : :;am against each domain descriptor k
(k=1;: : :;4) and associate atribute Ai with the atribute class Ck whose domain
descriptor k recognizes the majority of its occurrences.</p>
      <p>Although simple, this classi cation procedure is very e ective as we
demonstrate in experiments we carried out and report later in this paper.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Similarity Function</title>
      <p>
        Based on the general coe cient similarity proposed by Gower [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], we propose a
similarity function for comparing products as the sum of all non-missing
similarity scores sijk over the maximum number of attributes present in one of the
products according to Equation 1.
      </p>
      <p>Sij =</p>
      <p>K K
X mikmjksijk= max(X
k=1 k=1</p>
      <p>K
mik; X
k=1
mjk)
(1)</p>
      <p>In this equation, similarity scores sijk are computed for every atribute Ak
that has value for both products pi and pj being compared. Also, mik (mjk) is
0 when the value for attribute Ak is missing for products pi (pj ) and 1 when it
is not missing.</p>
      <p>The speci c functions used for computing the similarity score sijk depend on
the class of the attribute Ak. Recall from Section 3 that this class was already
de ned. For each one of the four atribute classes we de ned an appropriate
similarity function.</p>
      <p>For the Numerical class, the similarity function is de ned as the absolute
di erence between the values of the attribute in the two products, as shown in
Equation 2.
implying that objects having the same value get a similarity score of 1 and 0
otherwise.</p>
      <p>N
sijk = 1</p>
      <p>jvik vjkj
max (vik; vjk);
C
sijk = 1(vik = vjk)
where vik and vjk are, respectively, the values of the attribute k for products pi
and pj .</p>
      <p>For the Categorical class, the similarity function is de ned as
(2)
(3)
siMjk = jvik \ vjkj</p>
      <p>jvik [ vjkj
In this case vik and vjk denote the sets of individual categorical values composing
the actual values. For instance, vik would be fAuto, On, O , Slow Syncro, : : : g
for the attribute Flash Modes in the camera Sony W830 of Figure 1.</p>
      <p>For the Dimensional class, the similarity function is the normalized euclidean
distance over the dimension values, as described in Equation 5.</p>
      <p>
        For the Multicategorical class, the similarity function is computed using the
Jaccard coe cient [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] between the sets of
      </p>
      <p>D
sijk = 1</p>
      <p>D
[X((vidk)0
d=1
(vjdk)0 )2] 21
(4)
(5)
(6)
where D is the number of dimensions found in the values of the attribute and
vidk is the value for dimension d in vik (the same applies to vjdk). For computing
this function, each dimension is mean-centered and normalized using
(vxdk)0 = ((vxdk)
d)= d
d and d are, respectively, the mean and the standard deviation of the set of
values of dimension p in all values of atribute k, for the products in the category.</p>
      <p>
        As a nal comment, it is worth noting that the general coe cient similarity
proposed by Gower [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and latter used by Kagie et. al in [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ], is unsuitable to
deal with objects with few common attributes. For instance, if directly applied to
the problem of comparing products, when two products have just one common
attribute and this attribute have the same value in both products, the Gower
similarity measure will assign the highest similarity score between these two
products. Our function tries to overcome this problem by penalizing the score
when the products have few common attributes, as de ned in Equation 1.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Results</title>
      <p>In this section we report the results of experiments we performed to evaluate
the attribute classi cation strategy presented in Section 3, and the similarity
function described in Section 4.
5.1</p>
      <sec id="sec-4-1">
        <title>Experimental Setup</title>
        <p>For the experiments, we have used ve datasets provided by Neemu1, a company
that develops search and recommendation technology for major e-commerce sites
in Brazil. These datasets comprise ve di erent popular product categories,
namely: Cameras, Camcorders, Laptops, Smartphones and TVs. The product
1 http://www.neemu.com
descriptions available in these datasets often provide many attributes that are
not related to the product characteristics themselves. For instance, attributes
related to the packing of the products such as, packing dimension, package
contents, etc., are very common. Thus, we disregarded these attributes in our
experiments. In addition, we removed all atributes that are not found in at least
20% of the products in a given category. By doing so, we tried to increased
the percentage of attributes that can be e ectively compared to calculate the
similarity between products.</p>
        <p>Table 1 compares the number of attributes originally available in each dataset
and the nal number of attributes we considered in each category. Notice that,
even though many attributes were removed, still the number of attributes
considered is large to be handled manually by humans. This table also presents the
number of distinct products available in each dataset.</p>
        <p>In Table 2, we present the number of attributes in each of the classes of our
taxonomy. This classi cation was carried out manually to be used as a golden
standard. Notice that the large majority of the attributes are categorical. This
trend was observed in all categories. Also, a single dimensional attribute was
available in each category,.
In Table 3, we summarize the results obtained with our attribute classi cation
strategy. For this, we used the well known Precision, Recall and F-1 metrics.
In this table, each line corresponds to the results obtained with attributes of
a distinct classe, namely, \NUM" (Numerical), \CAT" (Categorical), \MCA"
(Multicategorical) and \DIM" (Dimensionall).</p>
        <p>As it can be notice, our strategy has correctly classi ed most of the attributes
from all categories we tested. We obtained perfect classi cation in many cases
and F-1 values equal or above 0.8 were obtained in all cases but one. This case is
the Multicategorical class in the Laptops category, which has a single attribute
(see Table 2), and our classi cation strategy missed it. In many cases, the value
of some attributes eventually presented noise our cleaning operation was unable
to identify and x Nevertheless, we believe the small number of failures does not
compromise the e ectiveness of our strategy and, as will see next, did not harm
the overall results of our method.
5.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Similarity Measure Evaluation</title>
        <p>Evaluating the e ectiveness of the similarity measure we described in Section 4
proved to be a challenge by itself. Indeed, carrying out a thorough evaluation to
obtain values of Precision, Recall and F-1 would require to compare hundreds of
products, examining the values of tens of attributes, some of them very technical.
Thus, we opted to evaluate our proposed similarity measure in a task close to
its intended application. This task consists in taking a product given as input,
using the similarity measure to compare this product with all others in the same
category, and verifying if the k products deemed as the most similar are indeed
similar to the input product, according to a human-based evaluation. The results
are reported in terms of the precision considering these top-k answers, a metric
often known as P@k. In our case we used k = 5, which is reasonable in terms of
recommender systems.</p>
        <p>For each of the ve product categories, we randomly selected 10 products,
which we refer to as query products, and, for each of them, we examine the 5
most similar products in the same category according to our similarity measure.
Thus, a total of 250 pairs of products were manually evaluated. The results are
presented in Figure 2.</p>
        <p>In Figure 2, each graph corresponds to a product category and shows the
P@5 values resulting from each of the 10 query products, along with the average
of the ten values. Our similarity measure led to P@5=1 in 22 out of the 50 query
products. Only in 8 cases, the P@5 values were below 0.5. In all categories, the
average of P@5 values was around 0.75. An average above 0.8 was observed
for the TVs category. Notice that the very low P@5 values obtained for some
queries (e.g., 0 for query 1 in Smartphones or 1 for query 9 in Cancorders) does
not necessarily implies that our similarity measure failed. For instance, it might
happen that the query product has very few or none similar product in the
catalog. In this case, our function just gave a low similarity score, but no similar
products would appear among the top-5 answers. To solve this, a threshold on
similarity score could be applied. However, there is no obvious way of imposing
this threshold. Thus, we leave this study for future work.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>In this paper we presented a preliminary study on the problem of nding
products similar to a product given as input. This problem, although important for
e-commerce sites, has been ill addressed so far both in the industry and in the
academy. We described and evaluated a similarity function we have proposed
for comparing products based on their attributes. Our function is generic in the
sense that it deals di erent types of attributes occurring in products from
distinct categories. Prior to its application, the function requires that each attribute
has been classi ed into to a class that determines an speci c similarity function
that handles this attribute. We demonstrate that this classi cation can be carry
out by a simple but highly e ective strategy we proposed, which relies of
regular expressions. Experiments we have performed with our similarity function
with datasests with real products, revealed that it is accurate in nding similar
products, achieving average F-1 values above 0.75 in 5 representative product
categories.</p>
      <p>
        Our plans for future work address two main aspects. First, we are working
on improving the e ectiveness of our function by considering that di erent
attributes may have di erent degrees of importante for users when comparing two
products of a given category. Thus, we are investiganting ways for capturing this
knowledge from the user and using it to improve our function. For this, we have
been working on machine learning techniques, which require training from user
data. Thus, the second aspect we are currently addressing is on how to obtain
training data without requiring users to label instances speci cally for this
problem. Another interesting future work we plan to address is considering additional
similarity functions for attributes. For instance, in the case of categorical data
it is worth investigating the metrics studied in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Al-Muhammed</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Embley</surname>
          </string-name>
          .
          <article-title>Ontology-based constraint recognition for freeform service requests</article-title>
          .
          <source>In IEEE 23rd International Conference on Data Engineering</source>
          , pages
          <volume>366</volume>
          {
          <fpage>375</fpage>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Boriah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chandola</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Kumar</surname>
          </string-name>
          .
          <article-title>Similarity measures for categorical data: A comparative evaluation</article-title>
          .
          <source>In Proceedings of the SIAM International Conference on Data Mining</source>
          , pages
          <volume>243</volume>
          {
          <fpage>254</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>R.</given-names>
            <surname>Burke</surname>
          </string-name>
          .
          <article-title>Knowledge based recommender systems</article-title>
          . In J.
          <string-name>
            <surname>Daily</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Kent</surname>
          </string-name>
          , and H.Lancour, editors,
          <source>Encyclopedia of Library and Information Science</source>
          , volume
          <volume>69</volume>
          .
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>S.-H.</given-names>
            <surname>Cha</surname>
          </string-name>
          .
          <article-title>Comprehensive survey on distance/similarity measures between probability density functions</article-title>
          .
          <source>International Journal of Mathematical Models and Methods in Applied Sciences</source>
          ,
          <volume>4</volume>
          (
          <issue>1</issue>
          ):
          <volume>300</volume>
          {
          <fpage>307</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Gower</surname>
          </string-name>
          .
          <article-title>A general coe cient of similarity and some of its properties</article-title>
          .
          <source>Biometrics</source>
          ,
          <volume>27</volume>
          (
          <issue>4</issue>
          ):
          <volume>857</volume>
          {
          <fpage>874</fpage>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>M.</given-names>
            <surname>Kagie</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. van Wezel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Groenen</surname>
          </string-name>
          .
          <article-title>Choosing attribute weights for item dissimilarity using clikstream data with an application to a product catalog map</article-title>
          .
          <source>In Proceedings of the 2008 ACM Conference on Recommender Systems</source>
          , pages
          <fpage>195</fpage>
          {
          <fpage>202</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>M.</given-names>
            <surname>Kagie</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. van Wezel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Groenen</surname>
          </string-name>
          .
          <article-title>A graphical shopping interface based on product attributes</article-title>
          .
          <source>Decision Support Systems</source>
          ,
          <volume>46</volume>
          (
          <issue>1</issue>
          ):
          <volume>265</volume>
          {
          <fpage>276</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J.</given-names>
            <surname>Schafer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Konstan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <article-title>E-commerce recommendation applications</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          -2):
          <volume>115</volume>
          {
          <fpage>153</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>