<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Leveraging Catalog to Resolve Conflicting Query Atributes in E-commerce Sites</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Suhas Ranganath</string-name>
          <email>suhas.ranganath@walmartlabs.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Walmart Labs</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <abstract>
        <p>Millions of people use online e-commerce platforms to search and buy products. Identifying attributes in a query is a critical component in connecting users to relevant items. However, in many cases, the queries have multiple attributes, and some of them will be in conflict with each other. For example, the query “maroon 5 dvds” has two candidate attributes, the color “maroon” or the band “maroon 5”, where only one of the attributes can be present. In this paper, we address the problem of resolving conflicting attributes in e-commerce queries. A challenge in this problem is that knowledge bases like Wikipedia that are used to understand web queries are not focused on the e-commerce domain. E-commerce search engines, however, have access to the catalog which contains detailed information about the items and its attributes. We propose a framework that leverages catalog information to resolve conflicting attributes in e-commerce queries. Our experiments on real-world queries on e-commerce platforms demonstrate that resolving conlficting attributes by leveraging catalog information significantly improves attribute identification, and also gives out more relevant search results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        E-commerce sites are being used by millions of people to buy
products in a fast and seamless manner. Users express their buying
needs through search queries, and an accurate understanding of
the query is necessary to return relevant items [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. A crucial part
of query understanding is to identify attributes inherent in the
query [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. For example, identifying that the query “maroon 5 dvds”
has product type “dvd” helps in returning relevant items. Research
on identifying query attributes in web search explores the use of
semantic information [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], user engagement [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and external
knowledge bases [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. There has been relatively less work in identifying
query attributes in the e-commerce domain.
      </p>
      <p>In many cases, query understanding systems have conflicting
candidate attributes for a given query. In the example query,
“maroon 5 dvds” the candidate attributes are the product type “dvds”
the color “maroon” and the band “maroon 5”. It is not
straightforward for query understanding systems to infer whether the
query is referring to the band “maroon 5” or the color “maroon”.
Designing algorithms to resolve conflicting query attributes can
Permission to make digital or hard copies of part or all of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed
for profit or commercial advantage and that copies bear this notice and the full citation
on the first page. Copyrights for third-party components of this work must be honored.
For all other uses, contact the owner/author(s).</p>
      <p>SIGIR 2018 eCom, July 2018, Ann Arbor, Michigan, USA
© 2018 Copyright held by the owner/author(s).</p>
      <p>ACM ISBN 978-x-xxxx-xxxx-x/YY/MM.
https://doi.org/10.1145/nnnnnnn.nnnnnnn
help e-commerce search systems to return more relevant items and
better satisfy the buying needs of users.</p>
      <p>
        This task faces several challenges. First, queries are short and
contain insuficient information for systems to identify attributes.
Second, knowledge bases like Wikipedia used to supplement query
text in web search [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are not focused on e-commerce domain, and
can lead to insuficient and noisy information. Third, e-commerce
sites have millions of users, and search algorithms have significant
issues of scalability.
      </p>
      <p>E-commerce search systems have access to a catalog which
contains the attributes of items sold by the system. Leveraging the
e-commerce catalog as a knowledge base to supplement the
textual information can help to resolve conflicting query attributes. In
the example query ”maroon 5 dvds”, we see from the catalog that
there is a significant number of items having the product type as
“dvds” and have a band attribute whereas very few of them have the
color attribute. This indicates that the catalog can provide valuable
information which can use to resolve conflicting query attributes.
Therefore, in this paper, we propose a framework to model catalog
information to better identify attributes in e-commerce queries.</p>
      <p>Specifically, we address the following questions: How to model
the catalog information to resolve conflicts in query attributes?
How to evaluate the impact of the framework on on e-commerce
search systems? The primary contributions of the work are
• Proposing the problem of resolving conflicts in query
attributes for e-commerce queries;
• Proposing a framework to model catalog information to
identify query attributes in e-commerce queries; and
• Presenting evaluations of the utility of catalog information
in identifying query attributes on real-world data.</p>
      <p>The rest of the paper is organized as follows. In Section 2, we
describe the proposed framework. In Section 3, we present
evaluations of the framework for identifying the query attribute and its
impact on ranking relevant items. We conclude in Section 4 along
with possible future directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>THE PROPOSED FRAMEWORK</title>
      <p>In this section, we present our proposed framework to identify
attributes in e-commerce queries. We first describe the notations used
and then define the problem statement. We then use the notations
to describe various aspects of the framework.</p>
      <p>We now present the notations used in the paper. Let q be the
query and A = {a1, a2, ..., an } be the set of candidate attributes in
the query. Let {A−ak } be the set of attributes of size Na without the
attribute ak . The problem can be formally stated as follows: “Given
the query q, attribute ak , and the attribute set {A − ak } determine
whether attribute ak is present in q.”</p>
      <p>We next present our framework to identify attribute values for
a given query. We explore two sets of metrics to model catalog
information to assist in query attribute identification. We next
describe two sets of metrics along with mathematical formulations,
one related to the presence of the attribute in the query, and the
second related to the presence of attribute value in the query.</p>
      <p>The first metric set computes p(m/q), the probability of an
attribute m being present in the query q. This is formulated as
(1)
(2)
p(m/n = x ) =
p(m/q) ∝ p(m/n = x ) ∗ p(n = x /q)
mn
nx
,
where p(m/n = x ) is the probability of attribute m being present
where n = x , and nx is the number of items in the catalog which
have value x for attribute n. Among such items mn is the number
of items having values for attribute m. For a given attribute m, we
compute Eq 1 for all attributes n ∈ {A − m}, resulting in a total of
Na − 1 feature values. According to Eq 1, if the query has attribute
value n = x , it is more likely to have attribute m if more items in the
catalog with n = x also contain attribute m. In the query “maroon
5 dvds”, very few items which have the product type “dvds” have
values for color, and Eq 1 metric gives a lower value for p(color /q),
the probability of color attribute being present in the query.</p>
      <p>The second metric computes p(m = l /q), the probability of
attribute m having a value l for the query q. This is formulated as
p(m = l /q) ∝ p(m = l /n = x ) ∗ p(n = x /q)
p(m = l /n = x ) =</p>
      <p>loд(nx )
where the number of items in the catalog having the value x for a
given attribute n is nx . Among these items, let the number of items
having value l for attribute m be denoted by mn . The score is higher
l
if more number of items having the value x for attribute n also have
value l for attribute m. We repeat this for all possible attributes
n ∈ {A − m} for a given attribute n resulting in an additional
Na − 1 feature values. The value set for a given attribute follow
a power law distribution where few values are prominent, so we
employ log smoothing to make the values linearly distributed.</p>
      <p>We integrate the scores derived from the catalog metrics into a
feature set. Our metrics are scalable and hence suited for handling
large-scale trafic common in an e-commerce site. We use an out of
the box classifier on the feature set to determine whether or not
the given query has the attribute value.</p>
    </sec>
    <sec id="sec-3">
      <title>3 EVALUATION</title>
      <p>We next evaluate our framework with the help of trafic weighted
random sample of 20000 queries on Walmart.com. To compare
our framework, we use the baseline Dict Lookup which identifies
attributes for a query by matching overalapping phrases in the
query with terms in the attribute dictionary. This baseline does
not address the possible conflicts that can arise between candidate
attributes. For the evaluation, we take color as the attribute that
has to be predicted for a query and the product type and brand as
the attributes whose values are known for the query.</p>
      <p>We design two evaluation tasks. The first task assesses the
effectiveness of the framework on identifying attributes of a given
S. Ranganath et al.</p>
      <sec id="sec-3-1">
        <title>Dict Lookup</title>
      </sec>
      <sec id="sec-3-2">
        <title>Framework</title>
      </sec>
      <sec id="sec-3-3">
        <title>Gain</title>
      </sec>
      <sec id="sec-3-4">
        <title>Precision</title>
        <p>1.0
1.05
+5.36%
query. We employ manual labeling by expert annotators for the
ground truth and use Precision, Recall, and F1 as the evaluation
metrics. The second task assesses the impact of the framework on
the ranking relevant items for the given query. We use the orders
of the query-item pair for the ground truth and nDCG@20 as the
evaluation metric. The evaluation results are illustrated in Table 1.</p>
        <p>From the table, we can see that the framework is significantly
better in identifying attributes of a given e-commerce query than
the baseline across all the Precision, Recall and F1 metrics. The
improvement in identifying query attributes is also reflected in
showing better ranking results as shown by the lift in nDCG@20.
The improvement in both the tasks demonstrates that the catalog
can be efectively leveraged as a knowledge base to identify
attributes for a given query in a better manner, and the ability of the
metric to efectively capture the relevant catalog information.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper, we address the problem of identifying attributes for
queries on e-commerce sites. General purpose knowledge bases
used in identifying attributes for web queries are not focused on
e-commerce needs. We design a framework that leverages catalog
as a knowledge base to resolve conflicts in query attributes. We
evaluate the framework on the set of queries from Walmart.com
and demonstrate that it significantly improves results in attribute
identification and ranking relevant items for e-commerce queries.
Future research directions can include leveraging query catalog
interactions and query sequences to design more involved metrics
for attribute identification. The utility of catalog in other query
understanding tasks such as query reformulation and type-ahead
is also an interesting avenue for researchers to explore.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Ricardo</given-names>
            <surname>Baeza-Yates</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Semantic Query Understanding</article-title>
          .
          <source>In Proceedings of the 40th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR '17)</source>
          . ACM, New York, NY, USA,
          <fpage>1357</fpage>
          -
          <lpage>1357</lpage>
          . https: //doi.org/10.1145/3077136.3096472
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jian</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Gang</surname>
            <given-names>Wang</given-names>
          </string-name>
          , Fred Lochovsky, Jian-tao
          <string-name>
            <surname>Sun</surname>
            , and
            <given-names>Zheng</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Understanding user's query intent with wikipedia</article-title>
          .
          <source>In Proceedings of the 18th international conference on World wide web. ACM</source>
          ,
          <volume>471</volume>
          -
          <fpage>480</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Hang</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Zhengdong</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep learning for information retrieval</article-title>
          .
          <source>In Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval. ACM</source>
          ,
          <volume>1203</volume>
          -
          <fpage>1206</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Yanen</given-names>
            <surname>Li</surname>
          </string-name>
          , Bo-June Paul Hsu, and
          <string-name>
            <given-names>ChengXiang</given-names>
            <surname>Zhai</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Unsupervised identification of synonymous query intent templates for attribute intents</article-title>
          .
          <source>In Proceedings of the 22nd ACM international conference on Conference on information &amp; knowledge management. ACM</source>
          ,
          <year>2029</year>
          -
          <fpage>2038</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Xiang</given-names>
            <surname>Ren</surname>
          </string-name>
          , Yujing Wang,
          <string-name>
            <surname>Xiao Yu</surname>
          </string-name>
          , Jun Yan, Zheng Chen, and Jiawei Han.
          <year>2014</year>
          .
          <article-title>Heterogeneous graph-based intent learning with queries, web pages and wikipedia concepts</article-title>
          .
          <source>In Proceedings of the 7th ACM international conference on Web search and data mining. ACM</source>
          ,
          <volume>23</volume>
          -
          <fpage>32</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Zhongyuan</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kejun Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Haixun</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiaofeng Meng</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ji-Rong Wen</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Query Understanding through Knowledge-Based Conceptualization</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>