<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Noise-aware Missing Shipment Return Comment Classification in E-Commerce</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Avijit Saha∗</string-name>
          <email>avijit.saha@flipkart.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vishal Kakkar∗</string-name>
          <email>vishal.kakkar@flipkart.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>T. Ravindra Babu</string-name>
          <email>ravindra.bt@flipkart.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E-Commerce, Text, Comment, Noise, Data Programming, Deep</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Flipkart Internet Private Limited</institution>
          ,
          <addr-line>Bangalore</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Learning</institution>
          ,
          <addr-line>Machine Learning</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <abstract>
        <p>E-Commerce companies face a number of challenges in return requests. Claims of missing-items is one such challenge, where customer claims that main product is missing from shipment through return comments. It is observed that dominant part of such claims are inadvertent given the limited literacy of customers. Some of them have fraud intent. At Flipkart, such claims are evaluated manually to examine whether the comment relates to missing item. Classification of the claim intent automatically saves human bandwidth and provides good customer experience by reducing the turn around time to customers. However, this is challenging as comments are replete with spell variations, non-English vernacular words, and are often incomplete and short. This is compounded by noisy labeling of such comments due to human bias and manual errors. To classify the claim intent, we apply conventional as well as deep learning methods. To handle label noise, we employed stateof-the-art noise-aware techniques, which fail to perform due to pattern specific label noise. Motivated by the wide pattern specific label noise, we encode domain heuristics as labeling functions (LFs) which label subsets of the data. However, LFs may conflict and prone to noise. We address the conflict by defining a conflict-score to rank the LFs. Proposed method of noise handling with LFs out performs all the state-of-the-art noise-aware baselines.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>E-commerce companies face a large number of return requests of
various types (reason-codes). Missing-item is one such reason-code
where customer claims that main product is missing from shipment.
∗Equal Contributions
Permission to make digital or hard copies of part or all of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed
for profit or commercial advantage and that copies bear this notice and the full citation
on the first page. Copyrights for third-party components of this work must be honored.
For all other uses, contact the owner/author(s).</p>
      <p>SIGIR 2018 eCom, July 2018, Ann Arbor, Michigan, USA
© 2018 Copyright held by the owner/author(s).</p>
      <p>ACM ISBN 123-4567-24-567/08/06.
https://doi.org/10.475/123_4
Unlike a usual return, product pick up from customer is avoided
in a missing-item return. Example of a missing-item return is
customer ordered a handset and received a stone in place of the
handset. A confirmed case of missing item results in loss to the
company since there is a definite fraud with one of the stakeholders
such as buyer, seller or delivery team. Hence, a careful scrutiny is
necessary before the approval of any missing-item returns.</p>
      <p>A missing-item return request is generated broadly for two
reasons. The first can be due to definite fraud with one of the
stakeholders. Secondly, given limited literacy levels of customers, it is
observed that the claims do not always refer to missing-item but
inadvertently claimed as missing-item. For example, a missing-item
return with customer comment ‘I did not like the item.’ clearly
indicates that the return belong to a return category other than
missingitem. Because here customer received the main product. Hence, the
return should be cancelled. We call this a comment-mismatch
(customer’s comment does not match with the return reason-code).
On the other hand, a missing-item return with customer comment
‘I ordered a phone but received an empty box.’ clearly indicates
that the return belong to the missing-item category. This return
should be approved. This is referred by non-comment-mismatch
(customer’s comment matches with the return reason-code).</p>
      <p>At Flipkart, a dedicated operation team assesses the compatibility
of customers’ comments on missing-item return with the
missingitem reason-code. Each missing-item return – passes through this
process and – is rejected when the return comment is
incompatible with the missing-item reason-code, otherwise approved. This
process wastes lot of human bandwidth and is prone manual
mistakes. Moreover, it hampers the customer experience due to the
lag between the return placement time and its status update to the
customer. Hence, we want to automate this process. In terms of
Machine Learning problem, given a missing-item return comment,
we want to predict whether it is a comment-mismatch (positive
class) or non-comment-mismatch (negative class).</p>
      <p>
        Customer comments are generally very noisy mainly because
of three reasons: a) spelling mistakes: empty is misspelled in
comment ‘Emety box’, b) usage of regional languages: comment ‘Galt
order ho gya h’, which means I ordered wrong product, uses Hindi
language, and c) varied comment length: token counts in a
comment ranges [
        <xref ref-type="bibr" rid="ref1">1, 323</xref>
        ]. To handle such dificulties, we employ word
embedding and meta features. Besides conventional classification
method (xgboost [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]), we use BLSTM [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] to capture the sequential
information in the comments.
      </p>
      <p>
        Also, the manual label generation process, which marks a
missingitem return comment as comment-mismatch/non-comment-mismatch,
is very noisy. The sources of noise are manual error, human bias,
and lack of well calibrated operation team. The label noise varies
depending on the patterns and is non-iid. This is clearly visible by
the fact that the overall label noise is ∼ 15% and a specific pattern
‘I ordered x quantity of an item but received y quantity’ has 50%
noise. Due to this reason, the state-of-the-art noise-aware [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]
baseline, which uses BLSTM as the base model, fails severely. In fact, we
observed a performance degradation of this noise-aware BLSTM
over vanilla BLSTM.
      </p>
      <p>
        The pattern specific noise variation and high noise on certain
patterns motivate us to use noise correction based on domain heuristics.
Like data programming[
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], we express weak supervision
strategies or domain heuristics as labeling functions (LFs) which label
subsets of the data. However, LFs may conflict and prone to noise.
To our best knowledge, no one has employed LFs to rectify label
noise.
      </p>
      <p>In a typical data programming setting, only data points are
available and LFs are created to generate labels. The LFs may conflict on
certain data points and have varying error rates. To handle it, data
programming defines a generative process over the LFs to learn
the correctness probability of each LF on each data instance. This
information is then used to fit a noise-aware classifier.</p>
      <p>The key diference between our setting and the data
programming is the availability of noisy labels (we call it as true LF) in our
case. Unlike data programming, we would like introduce less noisy
LFs compared to the true LF. Due to all these reasons, instead of
applying data programming directly, here resort to a simple method.
and promising methods to correct label noise using LFs, and leave
out the exploration of data programming as future work.</p>
      <p>We define multiple LFs to alter the noisy labels in our dataset.
However, applying them directly to flip the noisy labels is
impossible because the LFs conflict with each other - same data instance is
labeled as positive by a LF and negative by another LF. This requires
us to generate a ranked list of LFs. To do that, we define a conflict
score, which captures how less a LF conflicts with other LFs. Then,
the ranked LFs are applied to alter the true labels. In case of conflict
between LFs, the LF with the least conflict score is chosen to alter
the noisy label. We show that proposed method of noise handling
with LFs out performs all the state-of-the-art noise-aware baselines
as well as vanilla baselines.</p>
      <p>We integrate the following aspects in the paper.</p>
      <p>• Explanation of a real world problem and its challenges
• Exploratory data analysis and feature engineering
• State-of-the-art baselines - xgboost and LSTM and their
noise-aware variants
• Noise handling with labeling functions
• Proposed conflict-score to handle conflicts
• Shown superior performance of the proposed method
Section 2 contains insights into data. Related work, feature
engineering, modeling, experimentation, and conclusion are described
in Section 3, 4, 5, 6, and 7, respectively.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>DATA</title>
    </sec>
    <sec id="sec-3">
      <title>Description</title>
      <p>
        The data in our study comes from the customer comments on
missing-item return requests. We use one year of customer
comments: April 2017 - Mar 2018. We consider a comment-mismatch as
positive label and a not-comment-mismatch as negative label. Table
1 describes the dataset statistics. We can observe ∼ 41% are positive
labels. We count the number of words in a comment by tokenizing
it. Table 1 also shows that the word count of comments ranges
[
        <xref ref-type="bibr" rid="ref1">1, 323</xref>
        ]. The plot clearly shows the existence of widely variable
length comments in our dataset. To deep dive, we plot a histogram
of number of comments w.r.t. word count in Figure 1. We clip the
plot at word length 75 for better visualization. Clearly, it is a long
tail distribution - approximately 51% of comments has less than ten
words. However, there exists a fat tail of comments with very high
word count.
      </p>
      <p>Table 2 shows example of customer comments with diferent
word count. Interestingly, we can observe the presence of spell
error - ‘My mistek’ and regional language - ‘Sir khali box mila
h’ (Hindi). Often, comments are very noisy and do not adhere to
grammar rules - ‘Not working Good; tow time same product bye
mistack accepted so remove it’. To provide more insight on the
data, we show a word cloud of our dataset in Figure 2. Some of
the important keywords are cx (customer), product, mobile, box,
missing, and empty. The cx words occurs in the comments when
customer calls customer care to place the return request.
Recall, non-comment-mismatch implies a genuine return request of
type missing-item and comment-mismatch implies a return request
anything other than missing-item type. Given a missing item return
comment, our operation team at Flipkart mark it as either
commentmismatch (positive) or non-comment-mismatch (negative). This
label generation process is very noisy due to manual error, human
bias, and lack of well calibrated operation team. Table 3 shows
example of noisy labels for comments with diferent word count.
Actual label and true labels represents the label in the dataset and
the expected label, respectively. We show noisy labels from one,
two, four, and eight word count comments. For each word count,
we show one comment whose label should be positive but marked
as negative and one comment whose label should be negative but
marked as positive.
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>Noise Statistics</title>
      <p>To estimate the overall noise in our dataset, we manually relabeled
3k randomly chosen examples and consider it as test dataset. We
calculate the noise by considering mismatch between the actual
label and true label on the test dataset. The overall noise is estimated
to be 15.65%. Also, the positive and negative classes have 10.06%
and 19.45% noise, respectively.
3</p>
    </sec>
    <sec id="sec-5">
      <title>RELATED WORK</title>
      <p>
        Preprocessing is an important step for text classification. Two
important blocks [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] of preprocessing are - 1) Tokenization and 2)
Filtering. Tokenization [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], which is the initial step of
preprocessing, divides a text document into words known as tokens. In
Filtering, stop words are removed.
      </p>
      <p>
        Then, the preprocessed text is converted into feature vectors. One
of the widely used model for feature generation is the bag-of-words
(BOW) model [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. It represents a document to a k-dimensional
feature vector, where the individual co-ordinate represents the count
of a specific word in the document. Often, term frequency-inverse
document frequency (tfidf) [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] is used to penalize a frequently
occurring word. Recently, word2vec (w2v) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] model gained much
attentions. It embeds a word to a k-dimensional vector space by
preserving the property that words occurring in the same context
will have higher similarity score. Word2vec [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] features are used
for text classification in many ways - summation of the w2v
embedding vectors, mean of the w2v embedding vectors, and tf-idf
based weighted sum of the w2v embedding vectors corresponding
to all the words in a document.
      </p>
      <p>
        Support vector machines (SVMs) [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] are widely employed for
text classification. Athanasiou [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] has applied gradient boosting
machine for sentiment analysis task, and shown superior performance
over SVM, Naive Bayes, and neural network. Gradient boosting
machine is a boosting algorithm where each iteration fits a new model
to get better class estimation. Each newly added model is
correlated with the negative gradient of the loss function, and the loss is
minimized using gradient descent. Extreme gradient boosting
(Xgboost) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is another boosting algorithm with better regularization
and performs well in practice.
      </p>
      <p>
        Recently deep learning algorithms[
        <xref ref-type="bibr" rid="ref15 ref22">15, 22</xref>
        ] have shown
promising performance in text classification. Specifically, recurrent neural
networks (RNNs) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] are the widely used architectures to capture
the sequential information. Long term short memory networks
(LSTMs) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] is a variant of RNN which helps to overcome some of
the problem of RNN like, vanishing gradient problem and helps to
remember the context over long text. Many flavours of LSTMs are
proposed [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] to for text classification, such as Multilayer-LSTM,
Bidirectional-LSTM (BLSTM), and Tree-Structured LSTM. In
Multilayer LSTM, LSTMs are stacked over each other to capture the
non-linearity. In BLSTM, both past and future information are
preserved using two hidden states, and it helps to learn the context
better.
      </p>
      <p>
        Noisy label can be handled broadly by three approaches [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]: a)
label-noise robust models [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ], b) data cleansing methods [
        <xref ref-type="bibr" rid="ref13 ref5">5, 13</xref>
        ],
and c) label-noise tolerant learning algorithms [
        <xref ref-type="bibr" rid="ref11 ref3">3, 11</xref>
        ]. In label-noise
robust methods, label-noise is handled by reducing the overfitting.
Even though, theoretically learning algorithms are robust to
labelnoise, in practice performance varies from algorithm to algorithm,
such as bagging perform better than the boosting [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Data
cleansing methods filter out data points which appears to be mislabeled.
Filtering can be done with various approaches, such as outlier
detection [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], removal of all the misclassified data points by a
classifier [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and removal of any points which disproportionately
increases the model complexity [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In label-noise tolerant algorithm,
noise is handled explicitly in the modeling step. Label-noise robust
logistic regression [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] modifies the loss function to handle noise.
Recently, a probabilistic neural-network based framework [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is
developed, which views the true label as a latent variable and a
softmax layer is used to predict it. The noise is explicitly modeled
by an additional softmax layer that predict the noisy label based on
both the true label and the input features.
      </p>
      <p>
        As creating labeled training data is dificult and time consuming,
many approaches are developed to generate training data
automatically, such as distant supervision [
        <xref ref-type="bibr" rid="ref17 ref7">7, 17</xref>
        ] and data programming [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
Distant supervision heuristically maps a knowledge base of known
relations to an unknown domain to generate training data. Data
programming is a generic framework to create dataset pragmatically
using distant supervision. It expresses weak supervision strategies
or domain heuristics as labeling functions (LFs) which label
subsets of the data. The LFs may conflict on certain data points and
have varying error rates. To handle it, data programming defines a
generative process over the LFs to learn the correctness probability
of each LF on each data instance. This information is then used to
ift a noise-aware classifier.
4
      </p>
    </sec>
    <sec id="sec-6">
      <title>FEATURE ENGINEERING</title>
      <p>In this Section, we will discuss all the hand crafted features which
are used for conventional Machine Learning models.
4.1</p>
    </sec>
    <sec id="sec-7">
      <title>Meta Features</title>
      <p>We construct nine meta features as shown in Figure 3. Word count,
char count, alpha count, digit count, and non-alphanumeric count
compute the number of words, characters, alphabets, digits, and
non-alphanumeric characters in a comment. To show the
discriminating power of each meta feature, in Figure 3, we show histogram
of each feature w.r.t. both positive and negative class. In each
subplot, the blue and green histogram represents the distribution for
negative and positive class, respectively.</p>
      <p>In each subplot, the blue distribution is right shifted. It implies
that in general comments from the negative class are longer, have
more alphabets, digits and non-alphanumeric characters compare
to the comments from the positive class. This is explained by the
fact that short comments lack descriptive ability, and thus have
higher chance of being comment-mismatch. Interestingly, unique
character count and unique alphabet count are the most
discriminating features - there is a clear separation between the distribution
of positive and negative class for these two features.
4.2</p>
    </sec>
    <sec id="sec-8">
      <title>Word2vec Features</title>
      <p>
        We observed that our dataset is very noisy. For example, a keyword
like ‘missing’ has numerous variations - missing, misssing, misisng,
missig, etc. Moreover, there are semantically similar words, such as
‘empty’ and ‘khali’ both means empty. Also, customers use regional
language, such as hindi (‘mujhe product nahi mila’). To handle
these complexities, we train a word2vec model with 200 dimension
on one year customer’s return comments data from all the return
reason-codes. The number of comments are O(10M). We train the
word2vec by considering words which occurs at-least 25 times in
our corpus. For tokenization, we use Gensim [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] simple_preprocess
method.
      </p>
      <p>Table 4 shows similar words from the word2vec model for four
keywords - missing, mobile, product, and ordered. We can observe
that the word2vec model is able to capture spelling variations. It
is also able to capture semantically similar word, such as missing
and empty, mobile and handset, device and phone, and ordered
and purchased. Moreover, the regional language variation is also
captured, such as missing and khali (hindi), mobile and fone (hindi),
and missing and illa (tamil).</p>
      <p>4.2.1 Sum of Word2vec Features: We sum the individual
200dimensional embedding vector for each word in a comment and
use it as the final feature in model. To handle out-of-vocabulary
word, we fall back to the 200-dimensional zero vector.</p>
      <p>4.2.2 Weighted Word2vec Features: We fit a tfidf model on the
training data. Then, we calculate a weighted sum of the individual
200-dimensional embedding vector for each word in a comment.
Where weight of individual word embedding is assigned from the
tfidf score of the word.
4.3</p>
    </sec>
    <sec id="sec-9">
      <title>Bag-of-words (BOW) Features</title>
      <p>Each comment in the training dataset is preprocessed with Gensim
preprocessing. Then, a bag-of-words (BOW) model is trained with
5k vocabulary size and with English stop words removal.
5
5.1</p>
    </sec>
    <sec id="sec-10">
      <title>MODEL</title>
    </sec>
    <sec id="sec-11">
      <title>BASELINE</title>
      <p>We tried multiple conventional Machine Learning algorithms widely
used for text classification, such as SVM, naive Bayes, logistic
regression, random forest, and xgboost. In our data, around 50% of
the comments are long and we found that often long comments
have sequential information, such as ‘customer ordered 2 items but
received only 1’. To capture the sequential information, we
experimented with diferent RNN based models, such as RNN, LSTM,
Multi-layer LSTM, and Bi-directional LSTM (BLSTM). In below,
we only describe the best performing models from each of this
approach.</p>
      <p>5.1.1 Xgboost: A xgboost model is trained with all the features
described in Section 4. The model parameter is tuned by grid search.
While performing the grid search, we restrict the max depth of
individual tree to 8 and max number of estimators to 500. The best
parameters are chosen using 5-fold cross-validation.</p>
      <p>5.1.2 BLSTM:. We used BLSTM which avoids feature
engineering. In BLSTM, both past and future information are preserved
using two hidden states, and it helps to learn the context better.
We experimented with word2vec pretrained embedding from the
word2vec model as well as learning the embedding from scratch in
the network itself. We tune the number of neurons in the BLSTM.
We also experimented by adding fully connected relu layes in the
network before the output layer. The best parameters are tuned
based on a validation set.
5.2</p>
    </sec>
    <sec id="sec-12">
      <title>NOISE-AWARE BASELINE</title>
      <p>We tried two approaches to handle noisy labels: 1) data cleansing
method and 2) label noise-tolerant algorithm.</p>
      <p>
        5.2.1 Data Cleansing Method: The goal of such methods is to
iflter out data points which appears to be mislabeled. Here, we
apply a model prediction based filtering method [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In this filtering approach, we divide the training dataset into
k-folds (five-folds). We train a xgboost model (with the best
parameters found in 5.1.1) on k-1 folds and apply it to predict labels
for k-th fold. This process is repeated k times to get labels for
the entire training dataset. Then we filter out instances with
disagreement between the actual label and the predicted label. On
the filtered training data, a xgboost model (with same parameter)
is trained which forms the final model. We refer this method as
xgboost+filtering.</p>
      <p>
        5.2.2 Label Noise-Tolerant Algorithm: In label-noise tolerant
algorithm, noise is handled explicitly in the modeling step. Here, a
probabilistic neural-network based framework [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is considered to
handle label-noise. This framework views the true label as a latent
variable. A softmax layer is used to predict the true label. The noise
is explicitly modeled by an additional softmax layer that predict
the noisy label based on both the true label and the input features.
      </p>
      <p>Assuming the non-linear function applied on an input x be h =
h(x ), the true label y is modeled by:
p(y = i |x ; w) = Ík
l =1 exp(uTl h + bl )
exp(uTi h + bi )
, i = 1, ..., k
(1)</p>
      <p>Where k is the number classes and w is the network
parameterset (including the softmax layer). Next a softmax output layer is
added to predict the noisy label z based on both the true label and
the input features:
(2)
(3)
(4)
p(z = j |y = i, x ; wnoise) = Ík
l =1 exp(uTil h + bil )
exp(uTil h + bil )
p(z = j |x ) = p(z = j |y = i, x )p(y = i |x )</p>
      <p>Where, wnoise represents parameters in the second softmax layer.
Given n training data points with feature vectors x1, x2, ..., xn with
corresponding labels z1, z2, ..., zn and true labels y1, y2, ..., yn , the
log likelihood term of the model parameters is written as:
Õ log p(zt |xt ), t = 1, ..., n
t</p>
      <p>This modeling approach is named as c-model and learned using
a neural-network training. For our experiment, the h function is
considered to be a BLSTM, with the same parameter as in Section
5.1.2. This model is referred as BLSTM+noise-aware.
def lambdapartial(x ):
return 1 if # of positive integer tokens in x &gt;= 2 else 0
def lambdasoap(x ):
return -1 if IsTokenPresent(soap, x ) else 0
5.3</p>
    </sec>
    <sec id="sec-13">
      <title>NOISE HANDLING WITH WEAK</title>
    </sec>
    <sec id="sec-14">
      <title>SUPERVISION</title>
      <p>
        We encode domain heuristics as labeling functions (LFs) [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], which
label subsets of the data. However, LFs may conflict and prone to
noise. Assuming data point and class pair (x, y) are drawn from the
distribution X × {−1, 1}, a LF λi : X → {1, 0, −1} is a user-defined
function that encodes some domain heuristic, and provides
nonzero label for some subset of the data points. Where 1 and -1 refer
to the positive and negative class respectively. And 0 refers to the
case where LF can not label the instance. LFs collectively generate
a large but potentially overlapping set of training labels. LFs can be
created in many ways, such as leveraging domain specific patterns
to label data points or use existing knowledge bases to generate
labels.
      </p>
      <p>We consider that the actual labels came from a LF named λt rue ,
which has 100% coverage. As introduction of a more noisy LF than
λt rue will increase overall noise in the dataset, unlike data
programming, we introduce less noisy LFs than λt rue . Lets consider λ
denotes the m newly created LFs - {λi }im=1, each of which looks at
the domain specific patterns to label data points.</p>
      <p>A specific pattern is observed in the comments - ‘I ordered x
quantity of an item but received y quantity’. This is partial delivery,
as the customer has received part of the order, and should be marked
as comment-mismatch. With this pattern, a LF lambdapar t ial is
defined in Table 5. We also observed that customer often writes
they have received soap instead of a mobile phone. This is a
genuine missing-item request, and should be marked as
non-commentmismatch. With this pattern, a LF lambdasoap is also defined in
Table 5. The IsTokenPresent function returns true when soap is one
of the token of comment x . Table 6 describes the statistics of all the
LFs. Where coverage represents the percentage of instances with at
least one label, conflict represents the percentage of instances with
conflicting labels, and overlap depicts the percentage of instances
with more than one labels.</p>
      <p>To deep dive, we show conflict and overlap count in Table 7
and Table 8, respectively. Where a cell (i, j) in Table 7 defines the
number data instances where (λi == 1 and λj == −1) or (λi == −1
and λj == 1). Similarly, a cell (i, j) in Table 8 defines the number
data instances where (λi == 1 and λj == 1) or (λi == −1 and
λj == −1).</p>
      <p>As LFs conflicts with each other applying them directly to flip
the true labels is impossible. This requires us to generate a ranked
list of LFs. To do that, for each λi ∈ λ, we define a conflict-score
as below. Note, conflict count of λi is calculated by summing the
number of conflict between λi and each rule from λ − λi .
# of LFs
17</p>
      <sec id="sec-14-1">
        <title>Coverage (%)</title>
        <p>100</p>
      </sec>
      <sec id="sec-14-2">
        <title>Overlap (%)</title>
        <p>40</p>
        <p>Conflict (%)
9
true
partial
phone missing
soap
true
partial
phone missing
soap
# unique conflicts of λi with other LFs</p>
        <p>Coverage of λi</p>
        <p>Intuitively, the cf_score captures how much a LF conflicts with
other LFs. With this score, Algorithm 1 denoise the training data.
λsor t ed contains the sorted list of m newly created rules in
ascending order w.r.t the cf_scoreλi .</p>
        <p>Algorithm 1 Label Denoising with Labeling Functions (LFs)</p>
        <p>After denoising the training data with Algorithm 1, we apply the
xgboost and BLSTM model on the denoised training data. This two
approaches are named as xgboost+best-sequence and
BLSTM+bestsequence.</p>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>6 EXPERIMENTS</title>
    </sec>
    <sec id="sec-16">
      <title>6.1 Dataset</title>
      <p>The complete data consists of O(100k) instances out of which
randomly chosen 3k instances forms the test data and rest forms the
train data. The test dataset is manually relabeled to generate clean
labels. Table 9 shows train and test data statistics. Note, the test
dataset size is small because manual relabeling is time consuming.</p>
    </sec>
    <sec id="sec-17">
      <title>6.2 Experimental Setup</title>
      <p>We compare performance of xgboost+best-sequence and
BLSTM+bestsequence against the baselines - xgboost, BLSTM, xgboost+filtering,
and BLSTM+noise-aware. Moreover, to showcase the benefits of
the conflict handling with the cf_score, we compare the
performance of xgboost+best-sequence and BLSTM+best-sequence with
xgboost+random-sequence and BLSTM+random-sequence. For
randomsequence, λsor t ed consists of a of random permutation of m newly
created rules. Model performance varies for diferent random
sequences. Hence, for both xgboost+random-sequence and
BLSTM+bestrandom, we repeat experiments 10 times with diferent random
permutation of λ and report the mean and standard deviation. For
all the models, we fix random-state to 2018. As our dataset is well
balanced, we use accuracy as the evaluation metric.</p>
    </sec>
    <sec id="sec-18">
      <title>6.3 Results</title>
      <p>Table 10 shows the performance comparison between xgboost,
BLSTM, and their noise-aware variants. BLSTM is performing best
with an accuracy of 87.43. This proves that using sequence
information indeed benefits on our dataset. BLSTM and xgboost are
performing better than BLSTM+noise-aware and xgboost+filtering,
respectively. We can observe that the state-of-the-art noise-aware
algorithms are hurting the performance. We think that the
reason for such a performance degradation is due to the wide pattern
specific noise variation.
is quite low, and still BLSTM+best-sequence beats the
BLSTM+randomsequence, again proving conflict handling with cf_score helps.</p>
      <p>Overall, BLSTM+best-sequence performs best, proving the
efifcacy of our proposed approach. Best-sequence provide benefits
over random-sequence proving the benefits of conflict handling
among LFs by cf_score. Sequence model BLSTM is able to provide
benefits over xgboost. We were able to improve the accuracy from
86.90% to 90.04% with the help of BLSTM, LFs, and cf_score.</p>
    </sec>
    <sec id="sec-19">
      <title>7 CONCLUSION</title>
      <p>We discussed an important problem of classifying missing-item
return comments into comment-mismatch/non-comment-mismatch.
We highlighted the data and noise related challenges in both
comments and labels. We have experimented with the state-of-the-art
Machine Learning and Deep Learning methods as well as their
noise-aware variants. We have proposed a simple method with
labeling functions (LFs) to denoise the training dataset. A
conflictscore is defined to handle the conflicts between LFs. Empirically,
we have shown eficacy of our approach over the baselines.</p>
      <p>As future work, we intend to explore the complete data
programming framework to handle noisy labels.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Mehdi</given-names>
            <surname>Allahyari</surname>
          </string-name>
          , Seyed Amin Pouriyeh, Mehdi Assefi, Saied Safaei,
          <string-name>
            <given-names>Elizabeth D.</given-names>
            <surname>Trippe</surname>
          </string-name>
          , Juan B.
          <string-name>
            <surname>Gutierrez</surname>
            , and
            <given-names>Krys</given-names>
          </string-name>
          <string-name>
            <surname>Kochut</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Brief Survey of Text Mining: Classification, Clustering and Extraction Techniques</article-title>
          .
          <source>CoRR abs/1707</source>
          .02919 (
          <year>2017</year>
          ). arXiv:
          <volume>1707</volume>
          .02919 http://arxiv.org/abs/1707.02919
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Vasileios</given-names>
            <surname>Athanasiou</surname>
          </string-name>
          and
          <string-name>
            <given-names>Manolis</given-names>
            <surname>Maragoudakis</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A Novel, Gradient Boosting Framework for Sentiment Analysis in Languages where NLP Resources Are Not Plentiful: A Case Study for Modern Greek</article-title>
          .
          <source>Algorithms</source>
          <volume>10</volume>
          (
          <year>2017</year>
          ),
          <fpage>34</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Jakramate</given-names>
            <surname>Bootkrajang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ata</given-names>
            <surname>Kabán</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Label-Noise Robust Logistic Regression and Its Applications</article-title>
          .
          <source>In Proceedings of the 2012 European Conference on Machine Learning and Knowledge Discovery in Databases - Volume Part I (ECML PKDD'12)</source>
          . Springer-Verlag, Berlin, Heidelberg,
          <fpage>143</fpage>
          -
          <lpage>158</lpage>
          . https: //doi.org/10.1007/978-3-
          <fpage>642</fpage>
          -33460-3_
          <fpage>15</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Leo</given-names>
            <surname>Breiman</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <string-name>
            <given-names>Random</given-names>
            <surname>Forests</surname>
          </string-name>
          .
          <source>Mach. Learn</source>
          .
          <volume>45</volume>
          ,
          <issue>1</issue>
          (Oct.
          <year>2001</year>
          ),
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          . https: //doi.org/10.1023/A:1010933404324
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Carla E.</given-names>
            <surname>Brodley</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark A.</given-names>
            <surname>Friedl</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Identifying Mislabeled Training Data</article-title>
          .
          <source>J. Artif. Int. Res</source>
          .
          <volume>11</volume>
          ,
          <issue>1</issue>
          (
          <year>July 1999</year>
          ),
          <fpage>131</fpage>
          -
          <lpage>167</lpage>
          . http://dl.acm.org/citation.cfm?id=
          <volume>3013545</volume>
          .
          <fpage>3013548</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Tianqi</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>XGBoost: A Scalable Tree Boosting System</article-title>
          .
          <source>In Proceedings of the 22Nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD '16)</source>
          . ACM, New York, NY, USA,
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          . https://doi.org/10.1145/2939672.2939785
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Craven</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kumlien</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Constructing biological knowledge bases by extracting information from text sources</article-title>
          .
          <source>In Proceedings of the International Conference on Intelligent Systems for Molecular Biology.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Thomas</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Dietterich</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>An Experimental Comparison of Three Methods for Constructing Ensembles of Decision Trees: Bagging, Boosting, and Randomization</article-title>
          .
          <source>Machine Learning</source>
          <volume>40</volume>
          ,
          <volume>2</volume>
          (
          <issue>01</issue>
          <year>Aug 2000</year>
          ),
          <fpage>139</fpage>
          -
          <lpage>157</lpage>
          . https://doi.org/10.1023/A: 1007607513941
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Benoît</given-names>
            <surname>Frénay</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ata</given-names>
            <surname>Kaban</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A Comprehensive Introduction to Label Noise. i6doc.com</article-title>
          .publ.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Dragan</surname>
            <given-names>Gamberger</given-names>
          </string-name>
          , Rudjer Boskovic, Nada Lavrac, and
          <string-name>
            <given-names>Ciril</given-names>
            <surname>Groselj</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Experiments With Noise Filtering in a Medical Domain</article-title>
          .
          <source>In Proc. of 16 th ICML</source>
          . Morgan Kaufmann,
          <fpage>143</fpage>
          -
          <lpage>151</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Goldberger</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ehud</given-names>
            <surname>Ben-Reuven</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Training Deep Neural-networks Using a Noise Adaptation Layer</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long Short-Term Memory</article-title>
          .
          <source>Neural Comput. 9</source>
          ,
          <issue>8</issue>
          (Nov.
          <year>1997</year>
          ),
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          . https://doi.org/10.1162/neco.
          <year>1997</year>
          .
          <volume>9</volume>
          . 8.
          <fpage>1735</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Victoria</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hodge</surname>
            and
            <given-names>Jim</given-names>
          </string-name>
          <string-name>
            <surname>Austin</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>A survey of outlier detection methodologies</article-title>
          .
          <source>Artificial Intelligence Review</source>
          <volume>22</volume>
          (
          <year>2004</year>
          ),
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Thorsten</given-names>
            <surname>Joachims</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Text Categorization with Support Vector Machines: Learning with Many Relevant Features</article-title>
          .
          <source>In Proceedings of the 10th European Conference on Machine Learning (ECML'98)</source>
          . Springer-Verlag, Berlin, Heidelberg,
          <fpage>137</fpage>
          -
          <lpage>142</lpage>
          . https://doi.org/10.1007/BFb0026683
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Ji</surname>
            <given-names>Young</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            and
            <given-names>Franck</given-names>
          </string-name>
          <string-name>
            <surname>Dernoncourt</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Sequential Short-Text Classification with Recurrent and Convolutional Neural Networks</article-title>
          .
          <source>CoRR abs/1603</source>
          .03827 (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Joseph</surname>
            <given-names>Lilleberg</given-names>
          </string-name>
          , Yun Zhu,
          <string-name>
            <given-names>and Yanqing</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Support vector machines and Word2vec for text classification with semantic features.</article-title>
          .
          <string-name>
            <surname>In</surname>
            <given-names>ICCI</given-names>
          </string-name>
          *CC,
          <string-name>
            <surname>Ning</surname>
            <given-names>Ge</given-names>
          </string-name>
          , Jianhua Lu, Yingxu Wang, Newton Howard, Philip Chen, Xiaoming Tao, Bo Zhang, and Lotfi A. Zadeh (Eds.).
          <source>IEEE Computer Society</source>
          ,
          <fpage>136</fpage>
          -
          <lpage>140</lpage>
          . http: //dblp.uni-trier.de/db/conf/IEEEicci/IEEEicci2015.html#LillebergZZ15
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Emily</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Mallory</surname>
          </string-name>
          , Ce Zhang, Christopher Rï£¡, and
          <string-name>
            <surname>Russ</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Altman</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Largescale extraction of gene interactions from full-text literature using DeepDive</article-title>
          .
          <source>Bioinformatics</source>
          <volume>32</volume>
          ,
          <issue>1</issue>
          (
          <year>2016</year>
          ),
          <fpage>106</fpage>
          -
          <lpage>113</lpage>
          . https://doi.org/10.1093/bioinformatics/ btv476
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Christopher</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
            , Prabhakar Raghavan, and
            <given-names>Hinrich</given-names>
          </string-name>
          <string-name>
            <surname>Schütze</surname>
          </string-name>
          .
          <year>2008</year>
          . Introduction to Information Retrieval. Cambridge University Press, New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Tomas</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Kai Chen, Greg Corrado, and
          <string-name>
            <given-names>Jefrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Eficient Estimation of Word Representations in Vector Space</article-title>
          .
          <source>CoRR abs/1301</source>
          .3781 (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Alexander</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Ratner</surname>
          </string-name>
          ,
          <string-name>
            <surname>Christopher De</surname>
            <given-names>Sa</given-names>
          </string-name>
          , Sen Wu, Daniel Selsam, and
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Ré</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Data Programming: Creating Large Training Sets, Quickly</article-title>
          . In NIPS.
          <volume>3567</volume>
          -
          <fpage>3575</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Radim</given-names>
            <surname>Řehůřek</surname>
          </string-name>
          and
          <string-name>
            <given-names>Petr</given-names>
            <surname>Sojka</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Software Framework for Topic Modelling with Large Corpora</article-title>
          .
          <source>In Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks. ELRA</source>
          , Valletta, Malta,
          <fpage>45</fpage>
          -
          <lpage>50</lpage>
          . http://is.muni.cz/publication/ 884893/en.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Kai</given-names>
            <surname>Sheng</surname>
          </string-name>
          <string-name>
            <surname>Tai</surname>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Improved Semantic Representations From Tree-Structured Long Short-Term Memory Networks</article-title>
          .
          <source>CoRR abs/1503</source>
          .00075 (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Jonathan</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Webster</surname>
            and
            <given-names>Chunyu</given-names>
          </string-name>
          <string-name>
            <surname>Kit</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>Tokenization As the Initial Phase in NLP</article-title>
          .
          <source>In Proceedings of the 14th Conference on Computational Linguistics - Volume</source>
          <volume>4</volume>
          (
          <issue>COLING</issue>
          '92).
          <article-title>Association for Computational Linguistics</article-title>
          , Stroudsburg, PA, USA,
          <fpage>1106</fpage>
          -
          <lpage>1110</lpage>
          . https://doi.org/10.3115/992424.992434
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Peng</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Zhenyu Qi, Suncong Zheng, Jiaming Xu,
          <string-name>
            <given-names>Hongyun</given-names>
            <surname>Bao</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Bo</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Text Classification Improved by Integrating Bidirectional LSTM with Two-dimensional Max Pooling</article-title>
          .
          <source>CoRR abs/1611</source>
          .06639 (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>