<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recurrent Neural Networks for Customer Purchase Prediction on Twitter</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shigeyuki Sakaki Fuji Xerox Co.</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Japan sakaki.shigeyuki@ fujixerox.co.jp</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>CCS Concepts</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Francine Chen Yan-Ying Chen FX Palo Alto Laboratory, Inc.</institution>
          <addr-line>Palo Alto, California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mandy Korpusik Computer Science &amp; Artificial Intelligence Laboratory</institution>
          ,
          <addr-line>MIT Cambridge, Massachusetts</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The abundance of data posted to Twitter enables companies to extract useful information, such as Twitter users who are dissatis ed with a product. We endeavor to determine which Twitter users are potential customers for companies and would be receptive to product recommendations through the language they use in tweets after mentioning a product of interest. With Twitter's API, we collected tweets from users who tweeted about mobile devices or cameras. An expert annotator determined whether each tweet was relevant to customer purchase behavior and whether a user, based on their tweets, eventually bought the product. For the relevance task, among four models, a feed-forward neural network yielded the best cross-validation accuracy of over 80% per product. For customer purchase prediction of a product, we observed improved performance with the use of sequential input of tweets to recurrent models, with an LSTM model being best; we also observed the use of relevance predictions in our model to be more e ective with less powerful RNNs and on more di cult tasks.</p>
      </abstract>
      <kwd-group>
        <kwd>Deep learning</kwd>
        <kwd>Recommender systems</kwd>
        <kwd>Microblogs</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>In social media, popular aspects of customer relationship
management (CRM) include interacting with and
responding to individual customers and analyzing the data for trends
and business intelligence. Another mostly untapped aspect
is to predict which users will purchase a product, which is
useful for recommender systems when determining which
users will be receptive to speci c product recommendations.
Many people ask for input from their friends before buying
Permission to make digital or hard copies of part or all of this work for personal or
classroom use is granted without fee provided that copies are not made or distributed
for profit or commercial advantage and that copies bear this notice and the full citation
on the first page. Copyrights for third-party components of this work must be honored.
For all other uses, contact the owner/author(s).</p>
      <p>CBRecSys 2016, September 16, 2016, Boston, MA, USA.
c 2016 Copyright held by the owner/author(s).</p>
      <p>
        ACM ISBN 123-4567-24-567/08/06. . . $15.00
DOI: 10.475/123 4
Looking to buy a camera for X-mas, can I get some help?
Mann wish I had an iphone,I wanna be cooolll too! #sadtweet
Now thinking of getting a new phone
My baby arrived #panasonic #lumix - its waterproof but I
daren't try it hah :)
Got a #Xoom tablet today. Already rooted and master of my
domain. Thanks @koush
a higher-priced product, and users will often turn to social
media for that input [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Many users also post
announcements about signi cant purchases. Thus, social media posts
may contain cues that can be used to identify users who
are likely to purchase a product of interest as well as later
post(s) indicating that a user made a purchase.
      </p>
      <p>In this paper, we present our investigations of deep
learning methods for predicting likely buyers from microblogs,
speci cally Twitter tweets. While many social media posts
are private, semi-private or transient, and thus hard to
obtain, microblogs such as tweets are generally public,
simplifying their collection. With the ability to identify likely
buyers, as opposed to targeting anyone who mentions a product,
advertisements and products can be presented to a more
receptive set of users while annoying fewer users with spam.</p>
      <p>Microblogs cover a variety of genres, including
informative, topical, emotional, or \chatter." Many of a user's tweets
are not relevant, that is, indicative of whether a user is likely
to purchase a product. Thus, we hypothesize that
identifying tweets that are relevant to purchasing a product is useful
when predicting whether a user will purchase that product.
We also hypothesize that a sequence of tweets contains
embedded information that can lead to more robust prediction
than classi cation of individual tweets. For example, with
the given order of the second and third tweets in Table 1,
the user seems interested in buying a new phone, but if the
order is reversed, the inference is that the user was thinking
of buying a phone, but did not buy one. In this work, we
investigate the use of recurrent neural networks to model
sequential information in a user's tweets for purchase
behavior prediction. Our use of recurrent models
enables previous tweets to serve as context.
introduce relevance prediction into the model for
reducing the in uence from noisy tweets.</p>
      <p>This paper is organized as follows. In the next section,
we describe related work in deep learning and social media
product interest prediction. Afterward, we detail our data
collection and annotation process. Then, we explain the
deep learning methods, discuss the experiments, and analyze
results. Finally, we conclude and propose future work.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        There are several works related to identifying customers
with interest in a product. A system developed by [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for
predicting user interest domains was based on features
derived from Facebook pro les to predict from which category
of products an eBay user is likely to make a purchase. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
used a rule-based approach for identifying sentences
indicating \buy wishes" from forums with buy-sell sections. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
presented a method for predicting whether a single post
on Quora or Yahoo! Answers contained an expression of
\purchase intent." However, some of their Purchase Action
words, including \want" and \wish" only indicate interest in
a product. Although users may say they want a product,
many of these Twitter users will not buy the product in the
near future (see Figure 2 and Table 3). For our task, we
predict whether a user will actually make a purchase; this is
di erent from the task of predicting user interest in a
product. Often users with interest in a product may not have
the means to purchase soon or may say they want or need
something as an indication that they like something.
      </p>
      <p>Our approach also di ers from these works in that access
to Facebook pro les is not needed, and features are
automatically learned using neural networks. In addition, most
of the earlier works classify a single sentence or posting. In
contrast, we predict user behavior based on past postings,
that is, from a sequence of tweets, which enables preceding
tweets in a sequence to provide context for the current tweet.</p>
      <p>
        Our approach draws on the deep learning Recurrent
Neural Network (RNN) models which have been successfully
applied to sequential data such as sentences or speech. For
example, [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] used a combination of a convolutional
neural network to represent sentences and an RNN to model
sentences in a discourse. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] observed that long short-term
memory (LSTM) models, a type of RNN with longer
memory, performed better than an RNN on the TIMIT phoneme
recognition task. In our work, we compare RNN and LSTM
models for the task of predicting whether a user will buy a
product of the type that they have mentioned.
      </p>
    </sec>
    <sec id="sec-3">
      <title>DATA COLLECTION AND LABELING</title>
      <p>For the buy prediction task, we focused on two product
categories: (1) cameras and (2) mobile devices, i.e.,
mobile phones, tablets, and smart watches. These are
generally higher-priced products which users do not purchase
frequently and therefore are more likely to tweet about.</p>
      <p>
        We created a separate corpus for each category composed
of tweets by users who either: (1) bought a target product
or (2) wanted, but did not buy, a target product [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. To
collect tweets by each user, we rst identi ed from eBay
listings a set of model names for each product category. Similar
model names were merged, e.g., \iPhone4" and \iPhone5"
into \iPhone," resulting in 146 camera names and 80 mobile
device names. We also created a set of regular expressions
that may indicate a user bought or wanted one of the
products (see sample in Table 2). Tweets containing a bought
or want expression for one of the product names were then
collected using the Twitter search API, and the user of each
tweet was identi ed from the tweet meta-data. The tweets of
the identi ed users were collected using the Twitter search
and timeline APIs. We called users found with \bought"
regular expressions candidate buy users, and users identi ed
from \want" regular expressions candidate want users.
      </p>
      <p>Due to poor labeling performance by Mechanical Turkers,
who often were not familiar with many of the lesser-known
mobile devices and cameras, we used an \expert" annotator
to whom we gave many examples of labeled bought and want
tweets, including trickier cases. For example, a user did not
buy a camera if they were given one or if they retweeted (RT)
a user who bought a camera. A sample of the labels was
checked for accuracy by one of the authors. The annotator
determined whether each candidate want user tweeted that
they bought the product type of interest; if so, the candidate
want user was labeled a buy user (Table 3). Similarly, the
tweets of each candidate buy user were examined for at least
one tweet indicating that the user really bought the target
product type. In total, we annotated tweets from 2,403
mobile device users and 1,252 camera users. The annotator also
labeled a separate random sample of tweets as relevant/not
to predicting whether a target product was bought.
4.</p>
    </sec>
    <sec id="sec-4">
      <title>DEEP LEARNING METHODS</title>
      <p>
        In this section, we describe the neural network (NN)
models we implemented in Python's Theano toolkit [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for
classifying tweets as relevant/not and predicting whether a
Twitter user bought a product 60 days after tweeting about it.
      </p>
      <p>The logistic regression (LR) model combines the input
with a weight matrix and bias vector, feeding it through a
softmax classi cation layer that yields probabilities for each
class i. The class i with the highest probability is the output.</p>
      <p>A feed-forward (FF) network enables more complex
functions to be computed through the addition of a sigmoid
hidden layer below the softmax. A natural extension of the FF
network for sequences is a recurrent neural network (RNN),
in which the hidden layer from the previous timestep is fed
into the current timestep's hidden layer:
ht = (Wxxt + Whht 1 + b)
(1)
where ht is the hidden state, xt is the input vector, W is a
learned weight matrix, and b is a learned bias vector.</p>
      <p>
        Thus, information from early words/tweets is preserved
across time and is still accessible upon reaching the nal
word/tweet for making a prediction. However, typically the
error gradient vanishes as the sequence becomes increasingly
long, which in practice causes information loss over long time
spans. To compensate for this, long short-term memory [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
uses input, output, and forget gates to control what
information is stored or forgotten within a memory cell over longer
periods of time than a standard RNN. In the LSTM, the
input gate it, forget gate ft, and candidate memory cell Cft
are computed using input xt and previous hidden layer ht 1,
weight matrices (i.e., Wi and Ui for the input gate, Wf and
Uf for the forget gate, and Wc and Uc for candidate Cft),
and forget gate bias term bf as follows:
(2)
(3)
(4)
(5)
(6)
(7)
it = (Wixt + Uiht 1)
ft = (Wf xt + Uf ht 1 + bf )
      </p>
      <p>Cft = tanh(Wcxt + Ucht 1)
The new memory cell Ct is generated from the combination
of the candidate memory cell Cft, controlled by the input
gate it through element-wise multiplication, and the
previous memory cell Ct 1, modulated by the forget gate ft:</p>
      <p>Ct = it Cft + ft Ct 1
The output gate and new hidden layer are computed by:
ot = (Woxt + Uoht 1 + VoCt)</p>
      <p>ht = ot tanh Ct</p>
      <p>
        To preprocess the data, each tweet was rst tokenized
by the TweeboParser [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a dependency parser trained on
tweets. Then each token was converted into a vector
representation using word2vec [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which learns an embedded
vector representation from the weights of a neural network
trained on Google News. We also experimented with the
GloVe word embedding [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and with learning an embedding
from random initialization. In preliminary experiments on
predicting tweet relevance, word2vec consistently performed
best and so was used for all reported results.
4.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>Models for Predicting Tweet Relevance</title>
      <p>To predict tweet relevance, we compared the LR, FF,
RNN, and LSTM models. Since the RNN and LSTM are
sequential, the input was an array of token embedded vectors;
for the LR model and FF networks, we summed the token
embedded vectors as input. For regularization of the LR and
FF models, we employed an early stopping technique, and
for the RNN and LSTM networks, we incorporated dropout.
4.2</p>
    </sec>
    <sec id="sec-6">
      <title>Models for Predicting Purchase Behavior</title>
      <p>To predict whether a user will buy a product based on
their tweets, we propose a con guration of neural networks
that uses predicted tweet relevance in purchase prediction.
The input for each user is a sequence of tweets (instead of
words, as is more commonly used) enabling the preceding
tweets to provide context for the current tweet. To model
the information in a tweet sequence, a recurrent network
(e.g., RNN or LSTM) is intuitively a good choice.</p>
      <p>In our proposed joint model (Figure 1), tweets from a
user are input as a sequence where each tweet is represented
as the sum of the embedded vectors representing its words.
The (optional) lower sub-network predicts the relevance of
each tweet; we use the best of the four types of relevance
classi cation models, the feed-forward neural net.</p>
      <p>The buy sub-network at each time step (i.e., tweet) is
composed of either an RNN or an LSTM memory cell, which
is fed a tweet vector along with its predicted relevance. The
maximum across each dimension of all the tweets' hidden
layer outputs are fed into the softmax classi er for the nal
buy/not buy prediction. The softmax will generalize well to
future work on predicting other labels besides buy/not-buy.</p>
      <p>For all experiments, each data set was split into 10
partitions for 10-fold cross-validation. Within each fold, 10%
was for validation, 10% for testing, and the rest for training.
In order to incorporate the FF classi er for predicting tweet
relevance as a sub-network in a joint network, we trained a
separate classi er for each of the 10 cross-validation
partitions. We selected all the users in the training and validation
sets for that partition, and trained the relevance classi er on
all those users' tweets for which we had a relevance label.
The predicted relevance for each tweet was used as an
additional feature to predict the user's purchase behavior.</p>
    </sec>
    <sec id="sec-7">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>We conducted two sets of experiments: rst, predicting
whether a tweet is relevant to a user's purchase behavior;
second, predicting whether a Twitter user will eventually
purchase a product within 60 days of tweeting about it.</p>
      <p>
        For all the networks, we set the hidden layer size to 50.
The weight matrices were initialized randomly, and bias
vectors were initialized to zero. We used RMSprop [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and
negative log-likelihood for training, sigmoid nonlinear
activations, and batches of size 10 for up to 100 epochs.
5.1
      </p>
    </sec>
    <sec id="sec-8">
      <title>Tweet Relevance</title>
      <p>We observe from the results shown in Table 4 that the
FF model performed best. The task is harder for combined
mobile device and camera data than for the two individual
products, likely due to di erences between domain-speci c
relevance indicators for cameras versus mobile devices.</p>
      <p>Model
Logistic</p>
      <p>FF
RNN (25%)
LSTM (50%)</p>
      <p>Mobile</p>
      <p>Camera
79.7
81.2
80.1
80.2
78.8
80.4
79.2
77.0</p>
      <p>Both
74.7
78.0
77.7
77.0</p>
      <p>We explored RNNs and LSTMs for predicting whether or
not a Twitter user would buy a product, since these models
sequentially scan through a user's tweets. We incorporated
the FF tweet relevance prediction model as a sub-network in
a deep network; this model predicts a feature indicating each
tweet's relevance, which is appended to the input tweets.</p>
      <p>For each Twitter user, we used all tweets containing a
product mention within a 60-day span, limited to users who
wrote between ve and 100 product-related tweets; the
upper limit was used to lter out advertisers. The 60 days
are motivated by the assumption that companies are not
interested in promoting products for longer periods, and by
Figure 2, which shows that about half the users purchase
what they want within 60 days.</p>
      <p>We evaluated our models on mobile devices, cameras, and
the two combined. Negative examples included Twitter users
who wanted a product but did not mention buying it. We
trained on their tweets from within the 60-day window
before their most recent tweet that mentions wanting a
product. Positive examples included users who eventually tweeted
about buying a product (the buy users), but did not include
the \bought" tweet or any tweets written afterward.</p>
      <p>The 10-fold cross-validation results averaged over several
runs (due to random initialization) on the expert-labeled
training data are shown in Table 5. As the baseline, we
trained a FF model with sums of all tokens across all tweets
for each user (*-sum) as the input. We observe better
performance when tweet information is input sequentially to a
model with memory (*-seq) than when the sequence
information is lost by summing (*-sum) tweets. That is, it is
important to capture information embedded in the sequence of
tweets. As expected, the LSTM consistently outperformed
the RNN because it has the ability to retain information over
longer time spans. We also observe that the addition of
predicted relevance probabilities to the RNN model (*+Rel)
improved performance over the simple RNN; however, adding
tweets' relevance only improved the LSTM's performance
for the harder combined product task, which may indicate
that the vanilla LSTM is powerful enough to learn which
tweets are relevant or not when trained on a single product.</p>
      <p>Model
FF-sum</p>
      <p>RNN-seq
RNN-seq+Rel</p>
      <p>LSTM-seq
LSTM-seq+Rel</p>
      <p>Mobile
73.6
80.8
81.3
83.9
83.8</p>
      <p>Camera
66.3
78.0
80.5
81.4
80.9</p>
      <p>Both
73.4
79.3
80.1
81.5
81.7</p>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>In this work, we investigated deep learning techniques for
predicting customer purchase behavior from Twitter data
that recommender systems could leverage. We collected a
labeled corpus of buy/not buy users and their tweets.</p>
      <p>A FF neural network performed best at predicting whether
a tweet is relevant to purchase behavior, with an accuracy
of 81:2% on mobiles devices and 80:4% on cameras. We
found that the use of a deep learning model that
incorporates sequential information performed better than ignoring
sequential information for the purchase prediction task.</p>
      <p>Our initial work in this area has many possible extensions.
While we used a 60-day window, it would be interesting to
observe user purchase probability changes over time. We will
also predict related behaviors, e.g., product comparison.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Bastien</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lamblin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pascanu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bergstra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. J.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bergeron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bouchard</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          . Theano:
          <article-title>New features and speed improvements</article-title>
          .
          <source>Deep Learning and Unsupervised Feature Learning NIPS 2012 Workshop</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          , A.-r. Mohamed, and
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <article-title>Speech recognition with deep recurrent neural networks</article-title>
          .
          <source>In Proc. of ICASSP</source>
          , pages
          <volume>6645</volume>
          {
          <fpage>6649</fpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Varshney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jhamtani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kedia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Karwa</surname>
          </string-name>
          .
          <article-title>Identifying purchase intent from social posts</article-title>
          .
          <source>In Proc. of ICWSM. AAAI</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural Computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kalchbrenner</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Blunsom</surname>
          </string-name>
          .
          <article-title>Recurrent convolutional neural networks for discourse compositionality</article-title>
          .
          <source>In Proc. of the 2013 Workshop on Continuous Vector Space Models and their Compositionality. ACL</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Swayamdipta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhatia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Smith.</surname>
          </string-name>
          <article-title>A dependency parser for tweets</article-title>
          .
          <source>In Proc. of EMNLP. ACL</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In Proc. of NIPS</source>
          , pages
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Manning</surname>
          </string-name>
          . GloVe:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In Proc. of EMNLP. ACL</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ramanand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bhavsar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Pedanekar</surname>
          </string-name>
          .
          <article-title>Wishful thinking nding suggestions and `buy' wishes from product reviews</article-title>
          .
          <source>In Workshop on Computational Approaches to Analysis and Generation of Emotion in Text, page 54</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sakaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Korpusik</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Corpus for customer purchase behavior prediction in social media</article-title>
          .
          <source>LREC</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>N.</given-names>
            <surname>Srivastava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <article-title>Dropout: A simple way to prevent neural networks from over tting</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <year>1929</year>
          {
          <year>1958</year>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tieleman</surname>
          </string-name>
          and G.
          <source>Hinton. Lecture 6</source>
          .5
          <article-title>-rmsprop: Divide the gradient by a running average of its recent magnitude</article-title>
          .
          <source>COURSERA: Neural Networks for Machine Learning</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          .
          <article-title>Social media peer communication and impacts on purchase intentions: A consumer socialization framework</article-title>
          .
          <source>Journal of Interactive Marketing</source>
          ,
          <volume>26</volume>
          (
          <issue>4</issue>
          ):
          <volume>198</volume>
          {
          <fpage>208</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Pennacchiotti</surname>
          </string-name>
          .
          <article-title>Predicting purchase behaviors from social media</article-title>
          .
          <source>In Proc. of WWW</source>
          , pages
          <volume>1521</volume>
          {
          <fpage>1532</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>