<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>RARE : A Recurrent Attentive Recommendation Engine for News Aggregators</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Dhruv Khattar</institution>
          ,
          <addr-line>Vaibhav Kumar, Shashank Gupta, Manish Gupta</addr-line>
        </aff>
      </contrib-group>
      <issue>0</issue>
      <abstract>
        <p>With news stories coming from a variety of sources, it is crucial for news aggregators to present interesting articles to the user to maximize their engagement. This creates the need for a news recommendation system which understands the content of the articles as well as accounts for the users' preferences. Methods such as Collaborative Filtering, which are well known for general recommendations, are not suitable for news because of the short life span of articles and because of the large number of articles published each day. Apart from this, such methods do not harness the information present in the sequence in which the articles are read by the user and hence are unable to account for the speci c and generic interests of the user which may keep changing with time. In order to address these issues for news recommendation, we propose the Recurrent Attentive Recommendation Engine (RARE). RARE consists of two components and utilizes the distributed representations of news articles. The rst component is used to model the user's sequential behaviour of news reading in order to understand her general interests, i.e., to get a summary of her interests. The second component utilizes an article level attention mechanism to understand her speci c interests. We feed the information obtained from both the components to a Siamese Network in order to make predictions which pertain to</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Author had equal contribution. He can also be contacted
at vaibhav2@andrew.cmu.edu</p>
      <p>yAuthor is also a Principal Applied Researcher at Microsoft.
Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
BY 4.0).
the user's generic as well as speci c interests.
We carry out extensive experiments over three
real-world datasets and show that RARE
outperforms the state-of-the-art. Furthermore,
we also demonstrate the e ectiveness of our
method in handling the cold start cases.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>A news aggregator collects news from a variety of
sources and presents it to the user. It would be quite
cumbersome for a user to select articles of her choice
from a huge list of presented articles which may
pertain to a variety of subjects. Hence, it becomes crucial
for such aggregators to have a recommendation system
to point the user to the most relevant items and thus
maximize her engagement with the site and minimize
the time needed to nd relevant content.</p>
      <p>A popular approach to the task of recommendation
is collaborative ltering Bell and Koren (2007); Rennie
and Srebro (2005); Salakhutdinov et al. (2007), which
uses the user's past interaction with the item to
predict the most relevant content. Another common
approach is content-based recommendations, which uses
features between items and/or users to recommend
new items to the users based on the similarity between
features. However, amongst the various approaches
for collaborative ltering, Matrix Factorization (MF)
Koren (2008), is the most popular one, which projects
users and items into a shared latent space, using a
vector of latent features to represent a user or an item.
Thereafter, a user's interaction with an item is
modelled as the inner product of their latent vectors.</p>
      <p>However, Collaborative Filtering methods are not
suitable for news recommendation because news
articles have a short life span and expire quickly Zhong
et al. (2015). Such methods also require a
considerable number of interactions with an item (article)
before making predictions which is not desirable for
news recommendation because we would ideally want
to start recommending articles as soon as they are
published. Also, they do not directly harness the
information present in the sequence in which the
articles were read by the user and hence fail to account
for the generic as well as speci c interests of the user
which may keep changing with time. In order to
address these issues it becomes crucial to understand the
content of the news articles as well as the user's
preferences. We explain this through an example in the
following paragraph.</p>
      <p>As can be seen from Fig. 1(A), if a user reads four
di erent articles belonging to tennis and football, then
we would like our model to infer that the generic
interests of the users lie in reading articles about sports.
Hence, this would allow articles belonging to di erent
topics in the sports category to be recommended to
the user. However, since the user reads more articles
on tennis rather than football, we would like to give
more weight to the articles related to tennis as can be
seen in Fig. 1(B). Hence, in our overall list of
recommended articles to the user, we would like to present
news articles related to sports amongst which articles
related to tennis would be given more importance. It
may also happen that the user suddenly starts
reading articles related to business rather than sports. In
such a case we may also want to start recommending
articles related to business as well. This can be seen in
Fig. 1(C). However, it is important to note that in all
these cases the sequential reading history of the user
is very important while generating recommendations.</p>
      <p>To encode this intuition, we propose a novel neural
network framework namely Recurrent Attentive
Recommendation Engine (RARE). As illustrated in Fig. 3,
RARE consists of two components. The rst
component is based on a recurrent neural network and uses
the sequential reading history of the user as its
input. We call this the generic encoder. This helps us
to identify the generic/overall interests of the users,
i.e., it provides a summary of the user's interests. The
second component utilizes a recurrent neural network
with an attention mechanism to identify the speci c
interests of the user. We call this the speci c encoder.
The part dealing with attention allows the model to
attend to articles in a di erential manner, discriminating
the more from the less important ones. We then
concatenate the representations obtained from both these
components and call it the uni ed representation of the
users' interests. Limiting the size of the user reading
history used as inputs to both these components allows
us to adapt to the changing user preferences. We then
feed this uni ed representation along with the
representation of the candidate article to a Siamese
Network and compute an element wise product between
the outputs obtained at the nal layer of the sister
networks, as illustrated in Fig. 2. Finally, we use a
logistic unit to compute the score for recommendation.</p>
      <p>Using such a network enhances the model with further
non-linearity and enables it to capture the user-article
interaction in a better sense. It also allows the model
to learn an arbitrary similarity function instead of the
traditional metrics. The distributed representation of
each news article is used as input to our model. This
gives us the capability to recommend articles as and
when they are produced, without depending on any
prior user interaction with that article.</p>
      <p>To summarize, the main contributions of this work
are as follows.</p>
      <p>We present a neural network based architecture
(RARE) with the following capabilities.</p>
      <p>{ It utilizes the content of the news articles
giving it the ability to recommend articles
as soon as they are published.
{ It takes into account the users' generic as
well as speci c interests.
{ It adapts to the changing interests of the
user.</p>
      <p>We carry out extensive experiments over three
real world datasets to show the e ectiveness of
our model. The results reveal that our method
outperforms the state-of-the-art.</p>
      <p>We show the e ectiveness of our model for solving
the cold-start cases as well.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>There has been extensive study on recommendation
systems with a myriad of publications. In this section,
we aim at reviewing a representative set of approaches.
2.1</p>
      <sec id="sec-3-1">
        <title>Common Approaches for Recommendation Systems</title>
        <p>Recommendation systems in general can be
divided into collaborative recommendation systems and
content-based recommendation systems. In
collaborative ltering based recommendations, an item is
recommended to a user if similar users liked that item.
Collaborative ltering can be further divided into user
collaborative ltering, item collaborative ltering or
a hybrid of both user and item collaborative
ltering. Examples of such techniques include Bayesian
matrix factorization Salakhutdinov and Mnih (2008),
matrix completion Rennie and Srebro (2005), Restricted
Boltzmann Machine Salakhutdinov et al. (2007),
nearest neighbour modelling Bell and Koren (2007). In
user collaborative methods such as Bell and Koren
(2007), the algorithm rst computes similarity
between every pair of users based on the items liked by
them. Then, the scores of user-item pairs are
computed by combining scores of this item given by
similar users. Item-based collaborative ltering Sarwar
et al. (2001), computes similarity between items based
on the users who like both items. It then recommends
items to the user based on the items she has previously
liked. Finally, in user-item based collaborative
ltering, both the users and the items are projected into a
common vector space based on the user-item matrix
and then the item and user representation are
combined to nd a recommendation. Matrix factorization
based approaches like Rennie and Srebro (2005) and
Salakhutdinov and Mnih (2008) are examples of such
a technique. One of the major drawbacks of
collaborative ltering is its inability to handle new users and
new items, a problem which is often referred to as the
cold-start issue.</p>
        <p>Another common approach for recommendation is
content-based recommendation. In this approach,
features from user's pro le and/or item's description are
extracted and are used for recommending items to
users. The underlying assumption is that the users
tend to like items that they liked previously. In Liu
et al. (2010), each user is modeled by a distribution
over news topics that is constructed from articles she
liked with a prior distribution of topic preferences
computed using all users who share the same location. A
major advantage of using content-based
recommendation is that it can handle the problem of item cold-start
as it uses item features for recommendation. For user
cold-start, a variety of other features like age, location,
popularity aspects could be used. In the following we
discuss previous work on neural approaches for
recommendation systems.
Early work which used neural networks Salakhutdinov
et al. (2007) used a two-layer Restricted Boltzmann
Machine (RBM) to model users' explicit ratings on
items. The work has been later extended to model
the ordinal nature of ratings Phung et al. (2009).
Recently auto-encoders have become a popular choice for
building recommendation systems Chen et al. (2012);
Sedhain et al. (2015); Strub and Mary (2015). The
idea of user-based AutoRec Sedhain et al. (2015) is to
learn hidden structures that can reconstruct a user's
ratings given her historical ratings as inputs. In terms
of user personalization, this approach shares a similar
spirit as the item-item model Ning and Karypis (2011);
Sarwar et al. (2001) that represents a user in terms of
her rated item features. While previous work has lent
support for addressing collaborative ltering, most of
them have focused on observed ratings and modeled
the observed data only. As a result, they can easily
fail to learn users' preferences from the positive-only
implicit data.</p>
        <p>In Wu et al. (2016) a collaborative denoising
autoencoder (CDAE) for CF with implicit feedback is
presented. In contrast to the DAE-based CF Strub and
Mary (2015), CDAE additionally plugs a user node
to the input of auto-encoders for reconstructing the
user's ratings. As shown by the authors, CDAE is
equivalent to the SVD++ model Koren (2008) when
the identity function is applied to activate the hidden
layers of CDAE. Although CDAE is a collaborative
ltering model, it is solely based on item-item
interaction whereas the work which we present here is based
on user-item interaction. On the other hand in He
et al. (2017), authors have explored deep neural
networks for recommendation systems. They present a
general framework named NCF, short for Neural
Collaborative Filtering, that replaces the inner product
with a neural architecture that can learn an arbitrary
function from the given data. It uses a multi-layer
perceptron to learn the user-item interaction function.
NCF is able to express and generalize matrix
factoriza</p>
        <p>We then x a reading history size R, and use the
representations of the previous R articles read by
the user as inputs to the model.
3</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Model Architecture</title>
      <p>In this section we rst introduce the news article
recommendation task and then provide an elaborate
deWe come up with a uni ed representation of the
users' interests using recurrent neural networks
with an attention mechanism.</p>
      <p>Treating the uni ed representation of the user as a
query and the representation of the candidate
article as a document, we use a Siamese network to
make them undergo similar transformations and
supercharge them with non-linearities to discover
user-item interactions.
3.3</p>
      <sec id="sec-4-1">
        <title>Distributed Representation for News Articles</title>
        <p>We learn a 300-dimension distributed
representation Le and Mikolov (2014) for each news article by
combining the title and text of the news articles.
Learning such a representation allows us to
Capture the overall semantics of the news article.
Enables the model to come up with a
representation for new news articles as well as of articles
with varying lengths.</p>
        <p>News articles generally follow an inverted pyramid
structure where the title and the rst paragraph give
away the desired information. Hence, we only choose
the title and the rst paragraph because it usually
contains all the relevant information without delving into
detailed explanations. We also experimented by
choosing the entire news article but found better results with
just the rst paragraph.
3.4</p>
      </sec>
      <sec id="sec-4-2">
        <title>Generic Encoder</title>
        <p>The inputs for the generic encoder are the
representations of the articles previously read by the user.
Fig. 3(a) shows the graphical model of the network
used to identify generic interests in RARE. We use
Recurrent Neural Network (RNN) with Long-Short Term
Memory (LSTM) cells. LSTMs have been shown to be
capable of learning long-term dependencies Hochreiter
and Schmidhuber (1997); Sutskever et al. (2014). The
aim of this component is to understand the generic
(broader/overall) interests of the user. The last
hidden state of the RNN, i.e., ht encapsulates this
information, which we represent as cg. We can think of the
nal hidden state as the overall summary of the user's
interests.</p>
        <p>The state updates of the LSTM satisfy the following
equations.</p>
        <p>ft =
it =
ot =</p>
        <p>Wf ht 1; rt + bf
Wi ht 1; rt + bi</p>
        <p>Wo ht 1; rt + bo
lt = tanh V ht 1; rt + d
ct = ft ct 1 + it lt
ht = ot tanh(ct)
(1)
(2)
(3)
(4)
(5)
(6)</p>
        <p>Here is the logistic sigmoid function. ft, it, ot
represent the forget, input and output gates
respectively. rt denotes the input at time t and ht denotes
the latent state. Wf , Wi, Wo and V represent the
weight parameters respectively, while bf , bi, bo and d
represent the bias parameters respectively. The forget,
input and output gates control the ow of information
throughout the sequence.
3.5</p>
      </sec>
      <sec id="sec-4-3">
        <title>Speci c Encoder</title>
        <p>The architecture of Speci c Encoder is similar to that
of the Generic Encoder. The graphical representation
for this can be seen in Fig. 3(b). We use LSTM cells
here as well. To capture the speci c interests of the
users, i.e., to understand the deeper interests of the
user within her broader interests, we use an article
level attention mechanism. This provides us with a
context vector which encapsulates the speci c interests
of the user. This can be represented as,</p>
        <p>R
cs = X
j=1
j hj
(7)
where the attention weights, j , control the part of
the input sequence which should be emphasized or
ignored and hj stands for the output of the hidden units.
This attention mechanism gives RARE the capability
to adaptively focus more on the important items.
3.6</p>
      </sec>
      <sec id="sec-4-4">
        <title>RARE</title>
        <p>The complete architecture of the proposed model can
be seen in Fig. 4. The outputs obtained from the
speci c and the generic encoder are concatenated and
then used as inputs to a Siamese network along with
the candidate article.</p>
        <p>For the given task, the generic encoder captures the
overall interests of the user, i.e., it captures the
summary of the entire news articles read by the user. At
the same time, the speci c encoder adaptively selects
the important articles to capture the speci c interests
of the user. Hence to take advantage of both kinds
of information we concatenate the outputs of both the
encoders.</p>
        <p>As shown in Fig. 3, we can see that htg is
incorporated into cu to provide the summarized user
interests. Note that di erent encoding mechanisms will
be invoked in both the encoders when trained jointly.
The last hidden state of the generic encoder htg plays
a di erent role from that of hts. The former has the
responsibility to encode the information present in
the sequence in which the articles were read by the
user. While the latter is used for computing attention
weights. Information obtained from both the encoders
where cu represents the uni ed representation of users'
interests.</p>
        <p>We then use cu as inputs to one of the sister
networks in the Siamese network as shown in Fig. 3. The
input to the other sister network is the learned
representation of the candidate article. The Siamese
network supercharges RARE with further non-linearities
and makes the user representation and the article
representation go through similar transformations. In
Huang et al. (2013), an architecture similar to that of a
Siamese network has been used for ranking documents
with respect to a query with great e ectiveness. If
we try to draw a parallel between the query-document
problem with our task, one can see that a query in
our case is cu and the document is the representation
of the candidate news article. Hence, it seems apt to
use such a network if we were to project both of these
into the same geometric space to uncover the
underlying user-article interaction pattern. A similar sort of
technique has also been used by authors in He et al.
(2017) for modelling user-item interactions. Final
predictions are obtained from the Siamese network after
the logistic on the element-wise product between the
outputs obtained from the sister networks.</p>
        <p>Rather than using Siamese networks, the other
choice was to use a typical encoder-decoder framework.
However, a typical encoder-decoder framework is
unable to produce out-of-vocabulary (OOV) words. In
the news recommendation problem setting, each new
published article, that has not been interacted by any
user would act as an \OOV word". However, it is very
crucial for a news recommender to recommend articles
as soon they are published which is why we resort to
such a method as it allows us to handle such cases well.
Typically, to learn the model parameters, existing
point-wise methods Salakhutdinov and Mnih (2007)
perform regression with a squared loss. This is based
on the assumption that observations are generated
from a Gaussian distribution. However, in He et al.
(2017) it has been shown that such a method is not
very e ective when we have implicit data available.</p>
        <p>Given a user u and an article x, let y^ux represent
the predicted score at the output layer. Training is
performed by minimizing the point-wise loss between
y^ux and its target value yux. Considering the one-class
nature of implicit feedback, we can view the value of
yux as a label 1 meaning the item x is relevant to a
user u, and 0 otherwise. The prediction score y^ux then
represents how likely an item x is relevant to u. Hence
in order to constrain the values between 0 and 1, we
use the logistic function. We then de ne the likelihood
function as follows.</p>
        <p>p( +;
jI;
m) =</p>
        <p>Y
(u;i)2 +
y^ui</p>
        <p>Y
(u;j)2
(1
y^uj ) (9)
where + and represent the positive (observed
interactions) and negative (unobserved interactions)
articles respectively. I represents the input and m
represents the parameters of the model. The negative log
likelihood can then be written as follows (after
rearranging the terms).</p>
        <p>yui log y^ui +(1 yui)(1 log y^ui) (10)
L =</p>
        <p>X
u;i2 +[
The loss is similar to binary cross-entropy and can be
minimized using gradient descent methods.</p>
        <p>It is also worth noticing that the likelihood
function is such that it simultaneously adjusts the model's
parameters by maximizing the score of the relevant
articles and at the same time adjusts to minimize the
score of the non-relevant articles. This is similar to
0:8
0:6
0:2
0:6
In this section, we describe the datasets, the
state-ofthe-art methods, evaluation protocol along with the
settings used for learning the parameters of the model.
We use three real world datasets for evaluation. First,
we use the dataset published by CLEF NewsREEL
2017 Hopfgartner et al. (2016). CLEF shared a dataset
which captures interactions between users and news
stories. It includes interactions of eight di erent
publishing sites in the month of February 2016. The
recorded stream of events include 2 million noti
cations, 58 thousand item updates, and 168 million
recommendation requests. It also includes information
like the title and text of each news article. For this
dataset we considered all the users who had read more
than 10 articles after which we get a total of 22229
users. The other two datasets are provided by a
popular news aggregation website (name omitted for
review). The second dataset contains a list of articles
read by 10297 users in an Indian language,
Malayalam. The third dataset contains a list of articles read
by 22848 users in Indonesian. We make the code
publicly available 1.
4.2</p>
      </sec>
      <sec id="sec-4-5">
        <title>Baselines</title>
        <p>We compare our proposed approach with the following
methods.</p>
        <p>ItemPop. News articles are ranked by their
popularity judged by their number of interactions.
This is a non-personalized method to benchmark
the recommendation performance Rendle et al.
(2009).
1https://github.com/dhruvkhattar/RARE
eALS He et al. (2016). This is a state-of-the-art
matrix factorization method for item
recommendation. It optimizes the squared loss (between
actual item ratings and predicted ratings) and treats
all unobserved interactions as negative instances
and weighting them non-uniformly by item
popularity.</p>
        <p>NeuMF He et al. (2017). This is a
state-of-theart neural matrix factorization model. It treats
the problem of generating recommendations
using implicit feedback as a binary classi cation
problem. Consequently it uses the binary
crossentropy loss to optimize its model parameters.</p>
        <p>Our method is based on user-item interactions,
hence we mainly compare it with other user-item
models. We leave out the comparison with other models
like SLIM Ning and Karypis (2011) and CDAE Wu
et al. (2016) because these are item-item models and
hence performance di erence may be caused by the
user models for personalization.
4.3</p>
      </sec>
      <sec id="sec-4-6">
        <title>Evaluation Protocol</title>
        <p>To evaluate the performance of the recommended item
we use the leave-one-out evaluation strategy which has
been widely adopted in literature Bayer et al. (2017);
He et al. (2016); Rendle et al. (2009). For each user
we held-out her latest interaction as the test instance
and utilized the remaining data for training. Since it is
time consuming to rank all items for every user during
evaluation, we followed the popular strategy Elkahky
et al. (2015); Koren (2008) that randomly samples 100
items that the user has not interacted with, ranking
the test item among the 100 items. The performance of
a ranked list is judged by Hit Ratio (HR) and
Normalized Discounted Cumulative Gain (NDCG) He et al.
0:8
0:9
0:8
0:6
0:2
0:68
0:67
10@0:66
G
CD0:65
N
0:64
0:63
We use an Intel i7-6700 CPU @ 3.40GHz which has
a RAM of 32GB and a Tesla K40c GPU. We
implemented our proposed method using Keras
Chollet et al. (2015). We randomly divide the labeled
set into training and validation set in a 4:1 ratio.
We tuned the hyper-parameters of our model
using the validation set. The proposed model and all
its variants are learned by optimizing the log
likelihood given by Eq. 10. We initialize the fully
connected network weights with the uniform
distribution in the range between p6=(f anin + f anout) and
p6=(f anin + f anout) Glorot and Bengio (2010). We
used a batch size of 256 and used AdaDelta Zeiler
(2012) as the optimizer.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results and Analysis</title>
      <p>In this section we present the results obtained by
carrying di erent experiments with our method.
0:5
0:4
0:6</p>
      <p>0:4
LSTM
RNN</p>
      <p>GRU
0 1 2 3 4 5 6 7 8 9 10 11</p>
      <p>K
LSTM
RNN</p>
      <p>GRU
0 1 2 3 4 5 6 7 8 9 10 11</p>
      <p>K
For MF based methods like BPR and eALS, the
number of predictive factors chosen is equal to the number
of latent factors. We report the best performance in
this case. For NeuMF, we vary the size of the CF
layers (also latent factors) to choose the best t for our
model.</p>
      <p>In Figs. 5 to 6, we compare our method with the
baselines. Note that the performance of ItemPop
measure was very weak and hence it does not show up
clearly in the graphs. The Top-K recommended lists
are used where K varies from 1 to 10. It is very
clear from Figs. 5 and 7 that RARE outperforms other
methods by a signi cant margin across all positions on
the NewsREEL and the Malayalam datasets
respectively. Although, RARE outperforms the other
methods in case of Indonesian dataset as well (Fig. 6) but
the margin is not that large. Amongst the di erent
baselines, the trend in the performance can be seen
as follows: NeuMF &gt;eALS &gt;BPR (in terms of both
HR and NDCG). Although, in Rendle et al. (2009) it
has been shown that BPR can be a strong performer
for ranking performance owing to its pairwise
ranking aware learner, we did not see the trend for our
datasets. On the other hand RARE outperforms all
the other baselines in terms of NDCG as well.
We vary the size of the reading history R used as inputs
to our model. From Fig. 5, one can see that the Hit
Ratio slowly increases with the size of the reading history
until a certain point after which it decreases. However,
the NDCG keeps on increasing. We can attribute this
behaviour to the fact that users have diversi ed
reading interests which only get e ectively captured after a
substantial number of interactions have been observed.
However, after a while, increasing the user history
often leads to over-specialization where the generic
interests tend to overpower the speci c ones. This is also
an indicator of the fact that the preference of a user
keeps varying and hence a window size should be
chosen such that it helps the model to dynamically adapt
to the users changing behaviour.</p>
      <p>For all our methods, we chose a reading history of
12 for the users. We needed to make a choice between
12 and 14, and we chose 12 because we gave more
importance to the HR rather than the NDCG.
We rst note the e ects on RARE by varying the kind
of recurrent network used. We tested our model by
using LSTMs, GRUs (Gated Recurrent Units) Chung
et al. (2014) and Vanilla RNN. From Fig. 9, the trend
in the performance can be observed as follows: LSTM
&gt;GRU &gt;RNN although the di erences are not very
large. One of the reasons for this could be the fact
that an LSTM or a GRU is better able to encode the
interests of the user as they handle long-term
dependencies better.</p>
      <p>We also note the e ects when using di erent
variants of our own model, i.e., when we replace the uni ed
representation in RARE with solely the speci c or the
generic encoder. The results for this can be seen from
Table 1. We note the trend in performance as follows
RARE &gt;Generic Encoder &gt;Speci c Encoder. This
indicates that merely identifying the users' generic
interests (a summary of overall interests) is not su cient
for learning a good recommendation model. However,
when we use a combination of both in RARE, we nd
that the recommendation performance improves which
clearly indicates that identifying both the speci c and
generic interests are essential for better
recommendations.
5.4</p>
      <sec id="sec-5-1">
        <title>Performance on Cold Start Cases</title>
        <p>We then evaluated our model for the cold start cases
as can be seen in Fig. 10. For this task we segregated
users who had read a new news article in the end, i.e.,
they read articles which had never been seen before
they read it. We found out that the number of such
users were 74 in the CLEF dataset. There were very
0:6
0:3
0:1
0:2</p>
        <p>Cold News</p>
        <p>Cold User
0:10 1 2 3 4 5 6 7 8 9 10 11</p>
        <p>K
few such users in the other two datasets. Out of these
74 users, we see that the HR@10 is around 0.35. This
promises us that our model is well suitable for handling
the item cold-start problem.</p>
        <p>For user cold-start, we test our learned model over
users who had read articles in between 2 to 4
(inclusive) over the same dataset. Since we set the history
size to 12, we had to set the remaining inputs to 0s.
The HR@10 score was around 0.5. We see a gradual
increase in the hit rates as we increase the value of
K. The results promise the e ectiveness of our model
to handle the problem of user cold start as well.
Although this is not exactly the user cold start problem
because it still considers some number of user
interactions, still it is worth noticing the performance because
the baselines need a considerable amount of user
history before making predictions. On the other hand, in
our method, we can simply use the trained model for
recommending articles to users who have had very few
interactions.
We observe the performance of our model when we
vary the number of layers used in the Siamese Network
in our model. We experiment by varying the number
of layers along with the number of hidden units. We
experiment by using one layer with size 128, two layers
with sizes 128 and 64 and three layers with sizes 128,
64 and 32. From Table 2, we can see that the best
performance is observed in the second case.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we proposed the Recurrent Attentive
News Recommendation Engine (RARE) to address the
problem of news recommendation. We attempt to
encode both the generic and the speci c interests of the
users. For the former we use a recurrent neural
network while for the latter we use a recurrent network
with an attention mechanism. We use the uni ed
representations obtained from both these along with a
Siamese network to make predictions. We conducted
extensive experiments on three real-world datasets and
demonstrated that our method can outperform the
state-of-the-art methods in terms of di erent
evaluation metrics.</p>
      <p>Keras.</p>
      <p>Junyoung Chung, Caglar Gulcehre, KyungHyun Cho,
and Yoshua Bengio. 2014. Empirical Evaluation
of Gated Recurrent Neural Networks on Sequence
Modeling. arXiv preprint arXiv:1412.3555 (2014).
Ali Mamdouh Elkahky, Yang Song, and Xiaodong He.
2015. A Multi-View Deep Learning Approach for
Cross Domain User Modeling in Recommendation
Systems. In WWW. 278{288.</p>
      <p>Xavier Glorot and Yoshua Bengio. 2010.
Understanding the Di culty of Training Deep Feed-forward
Neural Networks. In AI-Stats, Vol. 9. 249{256.
Xiangnan He, Tao Chen, Min-Yen Kan, and Xiao
Chen. 2015. Trirank: Review-aware Explainable
Recommendation by Modeling Aspects. In CIKM.
1661{1670.</p>
      <p>Xiangnan He, Lizi Liao, Hanwang Zhang, Liqiang Nie,
Xia Hu, and Tat-Seng Chua. 2017. Neural
Collaborative Filtering. In Proceedings of the 26th
International Conference on World Wide Web (WWW
'17).</p>
      <p>Xiangnan He, Hanwang Zhang, Min-Yen Kan, and
Tat-Seng Chua. 2016. Fast matrix factorization for
online recommendation with implicit feedback. In
Proceedings of the 39th International ACM SIGIR
conference on Research and Development in
Information Retrieval. ACM, 549{558.</p>
      <p>Sepp Hochreiter and Jurgen Schmidhuber. 1997. Long
Short-Term Memory. Neural Comp. 9, 8 (1997),
1735{1780.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Immanuel</given-names>
            <surname>Bayer</surname>
          </string-name>
          , Xiangnan He,
          <string-name>
            <surname>Bhargav Kanagal</surname>
          </string-name>
          , and Ste en Rendle.
          <year>2017</year>
          .
          <article-title>A Generic Coordinate Descent Framework for Learning from Implicit Feedback</article-title>
          .
          <source>In WWW.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Robert M Bell</surname>
            and
            <given-names>Yehuda</given-names>
          </string-name>
          <string-name>
            <surname>Koren</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Improved Neighborhood-based Collaborative Filtering</article-title>
          .
          <source>In KDD</source>
          .
          <volume>7</volume>
          {
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Minmin</given-names>
            <surname>Chen</surname>
          </string-name>
          , Zhixiang Xu,
          <string-name>
            <given-names>Fei</given-names>
            <surname>Sha</surname>
          </string-name>
          , and
          <string-name>
            <surname>Kilian Q Weinberger</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Marginalized Denoising Autoencoders for Domain Adaptation</article-title>
          . In ICML.
          <volume>767</volume>
          {
          <fpage>774</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Francois</given-names>
            <surname>Chollet</surname>
          </string-name>
          et al.
          <year>2015</year>
          . https://github.com/fchollet/keras.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Frank</given-names>
            <surname>Hopfgartner</surname>
          </string-name>
          , Torben Brodt, Jonas Seiler, Benjamin Kille, Andreas Lommatzsch, Martha Larson, Roberto Turrin, and
          <string-name>
            <given-names>Andras</given-names>
            <surname>Sereny</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Benchmarking News Recommendations: The CLEF NewsREEL Use Case</article-title>
          .
          <source>In ACM SIGIR Forum</source>
          , Vol.
          <volume>49</volume>
          . 129{
          <fpage>136</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Po-Sen</surname>
            <given-names>Huang</given-names>
          </string-name>
          , Xiaodong He,
          <string-name>
            <surname>Jianfeng Gao</surname>
            , Li Deng,
            <given-names>Alex</given-names>
          </string-name>
          <string-name>
            <surname>Acero</surname>
            , and
            <given-names>Larry</given-names>
          </string-name>
          <string-name>
            <surname>Heck</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Learning Deep Structured Semantic Models for Web Search using Clickthrough Data</article-title>
          .
          <source>In CIKM</source>
          .
          <volume>2333</volume>
          {
          <fpage>2338</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Yehuda</given-names>
            <surname>Koren</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Factorization Meets the Neighborhood: A Multifaceted Collaborative Filtering Model</article-title>
          . In KDD.
          <volume>426</volume>
          {
          <fpage>434</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Quoc</given-names>
            <surname>Le</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Distributed Representations of Sentences and Documents</article-title>
          . In ICML.
          <volume>1188</volume>
          {
          <fpage>1196</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Jiahui</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Dolan</surname>
          </string-name>
          , and
          <string-name>
            <surname>Elin R nby Pedersen</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Personalized news recommendation based on click behavior</article-title>
          .
          <source>In Proceedings of the 15th international conference on Intelligent user interfaces. ACM</source>
          ,
          <volume>31</volume>
          {
          <fpage>40</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Xia</given-names>
            <surname>Ning</surname>
          </string-name>
          and
          <string-name>
            <given-names>George</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Slim: Sparse Linear Methods for Top-n Recommender Systems</article-title>
          . In ICDM.
          <volume>497</volume>
          {
          <fpage>506</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Dinh Q Phung</surname>
          </string-name>
          ,
          <string-name>
            <surname>Svetha Venkatesh</surname>
          </string-name>
          , et al.
          <year>2009</year>
          .
          <article-title>Ordinal Boltzmann Machines for Collaborative Filtering</article-title>
          .
          <source>In UAI</source>
          .
          <volume>548</volume>
          {
          <fpage>556</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Ste en Rendle</surname>
          </string-name>
          , Christoph Freudenthaler, Zeno Gantner, and
          <string-name>
            <surname>Lars</surname>
          </string-name>
          Schmidt-Thieme.
          <year>2009</year>
          .
          <article-title>BPR: Bayesian personalized ranking from implicit feedback</article-title>
          .
          <source>In Proceedings of the twenty- fth conference on uncertainty in arti cial intelligence</source>
          . AUAI Press,
          <volume>452</volume>
          {
          <fpage>461</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Jasson</surname>
            <given-names>DM</given-names>
          </string-name>
          <string-name>
            <surname>Rennie and Nathan Srebro</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Fast maximum margin matrix factorization for collaborative prediction</article-title>
          .
          <source>In Proceedings of the 22nd international conference on Machine learning. ACM</source>
          ,
          <volume>713</volume>
          {
          <fpage>719</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andriy</given-names>
            <surname>Mnih</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Probabilistic Matrix Factorization</article-title>
          .
          <source>In NIPS</source>
          , Vol.
          <volume>1</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andriy</given-names>
            <surname>Mnih</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Bayesian probabilistic matrix factorization using Markov chain Monte Carlo</article-title>
          .
          <source>In Proceedings of the 25th international conference on Machine learning. ACM</source>
          ,
          <volume>880</volume>
          {
          <fpage>887</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          , Andriy Mnih, and Geo rey Hinton.
          <year>2007</year>
          .
          <article-title>Restricted Boltzmann machines for collaborative ltering</article-title>
          .
          <source>In Proceedings of the 24th international conference on Machine learning. ACM</source>
          ,
          <volume>791</volume>
          {
          <fpage>798</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Badrul</given-names>
            <surname>Sarwar</surname>
          </string-name>
          , George Karypis, Joseph Konstan,
          <string-name>
            <given-names>and John</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Item-based collaborative ltering recommendation algorithms</article-title>
          .
          <source>In Proceedings of the 10th international conference on World Wide Web. ACM</source>
          ,
          <volume>285</volume>
          {
          <fpage>295</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Suvash</given-names>
            <surname>Sedhain</surname>
          </string-name>
          , Aditya Krishna Menon,
          <string-name>
            <given-names>Scott</given-names>
            <surname>Sanner</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Lexing</given-names>
            <surname>Xie</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Autorec: Autoencoders meet collaborative ltering</article-title>
          .
          <source>In Proceedings of the 24th International Conference on World Wide Web. ACM</source>
          ,
          <volume>111</volume>
          {
          <fpage>112</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Florian</given-names>
            <surname>Strub</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jeremie</given-names>
            <surname>Mary</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Collaborative Filtering with Stacked Denoising AutoEncoders and Sparse Inputs</article-title>
          .
          <source>In NIPS Workshop on ML for eCommerce.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Ilya</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , Oriol Vinyals, and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Sequence to Sequence Learning with Neural Networks</article-title>
          .
          <source>In NIPS</source>
          .
          <volume>3104</volume>
          {
          <fpage>3112</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Yao</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Christopher</surname>
            <given-names>DuBois</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alice X Zheng</surname>
            , and
            <given-names>Martin</given-names>
          </string-name>
          <string-name>
            <surname>Ester</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Collaborative Denoising AutoEncoders for Top-n Recommender Systems</article-title>
          . In WSDM.
          <volume>153</volume>
          {
          <fpage>162</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Matthew D</given-names>
            <surname>Zeiler</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>ADADELTA: An Adaptive Learning Rate Method</article-title>
          .
          <source>arXiv preprint arXiv:1212.5701</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Erheng</given-names>
            <surname>Zhong</surname>
          </string-name>
          , Nathan Liu, Yue Shi, and
          <string-name>
            <given-names>Suju</given-names>
            <surname>Rajan</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Building discriminative user pro les for large-scale content recommendation</article-title>
          .
          <source>In Proceedings of the 21th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM</source>
          ,
          <volume>2277</volume>
          {
          <fpage>2286</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>