<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Real-time Integration of Social Media Background in Dynamic Recommendation Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yihong Zhang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiu Susie Fang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Takahiro Hara</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Donghua University</institution>
          ,
          <addr-line>Shanghai</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Osaka University</institution>
          ,
          <addr-line>Osaka</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recommendation systems provide personalized product and service recommendations by learning latent user preferences. Most recommendation systems nowadays are static, and do not consider real-time factors when making recommendations. However, purchasing behaviors are easily influenced by realtime events happening in society. Such real-time events can be extracted from social media, as previous works have shown. In this paper, we propose using social media as a background information source to improve e-commerce recommendation models. In contrast to previous works that created shallow representations of social media, we propose two representations of real-time social media information, that captures the dynamics of word usage trends and evolving semantic word relations. Taking a popular neural recommendation system as the base system, we show that the attention mechanism allows us to integrate the rich, matrix-like representation of social media. We conduct experimental evaluations on a real-world e-commerce dataset and a Twitter dataset. The results show that our method of social media background representation and integration is efective in integrating social media predictiveness in recommendation models, and the representation is superior compared to several other representations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;recommendation system</kwd>
        <kwd>social media</kwd>
        <kwd>user behavior modeling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Using a recommendation system to provide personalized service has become a popular practice
in e-commerce and online shopping platforms [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Typically, the goal of a recommendation
system is to discover latent user preferences from data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In such a system, the data useful
for discovering user preferences include implicit feedback and explicit feedback, which are
usually user purchase records and user ratings, respectively [
        <xref ref-type="bibr" rid="ref1 ref3 ref4">1, 3, 4</xref>
        ]. These data are also called
interaction data, as they are generated from user-item interaction. Recently, it has been found
that interaction data-based recommendation systems have some inherent weaknesses. The first
is so-called the cold-start problem. Given some users and items in an e-commerce platform,
sometimes there is no past record of user-item interaction, because the user or the item is a new
one in the system. To make recommendations in such cases, using information other than
useritem interaction is necessary [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. A group of such information is called contextual information
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Contextual information that has been shown useful in cold-start recommendation includes
user demographic data [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], item attributes [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and item review texts [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        The second weakness is related to temporal context awareness. In the traditional interaction
data, there are records of users purchasing or rating items, but these data are not timed.
Intuitively, one might think that user preferences can be influenced by real-time events, and data
closer to the time of recommendation may better indicate the user’s current preference [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In
order to make time-aware recommendations, we need both timed data and a modification to
traditional recommendation models. Previously, we have shown that a neural recommendation
system can be modified to incorporate time [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. As the next step, we need to find temporal
context. Social media can be considered a universal real-time information source. Social media
updates repeatedly according to real-time events happening around the world, on macro and
micro scales. Therefore we propose an extension to dynamic recommendation systems based
on social media.
      </p>
      <p>
        Some existing works have proposed to use social media to improve recommendation systems.
However, these works rely on the assumption that common users exist in social media and the
recommendation domain and can be identified [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ]. This is a strong assumption as many
e-commerce platforms do not have user social media data. In contrast to these works, the social
media in our study is considered a temporal background that reflects the general social interests
of the moment. By analyzing temporal patterns, we have found some associations between
social media discussions and purchase behavior. For example, when some local natural disaster
happened and received attention on social media, people’s interest for disaster prevention would
temporarily increase, and disaster prevention products on an e-commerce site would become
temporarily fast-selling. With such associations, what was discussed in social media can have
an impact on e-commerce user preferences even though users were not linked across platforms.
      </p>
      <p>
        While some interesting cases can be observed, Given the large number of words used on
Twitter, it is hard to handpick which word correlates to which e-commerce product. There
will be a lot of noise. A found correlation may be fake (false positive) and real related words
may not be found (false negative). Instead of using heuristics to discover positive cases, we
propose a more general approach. We convert social media into matrix-like representation and
fuse them with a neural network recommendation system through the well-known attention
mechanism [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ]. In this way, our method allows the system to automatically select relevant
real-time context without explicitly specifying the association between the context and the item.
In the preliminary study, we proposed a method that captures changing trends of words [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
While shown to be efective, this method considers words as independent information unit,
and overlooks their inter-relationships. Thus in this paper, we propose a new representation to
capture semantic word relations that may also change in real-time. The method is based on
graph convolutional network, the state-of-the-art structural information representation [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
More specifically, we extend the Evolve-GCN method [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] with real-time capabilities.
      </p>
      <p>
        Our work focuses on an e-commerce dataset collected by a specific platform, provided by
our industry partner1. However, we argue that our proposed method has generality that can
be adopted and applied to other e-commerce platforms. First, social media platforms such
as Twitter make data access available to the general public, and data can be easily collected.
Second, nowadays more and more recommendation systems use a deep neural network as the
recommendation model [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], and our method that fuses temporal context information with a
1Due to our agreement with the industry partner, we cannot make the dataset publicly available.
specific model can be easily applied to other models. Our contribution with this paper can be
summarized as the following:
• We generalize the problem of using social media background to support dynamic
recommendation systems. This is a rare study that use social media to support recommendation
without common and identifiable users.
• In addition to the previously proposed frequency-based representation, we propose a
graph-based representation for social media. We extend Evolve-GCN so that the model
can be used to provide real-time representations.
• We evaluate our method extensively using real-world datasets. The evaluation results
show that both the frequency-based representation and the graph-based representation,
and the combination of them, have positive impacts on the recommendation performance.
      </p>
      <p>They are also shown to be superior to other representations.</p>
      <p>The remainder of this paper is organized as the following. In Section 2, we will discuss related
work. Section 3 will introduce the problem and the base solution. Then in Section 4, we will
present our method for representing social media background and use attention to integrating
it into the base solution. Section 5 will present our experimental evaluations, including an
analysis of the dataset and experiment setups, followed by result discussions. Finally Section 6
will conclude this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        A number of research eforts have been made to address the problem of cold-start
recommendation using contextual information [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We are most interested in the temporal context, which is
close to our work. Cebrian et al. proposed a music recommendation system that used time in
the day (morning, afternoon, evening) as the temporal context [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Similarly Dias and Fonseca
proposed a music recommendation system that considered time in the day, weekday, day of the
month, etc., as well as session information [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. They also clustered songs into latent topics
by treating sessions as documents to further improve recommendations. Xiao et al. proposed
a probabilistic matrix factorization technique that considers day of the week as the temporal
context [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Beutel et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] proposed another work that incorporates temporal context to
recurrent neural networks. In these works, though, the temporal context is only discrete values
of time of the day, and there is no other information associated with such times.
      </p>
      <p>
        In addition, some works attempted to use social media as the context in recommendations.
For example, Alahmadi and Zeng proposed using linked Twitter accounts to address cold-start
recommendation [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. In order to connect social media to purchase behavior, they explicitly
asked e-commerce users to provide their Twitter accounts. Gao et al. studied the problem of
location recommendation with location-based social networks [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. They modeled temporal
check-in preferences from users’ past check-in records. In their case, both the contextual
information and the recommendation target were on the same platform, thus explicit user
links were available. Yang et al. proposed a method to predict sudden raise of product sales
by studying social media user interest difusion [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Their method was based on product text
snippets that can link social media text to products. In this paper, however, we consider social
media purely as a background. We assume no explicit link is available for social media and the
e-commerce platform, either through user or item. This makes our problem harder, but also
increases the generality of our solution. A work with a similar goal as ours was proposed by
Deng et al., who use Twitter as a background to recommend YouTube video clips [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Their
solution is feasible because many Twitter posts and YouTube video clips are generated by the
same news event. However, in more common scenarios, such as Twitter and e-commerce, this
correlation is dificult to establish.
      </p>
      <p>
        Outside of the research field of recommendation systems, social media as a background has
been used in diferent kinds of data analysis and applications. For example, Wei et al. found
that Twitter volume spikes could be used to predict stock options pricing [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. They used the
tweets that contained the stock symbols. Asur and Huberman studied if social media chatter can
be used to predict movie sales [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. They conducted sentiment analysis on tweets containing
movie names, and found some positive correlations. Pai and Liu proposed to use tweets and
stock market values to predict vehicle sales [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. They found that by adding the sentiment score
calculated from the tweets, prediction model performance substantially increased. Broniatowski
et al. made an attempt to track influenza with tweets [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. They combined Google Flue Trend
with tweets to track municipal-level influenza. Tweets were put through three classifiers to
isolate health-related, influenza-related, and case-reporting tweets, and finally the count of
relevant tweets was added to the prediction model. These works, however, only used high-level
features of social media, such as message counts or aggregated sentiment scores. In contrast,
our proposed solution captures richer information from social media, while also keeping them
machine-readable.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Preliminaries</title>
      <p>In this section, we will first introduce the problem of using social media background to support
dynamic neural recommenders. Next we will briefly introduce an existing dynamic neural
recommendation system, in which we will integrate social media information.</p>
      <sec id="sec-3-1">
        <title>3.1. Problem Formulation</title>
        <p>Technically, the problem a dynamic neural recommendation system tries to solve is to rank
candidate items based on user preferences at a certain time. Given a product description (), a user
description (), and the time , the system should make a prediction  =  ((), (), ),
where  is a score indicating the strength of preference, and  is the recommendation model.
This model is normally learned through supervised learning. In the implicit feedback
recommendations, the training dataset normally contains a number of positive triples, ((), (), ) = 1,
if user  has purchased product  at time , and a number of randomly sampled negative triples
((), (), ) = 0, from all triples where user  have not purchased product  at time .</p>
        <p>To use social media background to support the system is to add the temporal background 
as an extra input to the model, so that the prediction becomes  =  ((), (), , ). We
assume preprocessing has been done on social media texts and  can be a vector or a matrix,
depending on the representation method.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Base Dynamic Recommendation System</title>
        <p>
          There is a large number of proposals for learning the recommendation model  . We select the
recommendation system proposed by Wang et al. [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] because their system is a context-based
cold-start recommendation system that takes user and item embeddings as the input, similar to
our scenario.
        </p>
        <p>
          Assuming for each item, there is no user-item interaction past records available, and also
assuming from the contextual data, vector representations have been learned for users and
items, which are treated as () and (). The task of the cold start recommendation model
is thus to learn preference relationships between users and items based on their embeddings.
Wang et al. generalize a neural matrix factorization (NeuMF) model [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] that excludes the part
of learning embeddings with latent vectors. Their model is shown in Figure 1.
        </p>
        <p>ŷui
Fuly-connected</p>
        <p>Concatenation
GMF</p>
        <p>MLP
GMFItem
Embedding</p>
        <p>GMFUser</p>
        <p>Embedding
MLP ItemEmbedding</p>
        <p>MLP UserEmbedding
ItemEmbedding</p>
        <p>UserEmbedding</p>
        <p>This model is an ensemble of generalized matrix factorization (GMF) and multi-layer
perceptron (MLP). Two copies of user embeddings and item embeddings are input into the GMF
and MLP components, both of which produce an output embedding in their last layer. NeuMF
concatenates the two output embeddings and runs them through a fully-connected layer to
produce a prediction. The functions in this process are defined as the following:
z = p ⊙
q,
︂[ p ]︂
q</p>
        <p>z = ( (− 1(...2(2</p>
        <p>+ 2)...)) + ),
ˆ =  (h · [︂ z ︂] ),
z
where p and p denote user embeddings for GMF and MLP, while q and q  denote item
embeddings for the two components. ˆ denotes prediction results.</p>
        <p>Since the dataset contains only observed interactions, i.e., user purchase records of items,
when training the model, it is necessary to bring up some negative samples, for example, by
randomly choosing some user-item pairs that have no interaction. They defined the loss function
as the following:
 =</p>
        <p>∑︁
(,)∈∪−
 log ˆ + (1 − ) log(1 − ˆ),
(1)
(2)
(3)
(4)
Social Media</p>
        <p>Frequency change
analysis
Word co-occurrence
graph construction</p>
        <p>TrendRepresentation</p>
        <p>Attention</p>
        <p>Attention
SemanticRepresentation</p>
        <p>TransformedItem</p>
        <p>Embedding
GMFItem
Embedding
Ful y-connected
Concatenation</p>
        <p>GMF</p>
        <p>MLP
ItemEmbedding
where  = 1 if user  purchased item , and 0 otherwise.  denotes observed interactions
and − denotes negative samples.</p>
        <p>Although it is possible to make recommendations at any moment when a purchase intention
is detected, we follow a more realistic scenario by changing the recommendation three times
a day, i.e., in the morning, afternoon, and evening, which correspond to hour 10, 16, and 22
of the day. Since in our e-commerce dataset, each purchase is associated with a time, we can
modify the target variable to incorporate timing. Specifically, we change  to  such that
 = 1 if user  purchased item  in the next time segment following , and 0 otherwise. The
length of next time segment following hour 10 and 16 is set to 6 hours, and for hour 22 it is set
to 12 hours2. In our dataset, these three time segments separate purchases records evenly, with
29,799, 28,266 and 27,898 purchases in each of the three segments.</p>
        <p>It is an important problem to determine whether the user has purchase intention at hour  or
not, before making the recommendation. Here we assume this information is already obtained,
for example, from the fact that the user visited the e-commerce website. We use the training
label  such that there is guaranteed to be an item  that user  will purchase for time . In
other words, time s when user  made no purchase at all are ignored.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Social Media Background Attention</title>
      <p>
        The social media background in our scenario is a collection of social media posts. They need to
be transformed before they can be added to a neural recommendation model. We fuse social
media background with the base recommendation system in two steps. First we convert the
social media background to machine-readable matrix inputs. Then we use attention [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] to fuse
this input with the neural recommendation model.
      </p>
      <sec id="sec-4-1">
        <title>2The hour of the day is taken as the remainder of /24.</title>
        <sec id="sec-4-1-1">
          <title>4.1. Frequency-based Representation</title>
          <p>
            We design a method that aims to capture the ℎ or  of the social media words
that comes from interesting phenomenon happening in real-world. When something interesting
or stimulating happens, some topics on social media may become trending. When this happens,
we say certain social media aspects are emerging. We capture this emergence by observing the
change in word frequency. The frequency of the social media words is taken as the count of
messages that contains the word. Thus we first obtain the frequency table of social media words
against time units. We then devise a method for emergence detection based on word frequencies.
Following an approach of previous works on social media event detection [
            <xref ref-type="bibr" rid="ref31">31</xref>
            ], our method
involves a foreground and a background. The foreground is a period closer to the current time,
and the background is a period farther from the current time, before the foreground. Suppose
the period for foreground is  , and for background is , so that word frequencies in these
periods are  = {− , ..., − 1} and  = {− − , ..., − − 1}. We set  to  ,
if − 1 &gt;  (), where  (· ) is the mean function, i.e., the frequency on the last day in the
foreground period increases compared to the mean of foreground period, and   otherwise.
Similarly we set  for the background period. Finally the emergence  of the word at time
 is set as:
{︃1, if  OR ( AND  () &gt;  ())
(5)
 =
0,
          </p>
          <p>otherwise</p>
          <p>With this formula, we aim to capture two phases of surges of words in social media. First,
 captures a new surge. Second,  AND  () &gt;  () captures the sustenance of a
previous surge. Both phases can be considered a part of an emergence. With this calculation,
we obtain for each time unit the emerging words in product sales and social media. This
representation can be called bag-of-emerging-words (BOEW).</p>
          <p>Practically, the time unit for the prediction can be set to a day, while the emergence calculation
can be done on a smaller unit, such as an hour. In this way, for each day, we obtain a vector of
length 24 for each word, indicating its emergence status in each hour of the day. Assuming we
have || words in the dictionary . For each day, we can have a matrix  that has || rows,
each row is an emergence vector of a word across 24 hours.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.2. Graph-based Representation</title>
          <p>The above method can capture the trends of individual word usages. However, words have
inter-relationships, and treating words independently would lead to loss of information. As the
semantics of words also change in real-time, and such changes need to be captured. Therefore
we propose a graph-based method to capture the inter-relationship between words that evolves
in real-time. Here we will first introduce the graph construction, and then propose a graph
embedding method that can be applied in real-time.</p>
          <p>Our graph is constructed based on the co-occurrence of words in tweets. Suppose at time 
we have a new collection of social media tweets  . We use mutual information [32] to represent
the co-occurrence behavior. Specifically, we have
(1, 2) = 
︂(  (1, 2)| | )︂ ,
 (1) (2)
where  (1, 2) is the frequency of co-occurrence of words 1 and 2, | | is the total number
of tweets, and  () is the frequency of occurrence of a single word . We use a threshold  to
determine the co-occurrence relation, such that if (1, 2) &gt; , we create (1, 2) =
_. For  time units, we would have  graphs, each of which can be represented
as an adjacency matrix .</p>
          <p>
            Graph convolutional network (GCN) [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] currently is the common method to generate
embeddings for graphs. Given an adjacency matrix  and the node embedding matrix of layer
, , we set a weight matrix  , so that the embedding in the next layer +1 is mapped
through graph convolution:
+1 = GCONV(, ,  )
          </p>
          <p>=  (ˆ )
where ˆ is a normalization of , and  is the activation function (typically ReLU) for all but the
output layer. Utilizing the spectral graph theory, the validity of this approach has been shown
in several existing works [33, 34].</p>
          <p>
            Given a series of graphs across several time units, a naive solution would be applying GCN
and generating a set of node embeddings in each graph. However, the embeddings in diferent
times would then be irrelevant to each other, and thus cannot be used for learning a single
recommendation model. Recently, a technique called EvolveGCN has been proposed to address
this problem [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ]. The basic idea of this technique is to transfer the weights in one time unit to
the next through some activation function. Specifically, it proposed:
(6)
(7)
(8)
 = LSTM(− 1)
where LSTM is a long short-term memory. The LSTM may be replaced by other recurrent
architectures, as long as the roles of  and − 1 are clear3.
          </p>
          <p>While EvolveGCN can be used to generate node embeddings peculiar to each time segment,
it has a drawback. The training of this network requires holistic data, with each training epoch
processing all time segments, thus it cannot directly be used in real-time. To address this
problem, we propose an incremental version of the EvolveGCN (IEGCN). The basic idea is that
instead of processing all time segments in one epoch, we process one time segment at a time.
Algorithm 1 shows the process.</p>
          <p>Updating parameters (line 5) can be done through some pseudo learning tasks such as link
prediction. While in the first few time segments, the embeddings are not accurate due to the lack
of information, in later time segments the embeddings should have similar representativeness
as the native EvolveGCN. With this algorithm, for each time , we can obtain a matrix , each
row of which is the embedding of a word in the dictionary .</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.3. Fusing Base System with Social Media Background Through Attention</title>
          <p>If the social media is represented as a vector, such as the average embedding of all hours, a
simple way to integrate it into the model is by concatenating the vector with outputs of a middle
3This is also known as the  version of EvolveGCN. There is also a less-known  version. We omit it here for
simplicity, but the details can be found in the referred paper.</p>
          <p>Algorithm 1 Learning an incremental EvolveGCN (IEGCN) model
INPUT:  for all  ∈ , nEpoch
OUTPUT: +1 for all  ∈ 
1: for each  ∈  do
2:  = LSTM(− 1)
3: for each  in nEpoch do
4: +1 = GCONV(, , )
5: update parameter through a loss function
6: end for
7: end for
and the loss function defined in Equation (4) is modified so that the temporal aspect is considered
layer. If we denote the vector as b, a possible position to add it is in the concatenation layer
where the outputs of GMF and MLP are jointed. If we do so, the prediction from the final layer
becomes
⎡z ⎤
ˆ =  (h · ⎣z ⎦),</p>
          <p>b
 =</p>
          <p>∑︁
(,,)∈∪−</p>
          <p>log ˆ + (1 − ) log(1 − ˆ).</p>
          <p>
            However, the above method aggregates information in a higher dimension to a lower
dimension (i.e., from a matrix to a vector), which can lead to information loss. If social media is
represented as a matrix, we can use better methods, for example, by using the attention. In recent
years, the attention mechanism in deep learning has been shown to be helpful by allowing
the model to focus on some aspects of input data [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ]. The goal of an attention module is to
produce a weighted average of candidate embeddings of a reference source, called keys, based
on their relationships with a query embedding. In our case, the keys are vectors of diferent
rows in the representation matrix and the queries are item embeddings. The item embeddings
can be obtained by concatenating the item embedding learned for GMF and MLP, q and q  .
          </p>
          <p>For clarity, we will focus on the frequency-based representation matrix first. Denote the
matrix as  for time , The output of an attention module is thus a context vector  for item 
c = ∑︁</p>
          <p>
            a = softmax(ℎ,  )
(9)
(10)
(11)
(12)
where  is the vector in row , and  is called attention weights. The attention weights can
be generally obtained using the following formula
where ℎ is the embedding of item , and  is an attention score function calculated on ℎ and
 . Several ways have been proposed to calculate attention weights. In this paper we choose a
simple approach called the  attention function [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. Basically, it is calculated as
(ℎ,  ) = ℎ⊺ 
where W is a randomized weight matrix. Although theoretically simple, it has been shown that
this function can capture the relevance of keys with respect to the query. Clearly, we can obtain
the attention output for graph-based representation matrix  in the same way.
          </p>
          <p>We consider that the attention mechanism is what we need for finding relevant real-time
aspects with respect to an item. The model after fusing the real-time information through the
attention mechanism is shown in Figure 2. The gray components are unchanged components
in the base model, while the bright components are new additions or components that have
changed because of the fusion.</p>
          <p>With this extension, the concatenation layer takes the output of the encoder and thus becomes
⎡z ⎤
ˆ =  (h · ⎢⎢⎣zcc ⎥⎥⎦),
(14)</p>
          <p>Essentially, fusing with real-time social background information this way allows the final
layers of the model to learn the latent relationship between the social media background and
user-item pairs. More specifically, the co-occurrence of real-time aspects and positive/negative
instances will be captured. In the case of social media word , for example, when the model
receives many times the co-occurrence of the emerging word  and the purchase behavior
of products containing the word , it will reinforce this pattern. By allowing attention
on diferent rows in the matrix, we also capture finer context such as diferent delays in the
causality between the social media background and purchase behaviors.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental Evaluation</title>
      <p>We conduct experiments with real-world data to test the efectiveness of our approach. For this
purpose, we implement the base model and the new model with the extension to incorporate
real-time social media background. The first set of evaluations is conducted by comparing
these two models. We then test the advantage of our model compared to other methods that
have diferent ways to represent the social media background. In this section we present the
experimental setup, including datasets and implementation details, and discuss the evaluation
results.</p>
      <sec id="sec-5-1">
        <title>5.1. Datasets</title>
        <p>We are provided with an e-commerce dataset by our industry partner for the purpose of testing
recommendation methods. The dataset is collected from a flash sales platform, which ofers
discount coupons that are made available for a limited period of time, usually between 7 and
14 days, essentially making them flash sales. The dataset contains the information of several
thousands of products and users who purchased the product during the flash sales events.
Since the available periods of the products are short, the market is rapidly changing, and thus
the majority of the products can be considered cold items that have no purchase records in
an earlier period. The products include several categories of items, such as food, cosmetics,
home appliances, hobby classes, and travel packages. All products are associated with text
descriptions written in Japanese. Each user is associated with some user information including
gender, age, and the prefecture code of their home address. The dataset is for a period of four
months, between June and September 2017.</p>
        <p>Since the products are associated with text descriptions written in Japanese, we can use word
embeddings to represent products [35]. We use a natural language processing package called
kuromoji4 to process the Japanese text. The package can efectively perform segmentation
and part-of-speech (POS) tagging for Japanese text. We use the package to tokenize the text
description and run POS tagging to select only nouns in the text. The vector representation of a
product is thus obtained as the average word embeddings of the nouns in the description.</p>
        <p>We obtain a social media dataset by collecting Japanese tweets through Twitter API5. To align
with the period of the e-commerce dataset, we develop a procedure to search past tweets. In
addition to the time requirement, it is also desirable that the tweets are talking about Japanese
domestic afairs, which reflects the background in which the e-commerce business was operated.
Our procedure is thus as the following. First, we collect a list of Japanese politician Twitter
accounts6. From them we remove a few top politician accounts such as Abe Shinzo as they
would attract foreign followers. Next we collect the follower of these politicians, who are
expected to be Japanese citizens. Then we select from these citizen accounts whose earliest
tweets are dated earlier than June 2017. This is to ensure that the accounts are active during
the entire period of the e-commerce dataset. Finally, we collect tweets in the said period from
these selected accounts. These tweets become our social media data in this study. In total this
dataset contains about 2,464,645 tweets from 33,443 accounts. Intuitively, this social media
dataset would only be weakly related to user purchase behaviors, since consumer products are
not its topic of interest. But messages in this dataset are more similar to typical social media
discussions. If we can use this dataset to improve recommendation performance, we can say
the totality of social media indeed contains predictive hints.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Implementation Details</title>
        <p>We use Pytorch7 to implement the base and the extended models. Model parameters are
randomly initiated. We follow the base model and use a tower pattern for the MLP component,
which halves the layer size for each successive higher layer. The sizes of three layers in the
MLP are thus [200, 100, 50]. The output size of GMF is set to 50. The size of the fully connected
layer is set to 100. The number of hours, , to consider as embeddings in time  is set to 24.</p>
        <p>We use the Adam optimizer with an initial learning rate of 0.001. We run 50 training epochs
for each model, before which model performances generally become stable. When training the
model, we randomly sample 4 negative instances for each positive instance. For the base model,
the negative instances (, ) are user  and item  that have no interaction. While for the new
model, the negative instances (, , ) are user  and item  that have no interaction for time .</p>
        <sec id="sec-5-2-1">
          <title>4https://github.com/atilika/kuromoji</title>
          <p>5https://developer.twitter.com/en/docs
6Such a list can be found online as political social media accounts are usually public. An example list is provided by
the website Meyou with the url https://meyou.jp/group/category/politician/
7https://pytorch.org/</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Evaluation Settings</title>
        <p>We divide that dataset into a training set and a test set. The training set is for a period from June
1 to September 16, 2017, and the test set is between September 17 to 30, 2017, a period of two
weeks. We remove what are so-called free items, which are time-limited discount coupons of 0
price, which anyone can get during the active period without paying a fee. These items occupy
a large portion of the dataset, but they do not reveal user-item preferences, so we consider them
noises. After removing the free items, the number of users, items, and interactions are shown
in Table 1.</p>
        <p>Considering the consistency of the evaluation, for each interaction in the test dataset (positive
instances), we randomly sample 99 negative instances from items available of the same time
segment. Combining the positive and negative instances, we have 100 candidate items in each
recommendation.</p>
        <p>We use hit-rate (HR) and Normalized Discounted Cumulative Gain (NDCG) to measure
recommendation performance. HR@K is calculated as
number of hits in top K recommendation
number of recommendations
.</p>
        <p>HR@K measures whether the correct item is in the recommended items, but it does not consider
the rank of the item. NDCG on the other hand counts the position of the correct item. It is
calculated as:
 @ = ∑︁ 2 − 1
=1 log2( + 1)
,
where  = 1 if the correct item is ranked in the -th position, and 0 otherwise. NDCG@K will
be higher if the correct item is ranked higher in recommended items. To simulate a realistic
scenario, where the recommended items are shown on a single web page to e-commerce website
visitors, we choose a number of K values between 3 and 10.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Comparing Base Model and Extended Models</title>
        <p>We compare the base model (base) with three variations of our extended model, the one
with only trend information captured by bag-of-emerging-words (BOEW), the one with only
semantic information captured by incremental Evolve-GCN (IEGCN), and the one with both
parts of information (BOEW + IEGCN). Recommendation performances measured as HR@K and
NDCG@K of diferent models are shown in Table 2. For each measurement of three variations
of the proposed model, the relative performance increases are shown below that measurement.
The best-performing results are highlighted in bold font.
(15)
(16)</p>
        <p>The main insight from these results is that considering social media background with our
approaches generally improves the recommendation prediction. Both the trend information and
semantic information are helpful in improving the accuracy, while the combination of the two
achieves even better accuracy than using them separately. When using the trend information, the
HR@K is improved by 1.65% to 5.9% for diferent K values. Using semantic information achieves
better accuracy, with HR@K improving between 5% to 11.89% for diferent K values. Using both
parts of information achieves the best accuracy in our evaluation, with HR@K improved between
7.5% to 15.7% with diferent K values. Similar trends are observed in HDCG@K results. Thus
from the results we can see that both trend information and semantic information contribute to
the predictiveness of social media background. Furthermore, incorporating the combination of
them achieves better accuracy than using them separately, indicating that their contributions
are complementary to each other.</p>
      </sec>
      <sec id="sec-5-5">
        <title>5.5. Comparing Diferent Methods of Social Media Background</title>
      </sec>
      <sec id="sec-5-6">
        <title>Representation</title>
        <p>We have shown that our method for representing and fusing the social media background can
improve recommendation accuracy compared to the base model. We acknowledge that there
are other ways to represent social media background, from simple to complex methods. In this
set of experiments, we compare our method to several other real-time representation methods
for social media background. We briefly introduce them as the following.</p>
        <p>Bag-of-words (BOW). One of the most common methods for text representation, BOW
has been the baseline in many previous researches in text mining [36, 37]. This method is
time independent and word independent. To integrate it with our model, we produce a BOW
representation for each word  and each time , so that , is the frequency count of the word
in that time period. This will generate a frequency vector in each time  of length ||, which
we concatenate in the concatenation layer in Fig 1.</p>
        <p>incremental tf-idf. Term frequency-inverse document frequency (tf-idf) has been shown to
be a more efective representation of texts, with the additional information of total document
count [38]. The formula for tf-idf is:
tf-idf(, ) = , · log</p>
        <p>| ∈  :  ∈ |
where , is the frequency of word  in the document , and  is the total
collection of documents. To apply it, we consider social media text posted in one time unit  as
a document. For the total set of documents, we set a look-back window of length ℎ so that
 = {− ℎ, ..., }. Again we can have a larger time unit and a small time unit, i.e., a
day and an hour. After applying tf-idf, we can have a matrix, each row is a vector representing
the tf-idf value of a word  across 24 hours of the day. The matrix has || rows. Then it is
integrated with the model using the same method as the trend matrix or the semantic matrix in
Section 4.3. This method is time-dependent because the total set of documents depends on the
current time.</p>
        <p>incremental bi-term topic model (IBTM). Over the past two decades, topic-modeling such
as Latent Dirichlet Allocation (LDA) has become a popular method for text representation [39].
Such methods consider the co-occurrence of words in texts, thus can capture word semantics.
Bi-term topic model (BTM) is a new variation of LDA that use bi-term instead of single-term
for calculation [40]. The representation is based on the generative assumption below:
1. Draw  ∼ Dirichlet( ).
2. For each topic  ∈ [1, ] draw  ∼ Dirichlet( ).</p>
        <p>3. for each biterm  ∈ , draw  ∼ Multinomial( ), and draw ,1, ,2 ∼ Multinomial( ).
The parameters in this generative process can be learned through techniques such as Gibbs
Sampling. As the result of learning, the parameter  represents a distribution of probabilities
on which a document is drawn from  topics. The parameter  represents a distribution of
probabilities a topic is drawn on || words. The incremental version of the algorithm, provided
by the same authors, trains a single model over a bi-term stream using an incremental Gibbs
sampler. When applying to our problem, a tweet is considered a document, and all tweets posted
in time  are considered the total collection of documents. As the result, for each time step , we
can get , which is a matrix with || rows, each of which is a vector of the probability value a
word  has for  topics. Since it is a matrix input, we can integrate it using the same method
presented in Section 4.3.</p>
        <p>
          Emergence with Embedding (E-EMB). In a previous work, we proposed a model that
tries to combined both trends and semantics [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. Similar to this work, it performs emergence
analysis on words. However, instead of generating bag-of-emerging-words, it uses pre-train
word2vec embeddings to represent words, and each time unit is represented as the average
embedding of the emerging words. In this way, it captures the trends and the global semantics.
It also operates on two levels of time units, i.e., day and hour. For each day, we can have a
matrix, each row is a vector of one dimension of the embedding across 24 hours.
        </p>
        <p>
          incremental graph convolutional network (GCN). GCN is the base form of Evolve GCN
[
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. We can obtain a GCN model from Evolve GCN by setting the look-back period to zero.
The incremental version of GCN is similar to Algorithm 1, and we only need to remove the
weight transfer step (line 2). The incremental GCN produces node embeddings in each time step
aligning to the same format as the IEGCN, and can be processed using the method described in
Section 4.3.
        </p>
        <p>We apply all baseline methods to the same Twitter data and generate diferent
representations. And then we input the representation, either in vector form or in matrix form, to the
recommendation model. Using the same test settings in the last set of experiments, the accuracy
results of all compared representations are shown in Table 3.</p>
        <p>As we can see from the table, our method steadily outperforms all other representations
in all measurements. The best compared model seems to be IBTM, which achieves 0.139 for
NDCG@10. Our method, however, outperforms this value by about 5%. Looking at other
baselines, we see that BOW is the least efective representation. E-EMB, the previous proposed
method, is comparable to GCN, but is worse than our new method. The tf-idf method, while
being simple to implement, achieves a good HR@10, only 1% less than our method, although
its HR@3 is poor. To conclude, while each representation has its strength and weakness, our
method achieves the best performance, mostly due to its capability to consider trends and
semantics at the same time.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this paper, we propose a method to integrate social media background in dynamic
recommender systems in real-time. Our method consists of a representation of social media, and an
attention-based fusing method. The representation takes into account both real-time trends
and evolving semantics of words. Experimental evaluations with real-world e-commerce and
social media datasets show that our method is feasible, with steadily improved accuracy results
achieved by the extended model. The representation is also shown to have an advantage over
several other real-time representations of the social media background. Our method is suitable
to be deployed practically because social media data can be easily obtained, and there is no
requirement to link user accounts. We would like to make further investigations, however,
because it is still dificult to tell which social media contexts are predictive for which products. To
make our method easier to explain, in the future, we plan to find ways to make the relationships
between social media background and product purchasing behavior more explicit. We also plan
to deploy our system on real e-commerce platforms and view its impact on product sales in
real-time.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>This research is partially supported by JST CREST Grant Number JPMJCR21F2.
microblogs, in: Proceedings of the 36th international ACM SIGIR conference on Research
and development in information retrieval, ACM, 2013, pp. 43–52.
[32] H. Peng, F. Long, C. Ding, Feature selection based on mutual information criteria of
max-dependency, max-relevance, and min-redundancy, IEEE Transactions on pattern
analysis and machine intelligence 27 (2005) 1226–1238.
[33] W. Dai, O. Jin, G.-R. Xue, Q. Yang, Y. Yu, Eigentransfer: a unified framework for transfer
learning, in: Proceedings of the 26th Annual International Conference on Machine
Learning, 2009, pp. 193–200.
[34] M. Schlichtkrull, T. N. Kipf, P. Bloem, R. Van Den Berg, I. Titov, M. Welling, Modeling
relational data with graph convolutional networks, in: European semantic web conference,
Springer, 2018, pp. 593–607.
[35] V. Lampos, B. Zou, I. J. Cox, Enhancing feature selection using word embeddings: The case
of flu surveillance, in: Proceedings of the 26th International Conference on World Wide
Web, International World Wide Web Conferences Steering Committee, 2017, pp. 695–704.
[36] W. Stokowiec, T. Trzciński, K. Wołk, K. Marasek, P. Rokita, Shallow reading with deep
learning: Predicting popularity of online content using only its title, in: International
Symposium on Methodologies for Intelligent Systems, Springer, 2017, pp. 136–145.
[37] Y. Zhang, M. Shirakawa, T. Hara, Predicting temporary deal success with social media
timing signals, Journal of Intelligent Information Systems (2021) 1–19.
[38] A. Ramisa, F. Yan, F. Moreno-Noguer, K. Mikolajczyk, Breakingnews: Article annotation by
image and text processing, IEEE transactions on pattern analysis and machine intelligence
(2017).
[39] D. M. Blei, A. Y. Ng, M. I. Jordan, Latent dirichlet allocation, Journal of Machine Learning</p>
      <p>Research 3 (2003) 993–1022.
[40] X. Cheng, X. Yan, Y. Lan, J. Guo, Btm: Topic modeling over short texts, IEEE Transactions
on Knowledge and Data Engineering 26 (2014) 2928–2941.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sarwar</surname>
          </string-name>
          , G. Karypis,
          <string-name>
            <given-names>J.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          ,
          <article-title>Item-based collaborative filtering recommendation algorithms</article-title>
          ,
          <source>in: Proceedings of the 10th International Conference on World Wide Web, ACM</source>
          ,
          <year>2001</year>
          , pp.
          <fpage>285</fpage>
          -
          <lpage>295</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <article-title>Latent aspect rating analysis on review text data: a rating regression approach</article-title>
          ,
          <source>in: Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining, ACM</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>783</fpage>
          -
          <lpage>792</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zha</surname>
          </string-name>
          ,
          <article-title>Learning binary codes for collaborative filtering</article-title>
          ,
          <source>in: Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>498</fpage>
          -
          <lpage>506</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Neural collaborative filtering</article-title>
          ,
          <source>in: Proceedings of the 26th International Conference on World Wide Web</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , Perth, Australia,
          <year>2017</year>
          , pp.
          <fpage>173</fpage>
          -
          <lpage>182</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X. N.</given-names>
            <surname>Lam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Vu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. D.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Duong</surname>
          </string-name>
          ,
          <article-title>Addressing cold-start problem in recommendation systems</article-title>
          ,
          <source>in: Proceedings of the 2nd international conference on Ubiquitous information management and communication</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>208</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Haruna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Akmar</given-names>
            <surname>Ismail</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Suhendroyono</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Damiasih</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Pierewan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chiroma</surname>
          </string-name>
          , T. Herawan,
          <article-title>Context-aware recommender system: A review of recent developmental process and future research direction</article-title>
          ,
          <source>Applied Sciences</source>
          <volume>7</volume>
          (
          <year>2017</year>
          )
          <fpage>1211</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Lika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kolomvatsos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hadjiefthymiades</surname>
          </string-name>
          ,
          <article-title>Facing the cold start problem in recommender systems</article-title>
          ,
          <source>Expert Systems with Applications</source>
          <volume>41</volume>
          (
          <year>2014</year>
          )
          <fpage>2065</fpage>
          -
          <lpage>2073</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <article-title>Addressing the item cold-start problem by attribute-driven active learning</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>32</volume>
          (
          <year>2019</year>
          )
          <fpage>631</fpage>
          -
          <lpage>644</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>McAuley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <article-title>Hidden factors and hidden topics: understanding rating dimensions with review text</article-title>
          ,
          <source>in: Proceedings of the 7th ACM conference on Recommender systems, ACM</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>N.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <article-title>Time-sensitive recommendation from recurrent user activities</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>3492</fpage>
          -
          <lpage>3500</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , T. Maekawa,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hara</surname>
          </string-name>
          ,
          <article-title>Using social media background to improve cold-start recommendation deep models</article-title>
          ,
          <source>in: Proceedings of 2021 IEEE International Joint Conference on Neural Networks IJCNN</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Y.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-R.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Connecting social media to ecommerce: Cold-start product recommendation using microblogging information</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>28</volume>
          (
          <year>2015</year>
          )
          <fpage>1147</fpage>
          -
          <lpage>1159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>Twitter is faster: Personalized time-aware video recommendation from twitter to youtube</article-title>
          ,
          <source>ACM Transactions on Multimedia Computing</source>
          , Communications, and
          <string-name>
            <surname>Applications</surname>
          </string-name>
          (TOMM)
          <volume>11</volume>
          (
          <year>2015</year>
          )
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>M.-T. Luong</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Pham</surname>
            ,
            <given-names>C. D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          ,
          <article-title>Efective approaches to attention-based neural machine translation</article-title>
          ,
          <source>in: Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>1412</fpage>
          -
          <lpage>1421</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>in: Advances in neural information processing systems</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Maciejewski</surname>
          </string-name>
          ,
          <article-title>Graph convolutional networks: a comprehensive review</article-title>
          ,
          <source>Computational Social Networks</source>
          <volume>6</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pareja</surname>
          </string-name>
          , G. Domeniconi,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          , T. Ma, T. Suzumura,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kanezashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kaler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Schardl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Leiserson</surname>
          </string-name>
          , Evolvegcn:
          <article-title>Evolving graph convolutional networks for dynamic graphs</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>34</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>5363</fpage>
          -
          <lpage>5370</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tay</surname>
          </string-name>
          ,
          <article-title>Deep learning based recommender system: A survey and new perspectives, ACM Computing Surveys (CSUR) 52 (</article-title>
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>38</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Cebrián</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Planagumà</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Villegas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Amatriain</surname>
          </string-name>
          ,
          <article-title>Music recommendations with temporal context awareness</article-title>
          ,
          <source>in: Proceedings of the fourth ACM conference on Recommender systems</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>349</fpage>
          -
          <lpage>352</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Fonseca</surname>
          </string-name>
          ,
          <article-title>Improving music recommendation in session-based collaborative ifltering by using temporal context</article-title>
          ,
          <source>in: 2013 IEEE 25th international conference on tools with artificial intelligence, IEEE</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>783</fpage>
          -
          <lpage>788</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Hsu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>A time-sensitive personalized recommendation method based on probabilistic matrix factorization technique</article-title>
          ,
          <source>Soft Computing</source>
          <volume>22</volume>
          (
          <year>2018</year>
          )
          <fpage>6785</fpage>
          -
          <lpage>6796</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Beutel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Covington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gatto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          , Latent cross:
          <article-title>Making use of context in recurrent recommender systems</article-title>
          ,
          <source>in: Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>54</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D. H.</given-names>
            <surname>Alahmadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-J.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <article-title>Twitter-based recommender system to address cold-start: A genetic algorithm based trust modelling and probabilistic sentiment analysis</article-title>
          ,
          <source>in: 2015 IEEE 27th International Conference on Tools with Artificial Intelligence (ICTAI)</source>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>1045</fpage>
          -
          <lpage>1052</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>H.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          , H. Liu,
          <article-title>Exploring temporal efects for location recommendation on location-based social networks</article-title>
          ,
          <source>in: Proceedings of the 7th ACM conference on Recommender systems, ACM</source>
          ,
          <year>2013</year>
          , pp.
          <fpage>93</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Chen,
          <article-title>Who's next: Rising star prediction via difusion of user interest in social networks</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Twitter volume spikes and stock options pricing</article-title>
          ,
          <source>Computer Communications</source>
          <volume>73</volume>
          (
          <year>2016</year>
          )
          <fpage>271</fpage>
          -
          <lpage>281</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>S.</given-names>
            <surname>Asur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Huberman</surname>
          </string-name>
          ,
          <article-title>Predicting the future with social media, in: Web Intelligence and Intelligent Agent Technology (WI-IAT)</article-title>
          , volume
          <volume>1</volume>
          , IEEE,
          <year>2010</year>
          , pp.
          <fpage>492</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>P.-F. Pai</surname>
          </string-name>
          , C.-H. Liu,
          <article-title>Predicting vehicle sales by sentiment analysis of twitter data and stock market values</article-title>
          ,
          <source>IEEE Access 6</source>
          (
          <year>2018</year>
          )
          <fpage>57655</fpage>
          -
          <lpage>57662</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Broniatowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dredze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Paul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dugas</surname>
          </string-name>
          ,
          <article-title>Using social media to perform local influenza surveillance in an inner-city hospital: a retrospective observational study</article-title>
          ,
          <source>JMIR public health and surveillance</source>
          <volume>1</volume>
          (
          <year>2015</year>
          )
          <article-title>e5</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amagata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Makeawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yonekawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kurokawa</surname>
          </string-name>
          ,
          <article-title>A dnnbased cross-domain recommender system for alleviating cold-start problem in e-commerce</article-title>
          ,
          <source>IEEE Open Journal of the Industrial Electronics Society</source>
          <volume>1</volume>
          (
          <year>2020</year>
          )
          <fpage>194</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Amiri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          , T.-S. Chua,
          <article-title>Emerging topic detection for organizations from</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>