<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Temporal Recurrent Activation Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giuseppe Manco</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giuseppe Pirro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ettore Ritacco</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ICAR - CNR</institution>
          ,
          <addr-line>via Pietro Bucci 7/11C, 87036 Arcavacata di Rende (CS)</addr-line>
          ,
          <country country="IT">ITALY</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>24</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>We tackle the problem of predicting whether a target user (or group of users) will be active within an event stream before a time horizon. Our solution, called PATH, leverages recurrent neural networks to learn an embedding of the past events. The embedding allows to capture in uence and susceptibility between users and places closer (the representation of) users that frequently get active in di erent event streams within a small time interval. We conduct an experimental evaluation on real world data and compare our approach with related work.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        an embedding of the past event history via Recurrent Neural Networks that also
cater for the di usion memory. The embedding allows to capture in uence and
susceptibility between users and places closer (the representation of) users that
frequently get active in di erent streams within a small time interval.
Related Work. We conceptually separate related research into: (i) approaches
like DeepCas [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and DeepHawkes [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] that tackle the problem of predicting the
length that a cascade will reach within a timeframe or its incremental popularity;
(ii) approaches like Du et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Neural Hawkes Process (NHP) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] model
and predict time event markers and time; (iii) approaches based on Survival
Factorization (SF) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] that leverage in uence and susceptibility for time and event
predictions; (iv) other approaches that do not use neural networks (e.g., [
        <xref ref-type="bibr" rid="ref12 ref3 ref5">5, 3,
12</xref>
        ]). PATH adopts a di erent departure point from these approaches: it focuses on
predicting the activation of (groups of) users before a time horizon. Di erently
from (i) PATH considers time and uses an embedding to capture both in uence
and susceptibility between users and predict future activations. Moreover, (i)
focuses on the prediction of cumulative values only (e.g., cascade size). Di erently
from (ii), we do not assume that time and event are independent. Besides, (ii)
focuses on predicting event types (e.g., popular users), which is not enough in the
scenarios targeted by PATH (e.g., targeted market campaigns) where one is
interested in predicting the behavior of speci c users and not their types. A for (iii),
it fails in capturing the cumulative e ect of history while PATH captures by using
an embedding. As for (iv), the main di erence is that PATH can automatically
learn (via neural networks) an embedding representing in uence/susceptibility.
The contributions of the paper are as follows: (i) PATH, a classi cation-based
approach based on recurrent neural networks allowing to model the likelihood
of observing an event as a combined result of the in uence of other events; (ii)
an experimental evaluation and a comparison with related work.
The remainder of the paper is organized as follows. We introduce the problem in
Section 2. We present PATH in Section 3. We compare our approach with related
research in Section 4. We conclude and sketch future work in Section 5.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Problem</title>
    </sec>
    <sec id="sec-3">
      <title>De nition</title>
      <p>We focus on network of individuals who react to solicitations along a timeline.
An activation network can be viewed as an instance of a marked point processes
on the timeline, de ned as a set X = f(th; sh)g1 h m. Here, th 2 R+ denotes the
events of the point process, and sh 2 M denote the marks in the measurable space
M. Relative to activation networks, the speci cation of sh occurs by means of the
realizations uh, ch and xh, where uh 2 V (with jVj = N ) represent individuals,
ch 2 I (with jIj = M ) represent solicitations and xh is side information which
characterizes of the reaction of the entity, described as an instance relative to a
feature space of interest. For example, V can represent users who are engaged in
online discussions I, and the tuple (th; (uh; ch; xh)) represents the contribution
of uh to discussion ch with the post xh. It is convenient to view the process
as a set of cascades: that is, for each c 2 I we can consider the subset Hc =
f(t; u; x)j(t; (u; c; x)) 2 Xg of elements marked by c, with mc = jHcj. Also, tc and
U c represent the projections on the rst and second column of Hc. We also denote
by H&lt;ct (resp. Hc t) the set of events ei 2 Hc such that ti &lt; t (resp., ti t). The
terms tc&lt;t and U &lt;ct can be de ned accordingly. The relationship u c v denotes
that both u and v are active in Hc and there are some events relative to u and
v such that u precedes v in some events. Finally, C = fH1; HM g denote a
collection of M cascades over V and I.</p>
      <p>Modeling di usion. We start from the observation that what is likely to
happen in the future (viz. which user will be active and when) depends on what
happened in the past (viz. the chain of previously active users). One
important point to take into account is the susceptibility of users, that is, the extent
to which they are in uenced by speci c previously activated users. Our model
should be exible enough to re ect both exciting and inhibitory e ects. While
the former boosts the likelihood of observing u active in c, the latter actually
could prevent it to do so. Given a cascade Hc, a timestamp t 0 and a user
u 62 U &lt;ct, the goal is to obtain an estimate of the density function f (t; ujH&lt;ct),
which can be used to model the following evolution scenario: given a time horizon
T c; how likely is it that u will become active in c within T c?</p>
      <p>h h
The challenge, at this point, is how to concretely formulate the density f . We
can decouple its speci cation as follows:
f (t; ujH&lt;ct) = g(tju; H&lt;ct) h(ujH&lt;ct);
(2.1)
where the rst component represents the likelihood that u becomes active within
t, given the current history, and the second component represents the likelihood
that u activates (independent of the time) as a reaction to the current history.</p>
      <p>As for H&lt;ct, explicit information includes features like the sequence of user
activations, their activation times, the relative activation speed, and possibly
c
the topic of the cascade. Nevertheless, our assumption is that H&lt;t can also
encode latent information including susceptibility and in uence between users
that can be derived, for instance, from neighborhood information in a network
(e.g., follower/followee relations in Twitter) or user behaviors (e.g., users that
retweet after a certain set of other in uential user (re)tweet). This is exactly
what we want to unveil in our modeling.</p>
      <p>
        Embedding history. We want to learn and embedding of users in a latent
K-dimensional space such that users in the same cascade are closer in the
embedding, and users within di erent cascades are distant. We make usage of two
matrices S = [s1; : : : ; sN ]; A = [a1; : : : ; aN ] 2 RN K that represent the
susceptibility and in uence, respectively. Matrices are computed by relying on the
standard network architecture borrowed from the word2vec paradigm [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]:
av = Wev
su = Veu
Here, u; v represents the one-hot encodings of u and v. The matrices We; Ve
represent the embeddings, obtained by minimizing an adapted form of contrastive
loss [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] that penalizes the distance of users within the same cascades and the
closeness of users in di erent cascades.
      </p>
      <p>
        Capturing the di usion memory. To encode temporal relationship within
c
H&lt;t we use recurrent neural networks (RNNs) An RNN is a recursive structure
that, at the current step, gets as input the previous network state (the outputs
form the hidden units) along with the current input to compute a new state.
The following picture provides an overview of a simple RNN cast to our context.
At each step k, we feed into the network
an event ek 2 Hc that encodes the current
user (uk) and its activation time (tk). The
learned hidden state (hk) represents the non- f(tk,uk) f(tk+1,uk+1)
linear dependency between these components hk 1 hk hk+1
and past events, which can be used to model
f (tk; ukjH&lt;ctk ). In the following, we adopt time
the LSTM instantiation of the RNN frame- ek ek+1
work [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The idea of an LSTM unit is to
reliably transmitting important information
many time steps into the future. At every time step, the unit modi es the
internal status by deciding which part to keep or replace with new information
coming from the current input. We use the shortcut hk = LSTM(zk; hk 1) to
denote a functional architecture that elaborates an input zk and outputs the
updated state.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>PATH: Predicting User Activation from a Horizon</title>
      <p>We now introduce PATH (Predicting User Activation from a Horizon), which
focuses on simplifying f (t; ujH&lt;ct) as the binary response function I(t Thju; H&lt;ct)
that denotes whether u becomes active in c within Th. We focus on events
ek 2 Hc where where the features of interest xk are limited to the time
delay k = tk tk 1 relative to the previous activation within the cascade. This
allows us to capture the property that cascades may have intrinsically di
erent di usion speeds causing some of them to concentrate users' activations in a
short timeframe while others in a more extended interval. Given a partially
observed cascade H&lt;tl (with tl &lt; Thc representing the timespan of the observation
c
window), our objective is to predict, for a given entity u 62 Utl , whether u 2 UThc .</p>
      <p>In order to uncover all the characteristics of the activations within cascades,
we consider a model built on all possible pre xes of the available cascades. Notice
that, in our reconstruction, we do not consider the rst element within the
cascade, which we assume becomes \spontaneously" active. Finally, for each
u 62 U c, we associate the cascades Hc tj [ f(tj ; u; j )g (with 1 j mc 1)
and Hc [ f(Thc; ui; (Thc tmc ))g with negative labels. Again, the intuition is
that, since u is not active no partial cascade provides the su cient intensity to
activate u within the given time horizon. Adding negative examples in the data
preparation represent an e ective data augmentation process, which enlarges
the training data by inferring new inputs in the training set. This is crucial to
let the approach better ne tune separation between active and inactive users,
as well as better characterize the true activation time of active users. Let TC
denote the set of all pairs partial sequence/associated label that can be built
from the above discussion. Our idea is to exploit the embedding and LSTM
tools described in the previous section to solve the supervised problem at hand.
Figure 1 illustrates the basic
architecture of the model. Given a
pair hHi; yii 2 TC with jHj=n and
by considering ek = (tk; uk; k) 2 y˜ yˆ
tHur(ewoitfht1he nketwonr)k, tchaenabrcehictaepc-- OLuatypeurt SimSciloarreity Aprcetdiviacttiioonn
tured by the following equations:
ak =Weuk
hk =LSTM ([ak; tk; k]; hk 1)</p>
      <p>(3.2)
y^i = (Wohn)
8
&lt;
y~i = exp</p>
      <p>an
:
n 1
X ak
k=1
(3.1)
(3.3)</p>
      <p>Here, y^i represents the
probability that yi is positive, as
provided by the network: that is, it
encodes the probability that un
becomes active within tn; y~i encodes the a nity between un and all users
preceding it within Hi. The distance kan Pkn=11 akk plays a crucial role here: since
the target user is on the tail of the cascade, the embedding should emphasize the
similarities with the predecessors that trigger an activation, and by the converse
minimize the similarities with those ones which do not trigger it. The loss is a
combination of cross-entropy and the embedding loss previously described:
Fig. 1: Overview of PATH.</p>
      <p>L =
where</p>
      <p>X
hHj H;yji=2nTC</p>
      <p>and
4</p>
    </sec>
    <sec id="sec-5">
      <title>Experiments</title>
      <p>fy ( log(y^) + log y~) + (1 y) ( log(1 y^) + log(1 y~))g
(3.5)
are weights balancing cross-entropy and embedding.</p>
      <p>We validate our approach by analysing the algorithm on real-life datasets. In
particular, we analyse the capability of the algorithm at predicting the activation
time of users within an information cascade. The implementation we use in the
experiments can be found at https://github.com/gmanco/PATH.
Datasets. We evaluated the prediction capability of PATH by exploiting two
real-world datasets containing propagation cascades crawled from the timelines
of Twitter2 and Flixster.3 In particular, Twitter includes 32K nodes with
9K cascades while Flixster includes 2K nodes with 5K cascades.</p>
      <sec id="sec-5-1">
        <title>2 http://www.twitter.com/ 3 http://www.flixster.com/</title>
        <p>The information propagation mechanism on Twitter is expressed by
retweeting, in other words a chain of repetitions and transmissions of a tweet from a set
of users to their neighbors in a recursive process. Each activation corresponds to
a retweet. An activation in Flixster happens when a user rates a movie, while
a cascade is composed by all the activations related to the same movie. The two
datasets di er essentially for the following characteristics: Twitter includes a
larger number of users and shorter delays than Flixster. In addition, retweets
intuitively highlight two relevant aspects, namely the importance of the topic
and the single in uence of the individual from which the retweet is performed.
By contrast, movie ratings are more likely to exhibit a cumulative e ect: popular
movies are more likely to be considered than unpopular ones.</p>
        <p>
          Evaluation Methodology. We evaluate PATH against two baseline models,
both relying on Survival Analysis [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The rst instantiation implements a Cox
proportional hazard model (CoxPh in the following). We implement the model
using the lifelines4 package and extract, for each event ek 2 Hc, the following
features: (1) size of the pre x; (2) last activation time; (3) average delay for
each active user so far; (4) number of neighbors in the history, and (5) coverage
percentage of them within the history; (6) the activation time of the most recent
neighbor, if any; (7) correlation between the activation of the current user an its
neighbors within the history, computed in previous cascades. This model
represents an intuitive baseline where features are manually engineered and include a
mix of external information (coming from the underlying network neighborhood)
and information derived from the cascade itself. The second instantiation is given
by the Survival Factorization (SF in the following) framework described in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
The comparison is important since SF relies on the same guiding ideas of PATH
(in uence/susceptibility) with the di erence that there is no cumulative e ect
c
of H&lt;tk , but instead an in uential user has to be detected for each activation.
        </p>
        <p>To evaluate the approaches we proceed as follows: given training and test
sets Ctrain and Ctest , we train the model on Ctrain and measure the accuracy
of the predictions on Ctest . The two sets are obtained by randomly splitting
the original dataset by ensuring that there is no overlap among the cascades of
the two sets, but there is no entity in the test that has not been observed in
the training. For the evaluation, we chronologically split each cascade c 2 Ctest
into c1 and c2 such that, for each u 2 c1 and v 2 c2, we have that u c v.
Next, we pick a random subsample c3 V U c. Then, given a target horizon
T c, we measure TP, FP, TN and FN by feeding the models on c1 and then
h
predicting the activation within Thc for each element in c2 [ c3. The choice of T c
h
can follow di erent strategies; Fixed horizon (Fixed horizon (FH): setting Tch as
the maximum observed activation time T test = maxftjt 2 tc; c 2 Ctest g; Variable
h
horizon (VH): varying Thc from the smallest to the largest activation time and
computing the activation probabilities associated to each possible value; Actual
Time (AT): a particular case of the VH strategy, where Thc , T u;c is relative to
h
the true activation time in c of each user u 2 c2 [ c3.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4 see http://lifelines.readthedocs.io for details.</title>
        <p>We plot the ROC and the F-Measure curves relative to the above alternatives
and report the AUC and F values. For PATH, the encoding of sequences as
described in section 3 already presumes that users are evaluated on intermediate
timestamps prior to their actual activation. Thus, VH and AT roughly coincide
in this case. Since both CoxPh and SF are capable of inferring, for each (u; H)
pair, the probability Su(tjH), the comparison with PATH is done by computing
1 Su(t~jH) where t~ is the horizon timestamp. The parameter space for PATH
was explored by grid-search, measuring the loss on a separate portion of the
training set by 5-fold cross-validation. We report 5 di erent instantiations, which
di er from the number of cells in the LSTM (32/64), the dimensionality of the
embedding (32/64) and the batch size in the training (128/256/512). Concerning
SF, the number of factors was set to 16 for both datasets.</p>
        <p>Evaluation Results. Figure 2 reports the ROC curves where we observe that
PATH outperforms the baselines and in particular exhibits a very good accuracy
on all con gurations. This is especially true on Flixster, where by the converse
SF does not seem capable of correctly correlating previous activations times.
The cumulative in uence e ect is evident here, as a natural consequence of the
underlying domain where cascading e ects are more likely as a consequence of a
\word of mouth" process. On Twitter, where the activation is more likely due
to the in uence of a single user (as testi ed by the good performance of SF),
PATH still achieves the best scores, thus proving the capability of the recurrent
layer to adapt the in uence to a single user. By analyzing Fig. 3, which displays
the F-measure curve for varying values of the threshold on the probabilities, we
can observe that, contrary to the baselines, higher thresholds do not cause a
signi cant drop of the recall. The only exception is CoxPh (FH), which seems
more stable on Flixster. This is a clear sign that the probabilities associated
with active and inactive users in PATH di er substantially, and in particular active
events are associated with signi cantly higher probabilities than inactive events.
5</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Concluding Remarks and Future Work</title>
      <p>We focused on the problem of predicting user activations in a given time horizon
and show that the embedding of the user activation history, where users that
0.96 0.96
0.88 0.88
0.80 0.80
-rsFaeeuM0000000000..........443567200180264240860.00 0.08 0.16 0.24 0.32 0.40Th0r.4e8sho0l.d56 0.64 0PPPPPSCC.AAAAAFoo7(TTTTTxx2VHHHHHPPHhh(((((0)AAAAA((.FATTTTT8HT,,,,,063333))422220,,,,,63363.8422428,,,,,2255155112066228.9)))))6 -rsFaeeuM0000000000..........443567210080264246080.00 0.08 0.16 0.24 0.32 0.40Th0r.4e8sho0l.d56 0.64 0PPPPPSCC.AAAAAFoo7(TTTTTxx2VHHHHHPPHhh(((((0)AAAAA((.FATTTTT8HT,,,,,033336))222240,,,,,63336.8422248,,,,,5521211525022686.9)))))6</p>
      <p>(a) Flixster (b) Twitter</p>
      <p>Fig. 3: F-Measure curves for PATH,CoxPh and SF on both datasets.
become active on the same cascades are placed close, can be e ectively learned
via recurrent neural networks. Experiments performed on real datasets show
the e ectiveness of the approach in accurately predicting next activations. It
is natural to wonder whether it is possible to cast the intuitions behind our
approach in a generative setting, to predict both which user is likely to become
active, and the time segment upon which s/he will become active.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>N.</given-names>
            <surname>Barbieri</surname>
          </string-name>
          , G. Manco, and
          <string-name>
            <given-names>E.</given-names>
            <surname>Ritacco</surname>
          </string-name>
          .
          <article-title>Survival factorization on di usion networks</article-title>
          .
          <source>In PKDD</source>
          , pages
          <volume>684</volume>
          {
          <fpage>700</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Q.</given-names>
            <surname>Cao</surname>
          </string-name>
          et al.
          <article-title>Deephawkes: Bridging the gap between prediction and understanding of information cascades</article-title>
          .
          <source>In CIKM</source>
          , pages
          <volume>1149</volume>
          {
          <fpage>1158</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Peng</given-names>
            <surname>Cui</surname>
          </string-name>
          , Shifei Jin, Linyun Yu,
          <string-name>
            <given-names>Fei</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Wenwu Zhu</surname>
            , and
            <given-names>Shiqiang</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Cascading outbreak prediction in networks: a data-driven approach</article-title>
          .
          <source>In KDD</source>
          , pages
          <volume>901</volume>
          {
          <fpage>909</fpage>
          . ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>N.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Trivedi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Upadhyay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gomez-Rodriguez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Song</surname>
          </string-name>
          .
          <article-title>Recurrent marked temporal point processes: Embedding event history to vector</article-title>
          .
          <source>In KDD</source>
          , pages
          <volume>1555</volume>
          {
          <fpage>1564</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A.</given-names>
            <surname>Guille</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Hacid</surname>
          </string-name>
          .
          <article-title>A predictive model for the temporal dynamics of information di usion in online social networks</article-title>
          .
          <source>In WWW</source>
          , pages
          <volume>1145</volume>
          {
          <fpage>1152</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>R.</given-names>
            <surname>Hadsell</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. Chopra.</surname>
          </string-name>
          , and
          <string-name>
            <surname>Y. LeCun.</surname>
          </string-name>
          <article-title>Dimensionality reduction by learning an invariant mapping</article-title>
          .
          <source>In CVPR</source>
          , pages
          <volume>1735</volume>
          {
          <fpage>1742</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural Computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ):
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>J.D. Kalb</surname>
            eisch and
            <given-names>L.P.</given-names>
          </string-name>
          <string-name>
            <surname>Ross</surname>
          </string-name>
          .
          <article-title>The Statistical Analysis of Failure Time Data</article-title>
          . Wiley Series in Probability and Statistics,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Guo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Mei</surname>
          </string-name>
          . Deepcas:
          <article-title>An end-to-end predictor of information cascades</article-title>
          .
          <source>In WWW</source>
          , pages
          <volume>577</volume>
          {
          <fpage>586</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>H</given-names>
            <surname>Mei</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.M.</given-names>
            <surname>Eisner</surname>
          </string-name>
          .
          <article-title>The neural hawkes process: A neurally self-modulating multivariate point process</article-title>
          .
          <source>In NIPS</source>
          , pages
          <volume>6757</volume>
          {
          <fpage>6767</fpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In NIPS</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. L.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Cui</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Song</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>From micro to macro: Uncovering and predicting information cascading process with behavioral dynamics</article-title>
          .
          <source>In ICDM</source>
          , pages
          <volume>559</volume>
          {
          <fpage>568</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>