<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>User Modeling and Churn Prediction in Over-the-top Media Services</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ajith Pudiyavitil</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lowe's</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>USA ajithkp</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@gmail.com</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jaideep Chandrashekar Interdigital AI Lab</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Reference Format: Vineeth Rakesh, Ajith Pudiyavitil, and Jaideep Chandrashekar. 2020. User Modeling and Churn Prediction in Over-the-top Media Services. In 3rd Workshop on Online Recommender Systems and User Modeling (ORSUM 2020), in conjunction with the 14th ACM Conference on Recommender Systems</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vineeth Rakesh Interdigital AI Lab</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <abstract>
        <p>We address the problem of customer retention (churn) in applications installed on over the top (OTT) streaming devices. In the first part of our work, we analyze various behavioral characteristics of users that drive application usage. By examining a variety of statistical measures, we answer the following questions: (1) how do users allocate time across various applications?, (2) how consistently do users engage with their devices? and (3) how likely are dormant users liable to becoming active again? In the second part, we leverage these insights to design interpretable churn prediction models that learn the latent characteristics of users by prioritizing the specifications of the users. Specifically, we propose the following models: (1) Attention LSTM (ALSTM), where churn prediction is done using a single level of attention by weighting on individual time frames (temporal-level attention) and (2) Neural Churn Prediction Model (NCPM), a more comprehensive model that uses two levels of attentions, one for measuring the temporality of each feature and another to measure the influence across features (feature-level attention). Using a series of experiments, we show that our models provide good churn prediction accuracy with interpretable reasoning. We believe that the data analysis, feature engineering and modeling techniques presented in this work can help organizations better understand the reason behind user churn on OTT devices.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>In recent years, users have increasingly taken to consuming
streaming video services via applications (i.e. Netflix, Hulu, Youtube, etc.)
on so called over the top (OTT) platforms (i.e. AppleTV, Roku,
Amazon FireTV, etc.). Given the very large (and still growing)
number of streaming services, there is fierce competition to attract
new customers, while maintaining customer satisfaction.
Unfortunately, there is significant cost to attracting new users; thus, service
providers are very invested in retaining end-users and keeping them
engaged with their products. These customer retention eforts
focus on providing exclusive and engaging content, personalized
recommendations and intuitive user interfaces. When such eforts
fail, the operators experience customer churn, wherein a subscriber
∗This work was done when the author was at Interdigital AI Lab.
stops using the service. By examining high level data collected on
one such OTT hardware platform, we propose feature engineering
techniques for modeling user behavior and leverage these features
to develop application-level churn prediction models. Specifically,
given a user u who installs an application a at a given time on their
OTT device, our model predicts whether u will be engaged (or not
engaged) with a after a particular time window.</p>
      <p>Users may decide to abandon a streaming service for any number
of reasons such as a limited time budget to consume content, an
increasing afinity for a diferent application or a lack of compelling
new content. Yet another reason may be that the user experiences
more hardware faults (i.e. reboots, poor wifi, high memory
usage, etc.) when a particular service is being used. In such cases,
the user perceives these faults as being caused by the application
and quits out of frustration. The key observation here is that both
application-level and device-level behavior can influence user churn.
Consequently, it is critical to model these heterogeneous factors
along with temporally correlated features reflecting usage of
diferent applications to accurately predict churn. To achieve this, we first
analyze a dataset that captures high level events on OTT devices
and examine the signals of application churn. These events could be
user-initiated (i.e. opening or closing an application, restarting the
device, putting the device to sleep, etc.) or device-specific events
(i.e. automatic reboots, wifi drops, software/firmware updates, etc.)
With this data, we examine questions such as: (1) how users allocate
time across applications on their OTT device, (2) how often and for
how long users engage with the device (specific applications), and
(3) how long do users go dormant, and how likely is it for a dormant
user to become reactive. In the second part of our work, we leverage
these statistical insights to design interpretable models that are
efective in predicting churn across a wide range of scenarios.</p>
      <p>
        The naive approach to building a model that predicts churn
would be to fix an observation time window T , extract a number of
features of interest |m| from this window and then deploy a suitable
classification algorithm that predicts whether a subscriber will quit
a service after a period of time T . While this is entirely viable, there
are two important drawbacks. First, the data is inherently noisy
and high-dimensional; OTT devices send out periodic device-level
and application level summaries (the start time and duration of an
application session), and events observed on the box. This results
in a feature space of size T × |m|. Second, there is significant
inherent temporal correlation in the data; if a user spends a significant
amount of time inside an application on successive days, there is a
strong signal that he/she will engage with the same application on
the next day. Flattening data over the entire window T into a single
representation vector will lead to this information being lost. An
alternative, more principled approach, is to learn latent attributes
of the data using time-series models such as recurrent neural
networks (RNN) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and use them as features for churn prediction. A
potential issue with this approach is that the compressed latent
vector is ineficient at capturing all the necessary information that
leads to churn. Furthermore, it is extremely dificult to interpret
the results of vanilla RNNs.
      </p>
      <p>
        In this paper, we propose a model that address the drawbacks of
the more conventional approaches. First, we introduce Attention
LSTM (ALSTM), where we modify the neural machine
translation (NMT) model [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] for churn prediction; ALSTM models the
local attention. Second, we propose Neural Churn Prediction Model
(NCPM), which incorporates two levels of attention i.e., local and
global. ALSTM uses temporal-level attention (local attention) where
diferent (sub) observation windows contribute diferent weights
towards predicting churn. For example, in a particular week we
might observe the subscriber slowly starting to watch more and
more content on a new application, and spending less and less time
in another in which they were previously engaged (and
eventually abandon). Thus, features collected in this week might require
higher priority when compared to other time frames. NCPM on
the other hand is a more comprehensive model that captures each
feature using a separate ALSTM. The individual ALSTMs are then
combined with a feature-level attention layer (global attention).
Global attentions are much better at prioritizing weights across
diferent temporal features. For example, churning could be more
influenced by device-level issues such as periodic reboots rather
than application engagement. Although attention-based RNNs have
been extensively used in the field of natural language processing
[
        <xref ref-type="bibr" rid="ref25 ref7">7, 25</xref>
        ], they have very rarely been applied to the problem of churn
prediction. To our knowledge, the only work that appears to
address this area is [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]; however, the attention mechanism used in
their work is quite diferent than ours. We summarize the major
contributions of our work as follows:
• Understanding user behavior: Through data engineering and
statistical analysis, we provide several insights that explain the
behavior of users in our OTT dataset.
• Predicting Churn: We propose attention-based RNN models
that learn the characteristics of users in a weighted low-dimensional
latent space.
• State-of-the-art performance: By conducting extensive
experiments on a real-world dataset, we show that NCPM outperforms
all other models over diferent test cases and achieves an
accuracy of upto 89% and AUC of 92%. Additionally, NCPM interprets
the reason for churning by emphasizing on features such as
interarrival time between the apps and consistency in app usage.
      </p>
      <p>We begin by introducing our dataset in Section 2 and in Section 3,
we model the behavior of OTT customers as observed in this dataset.
The churn prediction models ALSTM and NCPM are proposed in
Section 4 followed by the results of our experiments in Section 5.
Finally, we review related work in Section 6 and conclude our paper
in Section 7.
2</p>
    </sec>
    <sec id="sec-2">
      <title>DATASET</title>
      <p>Our dataset consists of high level application session data (start
time and durations) and device level events that spans from Sept
2018 to April 2019. The data was collected from a sample of 31k</p>
      <p>AndroidTV based OTT devices deployed in homes and operated
by a large provider. Each of these devices comes pre-installed with
a set of applications; apart from these, users (u) can download
others from a large catalog on the app store. We observed over
3.7k distinct applications being used in the dataset. However, some
devices and applications are used very sporadically and account
for very little data. To remove these, we pre-processed the data
to filter out devices that were active for less than 90 days in total,
and removed applications that were used on fewer than 15 devices
in our population. This filtering resulted in 14,082 unique users
(i.e., OTT devices) 1 and 462 unique apps; this dataset is denoted
by D. Note that the churn prediction model builds features over
the lifetime of the application on the device; this starts from the
time the application was installed, to when the user is deemed to have
abandoned it. Unfortunately, for the pre-installed applications such
as Netflix, YouTube and SlingTV there is no install date. One can
exclude such apps from the data; however, this leads to removing a
significant number of users. This is because a large proportion of
users tend to confine themselves to using the pre-installed apps (which
also happens to be the popular streaming services). Alternatively,
having all users in a single bin could lead to some serious bias in
our modeling since for default apps, we do not have a clear signal
on when it was downloaded. The user might have been using the
app well before the onset of our study. Therefore, besides D, we
create a separate dataset De that completely excludes the default
apps, while for (u, a) ∈ D that do not have a start date, we simply
take the first log entry of a by u as a proxy for the actual install
date. The subset De covers 8,223 unique devices using 397 distinct
applications.</p>
      <p>Clearly, the earlier a service provider is able to predict churn
(of u), the more efectively they can take action and address the
underlying reasons as to why a customer might be departing.
Consequently, we divide D and De into diferent days of activities A,
where A = {t |t ≤ T , T ∈ {5, 10, 15, 20}}. For instance, T = 5 means
we consider a maximum of five days of user activity to predict
the churn. Table 1 shows the characteristics of dataset D and De
across diferent activity days. In the upcoming section, we explain
the methodology of determining the churners and non-churners
(i.e., columns four and five in Table 1). Here, one can see that as T
increases, the number of users decrease. This is because the number
of users that continuously use the OTT for say 20 days is far less
than than those who use for just 5 days.
3</p>
    </sec>
    <sec id="sec-3">
      <title>ANALYZING USER CHARACTERISTICS</title>
      <p>App and Device Usage: Figure 1 (a) shows the ten most popular
apps seen in our dataset, based on the number of users that regularly
use them. Sling Tv, Netflix, Youtube and Google games occupy the
top four spots. Figure 1 (b) plots the distribution of daily time spent
inside each of the applications. We find that users spend 3-4 hours
on average with the OTT device, with a very small fraction of
users that spend more than 8 hours. In fig. 1 (c), we break this daily
spend into four diferent parts of a day, corrsponding to morning
(5am-12pm), afternoon (12pm-5pm), evening (5pm-10pm) and night
(10pm-5am). Unsurprisingly, we observe that users tend to spend
more time in the evening when compared to other periods of a</p>
      <sec id="sec-3-1">
        <title>1we use the words devices and users interchangeably</title>
        <p>
          T
5
10
15
20
#Users
13082
7954
5440
5009
#NonChurns
16044
8596
5932
4067
#Users
6390
5500
4929
4453
day (median of 2.2 hrs). However, it is not significantly higher
than afternoon, which has a median of 1.8 hrs and morning with
a median of 1.4 hrs. Another key statistic of interest is the inter
arrival time between application sessions. Note that there may not
be an explicit indicator of a user quitting an application. Very often,
users just stop using the application that is installed, or deactivate
their accounts but keep the application installed. Thus, churn must
be detected implicitly, i.e., by the fact of the application not being
started by the user for a suficiently long time. We calculate the
arrival time between successive start times of an application session
on a device (across all applications) and plot the maximum values,
across all the devices, in fig. 1 (d). We observe that users, after an
absence, return to the OTT applications within a median time of
6 days (100-200 hours). The 75%-ile value of this distribution is
about 10 days. Later in this section, we leverage this to establish an
inactivity threshold when we define churn more precisely.
User Engagement Patterns: Here we try to understand how users
spend time on their OTT device. We wish to explore the following
aspects: (a) Are there users who consistently use the box for same
number of hours every day? (b) Are there dormant users who don’t
use their device for a while, but then reactivate it? (c) Are there
users that engage with their device, but only intermittently and for
brief periods of time? To answer these, we carry out the following
analysis. First, for each day that our dataset spans, we compute the
cumulative time (in terms of cummulative distribution function
CDF) that the user engaged with the device. Specifically, we
compute (i, ci ) for each user, where i = 1, 2, .., |D | represents each day
in our dataset, and ci is the total number of hours spent on the
device upto day i. Next, we carry out a non-linear fit on this data
for each user, recording the learned slope, intercept and standard
error as derived features for each user. Subsequently, we cluster
the derived features using K-means; the number of clusters is
determined based on the silhouette score [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. This analysis yields
four main behavior patterns that cover the vast majority of users,
and are depicted in fig. 2. Each plot is based on the original data of
cumulative device engagement time (x-axis is days elapsed, y-axis
is cumulative time spent). These four patterns can be labeled as
follows: (1) mid bloomers (a), these are users who are initially silent
and do not use the OTT box heavily, but then suddenly start using
during the middle phase of their total period. (2) late bloomers (b),
these users remain dormant for a longer duration with minimal
activity, but suddenly start using the device towards the end. (3)
potential churners (c), these are users who are of interest to us. As
explained by the plot, these users start using the device heavily at
ifrst, but then stop using the box for various reasons. Please note
that since y-axis is the CDF, the flat line here indicates minimal or
no activity (also indicated by very low standard deviation, since
there is no activity). (4) consistent users (d), finally, these are users
who are consistent and regularly use their OTT boxes to watch
diferent shows.
        </p>
        <p>Understanding Churning Behavior: Before introducing our
prediction models, we briefly explain how we label a user (or device)
having churned, i.e., left the service. This is fundamentally a dificult
task because there is no explicit signal for this behavior. Further
complicating things, (a) some users don’t use the app for a few
days, but return back after a brief period of inactivity, and (b) some
users simply download the app once (or spend a brief amount of
time in it) and never use it again. Figure 3 depics, at a high level,
all the information for a user (u) application (a). The dotted lines at
either end capture the time data was collected and each of the green
vertical lines in the middle indicate the start of application sessions
(a1 is the first session, an is the last). Here, we see that the
application was downloaded after the start of the data collection and used
several times, the last instance is at t3. Somewhat infrequently, we
see the device itself disappears from the dataset; we consider this a
signal that the user has disconnected the device and is no longer
using it. In this scenario, we capture this event having occured at t4.
With this depiction, we can now define churn in very specific terms
by addressing the two challenges previously discussed. First, we
require that the application not be used for a period of time after the
last use. Following the example in fig. 3, we impose the condition
∆ 2 ≥ T3q . Here T3q – an inactivity threshold – is the 3rd quartile
of the distribution in fig. 1 (d) and turns out to be 10 days. Second,
we require a minimum number of sessions to be recorded for an
application and user. Specifically, u should have engaged with a at
least K times, i.e., n ≥ K and we set K = 3. Figure 4 illustrates the
characteristics of devices that are exclusively labeled as churn. We
see that the median inter-arrival times for the top 5 apps is around
30 days (Figure 4 (b)), which is significantly higher than the generic
inter-arrival characteristics shown in Figure 1 (d). In figure 4 (c)
we notice that users who download more apps tend to have higher
churning rate, we obtained a Pearson correlation coeficient of 0.67.
It is also interesting to observe that as the churn increases, users
tend to switch between apps more frequently, where the app switch
is indicated by the session feature (y-axis).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4 PREDICTING CHURN</title>
      <p>
        Given a user u and an app a, our objective is to predict if u will
continue or stop using a. We realize each user-app entity as a tuple
(X , Mx , Y ) where X = {x1, ..., xt } is a stream of events (or logs)
that spans a time t ∈ T . Each event x comprises of M features,
and Y = {y1, ..., yt } are the binary labels that indicate churn (or
non-churn) at t . When designing our churn prediction model we
had two main objectives. First, since our data is highly temporal, it
is important to learn the latent characteristics of churners (and
nonchurners) in such a way that it embeds the temporality of events.
Second, not only should we predict the churn with good accuracy,
but also produce highly interpretable results. In other words, we
should reason out as to why a user is churning. To achieve this, we
propose the following models: (1) attention LSTM (ALSTM), which
is a simple modification of the neural machine translation (NMT)
model [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and (2) neural churn prediction model (NCPM): a more
comprehensive model that incorporates temporal-level attention
(or local attention) and feature-level attention (or global attention).
Both the models are based on recurrent neural network (RNNs)
that have shown to be efective in modeling time-series data [
        <xref ref-type="bibr" rid="ref5 ref8">5, 8</xref>
        ].
RNNs take a series of temporally dependent inputs and learn their
latent representation (or hidden state vector) using the following
expression:
      </p>
      <p>
        ht = f (ht −1, xt )
where ht is the hidden layer at time t and f is some non linear
function. For our application, we model f using long short-term
memory network (LSTM) [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. LSTM has four states that are defined
as follows:
it = σ (Wi · [ht −1; xt ] + bi )
ft = σ (Wf · [ht −1; xt ] + bf )
ct = ft × ct −1 + it × tanh(Wc · [ht −1; xt ] + bc )
ot = σ (Wo · [ht −1, xt ] + bo )
ht = ot × tanh(ct )
where t is the time step (i.e, days) , ht is the hidden state at t , ct
is the cell state at t , xt is the hidden state of the previous layer at
time t , it , ft , ot are the input, forget and out gates, respectively.
(1)
(2)
      </p>
      <p>
        Attention LSTM (ALSTM): Obviously one can predict churn by
simply providing the input X to a vanilla LSTM, get the latent
vectors h from the final layer, and use them as features for prediction.
A potential issue with this approach is that the compressed (or
low dimensional) latent vector h is ineficient in capturing all the
necessary information that attributes to churn. As explained in
Section 1, in a particular week we might observe the subscriber
slowly starting to navigate towards a new application and spend less
and less time in one that they were previously engaged with (and
eventually abandon). So, it is important to give high priority to these
time windows when compared to other weeks. Modeling churn
using vanilla LSTM networks fails to prioritize such key events.
Inspired by recent developments in neural machine translation
(NMT) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], we incorporate attentions into LSTM to overcome this
issue. Since our application is very diferent from natural language
processing, we introduce two modifications over NMT. First, we
replace the decoder part with a single layer neural network (NN)
with sigmoid activation for churn prediction and second, we change
      </p>
      <p>(6)
(7)
(8)
the attention mechanism to suit our problem. The proposed ALSTM
model is shown in Figure 5 (a). Here, the attention block A outputs
a vector of weights α that emphasizes the importance of the latent
vector h for a given time frame t . The weighted latent vector p is
defined as follows:
In the above expression, β denotes the individual attention weights
that is defined as follows:
βm =
cmj =</p>
      <p>exp (cm )
PM</p>
      <p>m=1 exp (cm )
Z
X ziUi j
i=1
where the weight αj for a time instance t is defined by
p =</p>
      <p>T
X αt ht
t =1
αj =
sj =</p>
      <p>exp (sj )
PTt=1 exp (st )
K
X hkt · Wktj
k=1</p>
      <sec id="sec-4-1">
        <title>Neural churn prediction model (NCPM): One drawback of AL</title>
        <p>STM is that it is unable to prioritize across features. For example,
churning could be more influenced by the consistency of users (see
Figure 2), while the number of downloads might have little
importance. To overcome this problem, we incorporate both
temporallevel attention and feature-level attention. As depicted in Figure
5, instead of treating the features as a single vector, we decouple
the features and model them using individual ALSTMs. Similar
to ALSTM, the attention block pm of an a feature m captures the
influence (or weight) of the latent features from diferent slices of
time. On the other hand, the feature-level attention is capture by
the block B, which is defined by the following expressions:
д =</p>
        <p>M
X βmpm
m=1
(3)
(4)
(5)
where z = p1 ⊕ {pi }m2 is the concatenation (indicated by ⊕) of the
feature-level latent vectors . Finally, to predict the churn, a linear
projection with a sigmoid function is connected to the output of
the last layer to produce user churn prediction as follows:
yˆ = σ (Wд · д + bд )
The loss for both ALSTM and NCPM is computed using binary
cross entropy, that is defined as follows:</p>
        <p>L = X −yi loд(yˆi ) − (1 − yi )loд(1 − yˆi ) (9)</p>
        <p>i
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTS</title>
      <p>Obviously, from Table 1, one can notice that our dataset is biased
towards negative samples (i.e., #non-churns). Therefore, to create a
balanced dataset, for every positive data point (i.e, churns) for an
app a, we randomly sample a corresponding negative data point.
We test our models by varying the number of days in the training
sample (explained in Section 2). This helps us to see how quickly
can our models predict the churn. For all our experiments, we use
10 fold cross validation, where eight folds are used for training, one
for validation, and one for testing. The deep learning models are
implemented using Keras with Tensorflow as the back-end.
5.1</p>
    </sec>
    <sec id="sec-6">
      <title>Baselines</title>
      <p>We compare the performance of the proposed ALSTM and NCPM
with three baseline methods. Unlike the proposed models (i.e.,
ALSTM and NCPM) the following baselines do not capture the
temporality in data. Therefore, the inputs to these model are flattened
vectors across the time frames.
Logistic Regression: the classic model for binary classification
problem. Albeit simplistic, it helps us to understand if a linear
decision boundary is suficient to capture the churners and the
non-churners. We use L2 norm as the regularizer and Stochastic
Average Gradient (SAG) as the solver.</p>
      <sec id="sec-6-1">
        <title>Multi-layer Perceptron (MLP): We consider a simple two layer</title>
        <p>
          neural network and a dropout layer to avoid overfitting. A linear
projection with a sigmoid function is connected to the output of
the last layer to produce user churn. Similar to logistic regression,
MLP does not capture the temporal dependencies between the data.
The number of neurons are set to 80 for each intermediate layer.
Random Forest (RF): Despite the rapid advancements in the field
of deep learning, ensemble techniques such as RF [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] remain highly
competitive in producing excellent results on data with several
modalities. In our experiments, the number of decision trees are set
as 50 and the maximum depth as 10.
5.2
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Results</title>
      <p>Classification accuracy : Tables 2-3 shows that NCPM
outperforms all other models for both datasets De and D, achieving an
accuracy of upto 92%. As we increase the number of days the
accuracy increases for all models (except logistic regression). Here, CA-5
implies the classification accuracy with just 5 days of data, while
CA-20 implies 20 days of data. We can also see that the proposed
ALSTM is not as good as NCPM which proves the following: (1) it
is important to learn the latent attributes of each individual features
separately and (2) incorporating both global and local attention
is necessary. That being said, ALSTM clearly outperforms MLP,
which emphasizes the necessity of learning the temporal actions of
OTT users. The worst performing model is the logistic regression,
which is just slightly better than a random selection. This illustrates
the dificulty of our churn prediction task. The performance of RF is
very close to that of ALSTM, which indicates that ensemble models
are still a strong candidate for our problem.</p>
      <p>In general, the performance of models over non-continuous data
is much better than its continuous counterpart, this can be explained
using the following example. let us say that u uses an app for
six days before churning and we have the following data for u
{m1, m4, m6, m7, m8, m10}, where m is some feature and the sufix
indicates the day. Our objective is to predict the outcome on sixth
day, using the first five days. Since the user does not use the OTT box
for days two, three, and five, the continuous data that is fed to our
models (i.e., both NCPM and ALSTM) is essentially a sparse vector
{m1, 0, 0, m4, 0}, which has several missing values. On the contrary,
for non-continuous dataset, we will have the actual usage values
for five days. This obviously means that the model gets to train
with more observed data points, which leads to better prediction
accuracy. Another interesting observation is that the performance
of all models (except logistic regression) is noticeably better on the
all-apps dataset. One key reason for this outcome is the popularity
of the default apps. Apps such as Sling TV, Netflix and Youtube
are significantly popular than other non-default apps. Therefore,
the models are able to efectively learn the churn patterns for such
apps more efectively.</p>
      <p>AUC and ROC characteristics: Figures 6 and 7 compare the ROC
characteristics of the proposed models with other baselines. The
corresponding AUC values are listed in Table 4, due to space
constraints, only the non-default case is furnished. Similar to the
accuracy scores, for most scenarios, NCPM remains dominant over
other models. We also notice that RF tends to perform better than
NCPM and ALSTM when the temporal length of data is low (i.e.,
with just five days). However, as we incorporate more days for
training, there is a clear increase in the performance of our models.
The outcome for dataset D is much diferent than De where we
are able to achieve an AUC of almost 89% with just five days of
data; additionally, ALSTM seems to perform very similar to NCPM.
Interpreting the churn prediction: One of the key strengths of
our model is interpretability. As explained in Section 1, ALSTM
provides single level of interpretability, which indicates which days
are important when predicting churn. NCPM on the other hand, has
two levels; besides telling the important days, it also tells us which
features are important. We present the interpretability scores as
heatmaps in Figure 8. Due to the lack of space, we only furnish the
results of non-continuous dataset. Heat maps (a)-(d) explains that
the influence of features are not uniform across apps; for instance,
when we have less days for prediction, the churn is influenced
by two main attributes namely, the number of downloads and the
cluster types (Figures 8 (a) and (c)). As we incorporate more data
for training (i.e., the number of days) the attention tends to get
more focused towards a few key features. For non-default apps,
there seems to be more attention on the inter-arrival time, while
for all-apps the influence seems to be more towards the number
of reboots. It could be possible that these apps experience a higher
number of app crashes, which could lead to user rebooting the
device. For non-default apps, Hulu, Plot Tv and Kodi is heavily
influenced by the cluster id feature that we engineered in Section 3.
When it comes to temporal attention (Figures 8 (e)-(h)), for dataset
D, the influence is mainly concentrated on a few selective days,
i.e., day 4 for non-continuous case, and day 2 for continuous case.
Contrary to this, for De this influence is spread across almost all
days.</p>
      <p>In Section 3 we explained that the engagement pattern of users
could have a strong impact on churn. To show this efect, for each
user, we get the final attention score from individual RNNs of
NCPM and plot the outcome in Figure 9 (a). Here, we can see
that consistent users are the highest indicators of non-churn, while
potential churners are the highest indicators of churn. Interestingly,
mid bloomers seem to have higher attention over late bloomers
when it comes to predicting non-churners, while the opposite is
true for churners. Figure 9 (b) provides a more deeper look into this
outcome by emphasizing on the importance of temporal progression
on the user types. Unsurprisingly, during the initial phase (elapsed
duration of 10-20%) almost all types have less attention weights.
This is because, during the early phase, we do not have enough
data about the user type. As the time progresses, around 20-50%
of the elapsed duration, we see that potential churners have the
strongest impact on the outcome followed by consistent users and
late bloomers. Around 50-80% , the impact of potential churners
drastically reduces, while mid and late bloomers increase. At the
ifnal stage (i.e., 80-100%) almost all user types have less importance
(or attention). This is because, during the last phase, there is more
available data in the form of other features such as number of</p>
      <p>Model
Logistic
RF
MLP
ALSTM
NCPM
downloads and inter-arrival time between apps. Consequentially,
the model is able to rely on better indicators at the later stage.</p>
    </sec>
    <sec id="sec-8">
      <title>6 RELATED WORK</title>
      <p>The problem tackled in this paper is related to the following topics:
(1) churn prediction (2) user behavior modeling and (3) interpretable
neural networks. We now detail some existing research that
correspond to these topics.</p>
      <p>
        Churn Prediction: User retention (or churn) has been extensively
studied in the field of social computing and human computer
interaction (HCI) [
        <xref ref-type="bibr" rid="ref14 ref27 ref9">9, 14, 27</xref>
        ]. However, developing predictive models for
churn is still at infancy. Au et. al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] adopt a rule based learning
technique for early churn prediction. In [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], the authors tackle
the problem of churn prediction in mobile apps. They find that
application performance such as energy consumption and latency
have a significant impact on retention. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] use a social influence
based approach for churn prediction. Recently, [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] develop an
interpretable framework that constraints the objective of RNN with
the outcome of K-means clustering to predict the retention of users
in Snap Chat.
      </p>
      <p>
        Modeling user behavior: There are numerous research on
behavior modeling [
        <xref ref-type="bibr" rid="ref10 ref15 ref21">10, 15, 21</xref>
        ]. For example, [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] predict user intents
      </p>
      <sec id="sec-8-1">
        <title>Model</title>
        <p>Logistic</p>
        <p>RF</p>
        <p>
          MLP
ALSTM
NCPM
by leveraging the activity logs in Pinterest. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] predict the
likelihood of a successful search in web search queries. They show
that user behavior are more predictive of goal success than those
using document relevance.[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] model the behavior of users in the
Kickstarter crowdfunding domain using a heterogeneous
combination of social communities, popularity of projects and the impact
of reward categories. Studies such as [
          <xref ref-type="bibr" rid="ref11 ref24">11, 24</xref>
          ] and [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] model user
behavior from sequencial actions such as click streams and social
network activities. [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ] use a combination of Mahalanobis distance
(for detecting outlines) and Markov Chains to model sessions in
click streams, while [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] use a temporal LDA based approach for
tour recommendation in Foursquare.
        </p>
        <sec id="sec-8-1-1">
          <title>Interpretable Sequence Modeling: RNNs have become the state</title>
          <p>
            of-the-art technique for sequential modeling [
            <xref ref-type="bibr" rid="ref13 ref5">5, 13</xref>
            ]. Albeit a plethora
of research in the NLP domain [
            <xref ref-type="bibr" rid="ref19 ref2">2, 19</xref>
            ], extending interpretable RNNs
for other real world applications is still an emerging field of research.
In a recent work, [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] predict the engagement of users in the Snap
Chat app by capturing the in-app action transition patterns as a
temporally evolving action graph. [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] develop an interpretable
LSTM to learn multi-level graph structures in a progressive and
stochastic manner. [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ] propose a dual stage attention model for
medical diagnostics such as heart failure prediction. Zhou et. al.
[
            <xref ref-type="bibr" rid="ref28">28</xref>
            ] propose an attention-based RNN that predicts the purchase
probability of users for targeted ads. Albeit having a similar NN
architecture as ours, their problem is quite diferent than churn
prediction. Additionally, they modeling of local and global
attention is quite diferent than ours. To the best of our knowledge, the
only research that closely resembles our work is the churn
prediction model proposed by Yang et. al. [
            <xref ref-type="bibr" rid="ref26">26</xref>
            ]. However, the attention
mechanism used in their work is quite diferent than ours.
7
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>CONCLUSION</title>
      <p>In this paper we proposed interpretable recurrent neural network
based models for prediction churn in over the top media (OTT)
devices. In the first part of the paper, we analyzed the behavioral
characteristics of users and found that they can be categorized into
four main types: mid bloomer, late bloomers, potential churners
and consistent users. In the second part, we introduced two models
for churn prediction, namely Attention LSTM (ALSTM) and Neural
Churn Prediction Model (NCPM). In ALSTM, the prediction of
churn was done by weighting on individual time frames
(temporallevel attention) and (2) NCPM, we used two levels of attentions
namely, feature-level and temporal-level. We showed that NCPM
outperforms all other models over a wide range of test cases and
achieves an accuracy of upto 89% and AUC of 92%.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Wai-Ho</surname>
            <given-names>Au</given-names>
          </string-name>
          ,
          <article-title>Keith CC Chan, and</article-title>
          <string-name>
            <given-names>Xin</given-names>
            <surname>Yao</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>A novel evolutionary data mining algorithm with applications to churn prediction</article-title>
          .
          <source>IEEE transactions on evolutionary computation 7</source>
          ,
          <issue>6</issue>
          (
          <year>2003</year>
          ),
          <fpage>532</fpage>
          -
          <lpage>545</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Dzmitry</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          , Kyunghyun Cho, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
          <source>arXiv preprint arXiv:1409.0473</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          , Patrice Simard,
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Frasconi</surname>
          </string-name>
          , et al.
          <year>1994</year>
          .
          <article-title>Learning long-term dependencies with gradient descent is dificult</article-title>
          .
          <source>IEEE transactions on neural networks 5</source>
          ,
          <issue>2</issue>
          (
          <year>1994</year>
          ),
          <fpage>157</fpage>
          -
          <lpage>166</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Leo</given-names>
            <surname>Breiman</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Random forests</article-title>
          .
          <source>Machine learning 45, 1</source>
          (
          <year>2001</year>
          ),
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Zhengping</given-names>
            <surname>Che</surname>
          </string-name>
          , Sanjay Purushotham, Kyunghyun Cho, David Sontag, and Yan Liu.
          <year>2018</year>
          .
          <article-title>Recurrent neural networks for multivariate time series with missing values</article-title>
          .
          <source>Scientific reports 8</source>
          ,
          <issue>1</issue>
          (
          <year>2018</year>
          ),
          <fpage>6085</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Justin</surname>
            <given-names>Cheng</given-names>
          </string-name>
          , Caroline Lo, and
          <string-name>
            <given-names>Jure</given-names>
            <surname>Leskovec</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Predicting intent using activity logs: How goal specificity and temporal range afect user behavior</article-title>
          .
          <source>In Proceedings of the 26th International Conference on World Wide Web Companion. International World Wide Web Conferences Steering Committee</source>
          ,
          <fpage>593</fpage>
          -
          <lpage>601</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Kyunghyun</given-names>
            <surname>Cho</surname>
          </string-name>
          , Bart Van Merriënboer,
          <string-name>
            <surname>Caglar Gulcehre</surname>
            , Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and
            <given-names>Yoshua</given-names>
          </string-name>
          <string-name>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Learning phrase representations using RNN encoder-decoder for statistical machine translation</article-title>
          .
          <source>arXiv preprint arXiv:1406.1078</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Edward</given-names>
            <surname>Choi</surname>
          </string-name>
          , Andy Schuetz, Walter F Stewart,
          <string-name>
            <given-names>and Jimeng</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Using recurrent neural network models for early detection of heart failure onset</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>24</volume>
          ,
          <issue>2</issue>
          (
          <year>2016</year>
          ),
          <fpage>361</fpage>
          -
          <lpage>370</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Luca</surname>
          </string-name>
          Ciampaglia and
          <string-name>
            <given-names>Dario</given-names>
            <surname>Taraborelli</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>MoodBar: Increasing new user retention in Wikipedia through lightweight socialization</article-title>
          .
          <source>In Proceedings of the 18th ACM Conference on Computer Supported Cooperative Work &amp; Social Computing. ACM</source>
          ,
          <volume>734</volume>
          -
          <fpage>742</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Gideon</surname>
            <given-names>Dror</given-names>
          </string-name>
          , Dan Pelleg, Oleg Rokhlenko, and
          <string-name>
            <given-names>Idan</given-names>
            <surname>Szpektor</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Churn prediction in new users of Yahoo! answers</article-title>
          .
          <source>In Proceedings of the 21st International Conference on World Wide Web. ACM</source>
          ,
          <volume>829</volume>
          -
          <fpage>834</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Şule</given-names>
            <surname>Gündüz</surname>
          </string-name>
          and
          <string-name>
            <given-names>M Tamer</given-names>
            <surname>Özsu</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>A web page prediction model based on click-stream tree representation of user behavior</article-title>
          .
          <source>In Proceedings of the ninth ACM SIGKDD international conference on Knowledge discovery and data mining. ACM</source>
          ,
          <volume>535</volume>
          -
          <fpage>540</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Ahmed</surname>
            <given-names>Hassan</given-names>
          </string-name>
          , Rosie Jones, and Kristina Lisa Klinkner.
          <year>2010</year>
          .
          <article-title>Beyond DCG: user behavior as a predictor of a successful search</article-title>
          .
          <source>In Proceedings of the third ACM international conference on Web search and data mining. ACM</source>
          ,
          <volume>221</volume>
          -
          <fpage>230</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9</source>
          ,
          <issue>8</issue>
          (
          <year>1997</year>
          ),
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Selim</surname>
            <given-names>Ickin</given-names>
          </string-name>
          , Katarzyna Wac, Markus Fiedler, Lucjan Janowski,
          <string-name>
            <surname>Jin-Hyuk Hong</surname>
          </string-name>
          , and
          <article-title>Anind</article-title>
          K Dey.
          <year>2012</year>
          .
          <article-title>Factors influencing quality of experience of commonly used mobile applications</article-title>
          .
          <source>IEEE Communications Magazine</source>
          <volume>50</volume>
          ,
          <issue>4</issue>
          (
          <year>2012</year>
          ),
          <fpage>48</fpage>
          -
          <lpage>56</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Marcel</surname>
            <given-names>Karnstedt</given-names>
          </string-name>
          , Matthew Rowe, Jefrey Chan,
          <string-name>
            <given-names>Harith</given-names>
            <surname>Alani</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Conor</given-names>
            <surname>Hayes</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>The efect of user features on churn in social networks</article-title>
          .
          <source>In Proceedings of the 3rd International Web Science Conference. ACM</source>
          ,
          <volume>23</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Jaya</surname>
            <given-names>Kawale</given-names>
          </string-name>
          , Aditya Pal, and
          <string-name>
            <given-names>Jaideep</given-names>
            <surname>Srivastava</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Churn prediction in MMORPGs: A social influence based approach</article-title>
          .
          <source>In 2009 International Conference on Computational Science and Engineering</source>
          , Vol.
          <volume>4</volume>
          . IEEE,
          <fpage>423</fpage>
          -
          <lpage>428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Xiaodan</surname>
            <given-names>Liang</given-names>
          </string-name>
          , Liang Lin,
          <string-name>
            <given-names>Xiaohui</given-names>
            <surname>Shen</surname>
          </string-name>
          , Jiashi Feng, Shuicheng Yan, and Eric P Xing.
          <year>2017</year>
          .
          <article-title>Interpretable structure-evolving LSTM</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>1010</fpage>
          -
          <lpage>1019</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Yozen</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Xiaolin Shi,
          <string-name>
            <given-names>Lucas</given-names>
            <surname>Pierce</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xiang</given-names>
            <surname>Ren</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Characterizing and Forecasting User Engagement with In-app Action Graph: A Case Study of Snapchat</article-title>
          . arXiv preprint arXiv:
          <year>1906</year>
          .
          <volume>00355</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Tomáš</surname>
            <given-names>Mikolov</given-names>
          </string-name>
          , Martin Karafiát, Lukáš Burget, Jan Černock y`, and
          <string-name>
            <given-names>Sanjeev</given-names>
            <surname>Khudanpur</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Recurrent neural network based language model</article-title>
          .
          <source>In Eleventh annual conference of the international speech communication association.</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Yao</surname>
            <given-names>Qin</given-names>
          </string-name>
          , Dongjin Song, Haifeng Chen, Wei Cheng, Guofei Jiang, and
          <string-name>
            <given-names>Garrison</given-names>
            <surname>Cottrell</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>A dual-stage attention-based recurrent neural network for time series prediction</article-title>
          .
          <source>arXiv preprint arXiv:1704.02971</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Vineeth</surname>
            <given-names>Rakesh</given-names>
          </string-name>
          , Niranjan Jadhav, Alexander Kotov, and
          <article-title>Chandan</article-title>
          K Reddy.
          <year>2017</year>
          .
          <article-title>Probabilistic social sequential model for tour recommendation</article-title>
          .
          <source>In Proceedings of the Tenth ACM International Conference on Web Search and Data Mining. ACM</source>
          ,
          <volume>631</volume>
          -
          <fpage>640</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Vineeth</surname>
            <given-names>Rakesh</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang-Chien Lee</surname>
          </string-name>
          , and
          <article-title>Chandan</article-title>
          K Reddy.
          <year>2016</year>
          .
          <article-title>Probabilistic group recommendation model for crowdfunding domains</article-title>
          .
          <source>In Proceedings of the Ninth ACM International Conference on Web Search and Data Mining. ACM</source>
          ,
          <volume>257</volume>
          -
          <fpage>266</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Peter J Rousseeuw</surname>
          </string-name>
          .
          <year>1987</year>
          .
          <article-title>Silhouettes: a graphical aid to the interpretation and validation of cluster analysis</article-title>
          .
          <source>Journal of computational and applied mathematics 20</source>
          (
          <year>1987</year>
          ),
          <fpage>53</fpage>
          -
          <lpage>65</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Narayanan</given-names>
            <surname>Sadagopan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jie</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Characterizing typical and atypical user sessions in clickstreams</article-title>
          .
          <source>In Proceedings of the 17th international conference on World Wide Web. ACM</source>
          ,
          <volume>885</volume>
          -
          <fpage>894</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Ilya</surname>
            <given-names>Sutskever</given-names>
          </string-name>
          , Oriol Vinyals, and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Sequence to sequence learning with neural networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>3104</volume>
          -
          <fpage>3112</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Carl</surname>
            <given-names>Yang</given-names>
          </string-name>
          , Xiaolin Shi,
          <string-name>
            <given-names>Luo</given-names>
            <surname>Jie</surname>
          </string-name>
          , and Jiawei Han.
          <year>2018</year>
          .
          <string-name>
            <given-names>I</given-names>
            <surname>Know</surname>
          </string-name>
          <article-title>You'll Be Back: Interpretable New User Clustering and Churn Prediction on a Mobile Social Application</article-title>
          .
          <source>In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining. ACM</source>
          ,
          <volume>914</volume>
          -
          <fpage>922</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Igor</surname>
            <given-names>Zakhlebin</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Em</given-names>
            <surname>Horvát</surname>
          </string-name>
          , et al.
          <year>2019</year>
          .
          <article-title>Investor Retention in Equity Crowdfunding</article-title>
          .
          <source>In Proceedings of the 10th ACM Conference on Web Science. ACM</source>
          ,
          <volume>343</volume>
          -
          <fpage>351</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Yichao</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Shaunak Mishra, Jelena Gligorijevic, Tarun Bhatia, and
          <string-name>
            <given-names>Narayan</given-names>
            <surname>Bhamidipati</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Understanding Consumer Journey using Attention based Recurrent Neural Networks</article-title>
          .
          <source>In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining. ACM</source>
          ,
          <volume>3102</volume>
          -
          <fpage>3111</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Agustin</surname>
            <given-names>Zuniga</given-names>
          </string-name>
          , Huber Flores, Eemil Lagerspetz, Petteri Nurmi, Sasu Tarkoma, Pan Hui, and
          <string-name>
            <given-names>Jukka</given-names>
            <surname>Manner</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Tortoise or Hare? Quantifying the Efects of Performance on Mobile App Retention</article-title>
          .
          <source>In The World Wide Web Conference. ACM</source>
          ,
          <volume>2517</volume>
          -
          <fpage>2528</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>