<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Telecommunication customer churn prediction using machine learning methods</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Monika Zdanavičiūtė</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rūta Juozaitien</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomas Krilavičiu</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centre for Applied Research andDevelopment</institution>
          ,
          <country country="LT">Lithuania</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vilnius University</institution>
          ,
          <addr-line>Vilnius</addr-line>
          ,
          <country country="LT">Lithuania</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Vytautas Magnus University, Faculty of Informatics</institution>
          ,
          <addr-line>Vileikos street 8, LT-4440K4aunas</addr-line>
          ,
          <country country="LT">Lithuania</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>These days telecommunication sector has grown significantly due to the use of smart technologies, and it is likely to continue to grow. The main resource of telecommunications companies is customers, but due to the relatively high level of competition in this field, most customers are not tied to a single service company. To understand the key factors contributing to customer churn rate, we have analysed the real data of one telecommunication company. The data from 2020-01-01 to 2022-03-07 consisted of information on 21128 users, 140970 payments and 350379 calls. The main contribution of our work was to develop a churn prediction model which identifies customers who are most likely subject to churn. We performed experiments using k-nearest neighbours, support vector machine, decision trees, random forest, naive Bayes classifiers and the Cox proportional hazard model with time-varying covariates. Results showed that the Cox regression model with time-varying covariates was superior to classical classification methods because it can take into account static user parameters and reflect their changes over time.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Churn prediction</kwd>
        <kwd>telecommunication churn</kwd>
        <kwd>survival analysis</kwd>
        <kwd>churn</kwd>
        <kwd>telecommunications</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>IVUS 2022: 27th International Conference onInformation
Technology, May 12, 2022, Kaunas, Lithuania
$ monika.zdanaviciute@vdu.lt (M. Zdanavičiūtė);
ruta.juozaitiene@vdu.lt (R. Juozaitienė);
tomas.krilavicius@vdu.lt (T. Krilavičius)
© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).</p>
      <p>CEUR Workshop Proceedings (CEUR-WS.org)
into two groups according to this variable. Accu- arates customers into two groups.
racy and F-score measures were used to assess the The study [6] uses telecommunication users data
accuracy of the models, which showed that the XG- for customer analysis, which stores basic user
inforBoost method was the best classifier. In addition, mation (age, gender, etc.), plan order information
this method was used to find out which variables (payment method, monthly fee, full-time fee, etc.).
most influence customer exit. This study found It also provides information about the services
that customers with higher monthly charges are (telephones, internet, television, insurance, etc.)
more likely to churn. and information on whether the customer is
ac</p>
      <p>Data from the SyriaTel telecommunications ser- tive or has already churned. Clustering (k-Means,
vice provider were used for the [4] study. The DBSCAN) and classification methods (Multi-Layer
analyzed period is 9 months (about 10 million Perceptron, Back Propagation algorithm, Decision
users), and the available information includes data Trees, Logistic Regression, Support Vector
Maabout client (age, gender, place of residence, type chine) were used to analyze this data. The
classifiof contract concluded, services received), his ac- cation models were evaluated with several measures
tions (calls, messages, and internet usage), mobile of Precision and Accuracy, and the Back
Propadevice (device type, brand, model), and telecom- gation algorithm and Multi-Layer Perceptron best
munications tower infrastructure. To better de- predicted the customer’s retraction. When
clusterscribe users, the available data was used to create ing analysis was applied to groups, better active
a social network for all customers and to calculate and inactive clients were separated using the
DBvariables such as degree centrality measures, simi- SCAN algorithm.
larity values, and customer’s network connectivity The customer loyalty task is usually formulated
for each customer. For model training and testing, as a classification task, the data set of which
condata were separated into training (70%) and testing sists of active and churned users. To solve this
(30%) sets. Because the data sets were unbalanced problem, the literature suggests the use of –
(there were significantly more outgoing customers Nearest Neighbors [3], Neural Networks [5][6],
Supthan active ones), the classification was done in port Vectors [6][5], and Bayesian classifiers [5].
Detwo ways: by balancing the sets and applying the cision Tree, Random Forest, and XGBoost
algodata as it is. Using Decision Trees, Random For- rithms can also be used to analyze customer loyalty
est, GBM (Gradient Boosted Machine Tree), and [3][4]. However, there are also cases when this
probXGBoost algorithms, customers were classified into lem is solved by applying clustering methods, e.g.
two classes: churned and existing customers. The Genetic [2] or –means algorithm [6]. This type of
AUC (Area under curve) was used to determine ac- task is usually based on user information, as well
curacy. The obtained results showed that the XG- as payment history and calls data. Research shows
Boost algorithm classifies customers best according that customers who pay more for services tend to
to the available data. change telecommunications operators.</p>
      <p>The data set for the study [5] consists of call
records obtained from the University of California,
Department of Information and Computer Science. 3. Methods
The data set provides information on the use of the
3333 customer mobile system, which consists of 15 1. k-Nearest Neighbors is an algorithm that
quantitative, 5 categorical variables and a binary stores all available cases and classifies new
variable, describing whether the customer has left cases based on a similarity measure
(disthe customer base of the telecommunications ser- tance functions). Euclidian distance
funcvice provider. In the analysis of the available call tion [7]:
data, each user is assigned variables describing his ⎯⎸ 
call habits and, using classification methods, these ⎸⎷∑︁( − )2 (1)
customers are divided into two classes according =1
to said binary variable. The research uses Neural 2. Support Vector Machine (SVM) performs
Networks, Support Vector Machine and Bayesian classification by finding the hyperplane that
classification methods. The data set is divided into maximizes the margin between the two
training (80% of all data) and testing (20% of all classes [7]. Hyperplane equation:
data) sets so that the training set is balanced. Then
95% of the customers who leave and 5% of the ex-   +  = 0 (2)
isting ones remain in the testing set. The study
revealed that the support vectors method best sep- To define an optimal hyperplane we need to
maximize the width of the margin ():
max</p>
      <p>2
‖‖
3. Decision Tree is a flowchart-like structure in
which each internal node represents a "test"
on an attribute, each branch represents the
outcome of the test, and each leaf node
represents a class label [8]. A quantitative
measure of randomness, entropy, is used to select
a feature in a node. The initial entropy of
the set :
() = − ∑︁  (|) log2  (|), (4)</p>
      <p>4. Data set
where</p>
      <p>∈
mean 1, . . . ,  entropy after division:
 (|) = | :  ∈ ,  ∈ | ,</p>
      <p>||

(, ) = ∑︁  (|)(),</p>
      <p>=1
where
 (|) =  . (7)</p>
      <p>4. Random Forest is an ensemble learning
method for classification tasks that operates
by constructing a multitude of decision trees
at training time. The output of the random
forest is the class selected by most trees [9].
5. Na¨ıve Bayes classifier assume that the
effect of the value of a predictor () on a
given class () is independent of the values
of other predictors [7]. This assumption is
called class conditional independence.</p>
      <p>(|) =
 (|) ()
 ()
.</p>
      <p>(8)</p>
    </sec>
    <sec id="sec-2">
      <title>6. Cox proportional hazard model with time</title>
      <p>varying covariates is method for
investigating the efect of several variables upon the
time a specified event takes to happen. In a
Cox proportional hazards regression model,
the measure of efect is the hazard rate.</p>
      <p>Hazard function for individual :
ℎ() = ℎ0 exp(11 + 22 + . . . +  )
(9)
where ℎ0() is the baseline hazard function,
1 , 2 . . . . ,  – covariates, 1, 2, . . . , 
– regression coeficients.
(3)
(5)
(6)
• TP (true positive) - The user is
expected not to churn and he remains.
• TN (true negative) - The user is
expected to churn and he churns.
• FP(false positive) - The user is
expected to remain but he churns.
• FN (false negative) - The user is
expected to churn but remains.</p>
      <p>The analyzed data consists of three data sets in the
range from 2020-01-01 to 2022-03-07:
1. Users data set. Individual users
information, which includes demographic and other
data provided during registration. This
study analyzes 21128 users.
2. Payments data set. 140970 payment records
showing when and what type of plan was
purchased and how much it cost. There are
two types of plans: monthly and yearly.
3. CDR (call detail record) data set. A
realtime data records documenting telephone
calls or other telecommunications operations
(3350379 records).</p>
    </sec>
    <sec id="sec-3">
      <title>After the data transformations, a list of variables describing the users was created (Table 1).</title>
      <sec id="sec-3-1">
        <title>5. Churn definition</title>
        <p>In order to assess the risk of customer’s churn, the
definition of churn must first be de fined. Since user
leaving the customer base can be described in
several ways, it is necessary to monitor client behavior
and changes in activity and decide which definition
best describes churn. In the study, user churn is
described in two diefrent w ays. Diferent
problemsolving methods are used for each of these two
options.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>1. The user is classified as a churned customer</title>
      <p>if he has not purchased a new plan 35 days
after the first plan purchase.</p>
      <p>Figure 1 shows a bar graph showing the
distribution of the number of plans purchased
by customers. It shows that most customers
have only bought one plan.
The distribution of intervals between plan
orders for users who have purchased more
than one plan is shown in Figure 2. It shows
that most plans are ordered every 30 days, in
other words, most plans are ordered on a
regular monthly basis. There are also some
users who order multiple plans on the same
day. The data set for the classification
models consists of variables describing user
behavior (Table 1), calculated on the 25th day
after the purchase of the first plan. Class
labels indicate whether the customer has
purchased a second plan within 35 days after
the first plan. Five different methods are
used for classification: k-Nearest Neighbors,
Sup-port Vectors Machine, Decision Tree,
Ran-dom Forest and Na¨ıve Bayes classifier.
This definition of churn can only be used to
predict consumers purchasingmonthly
plans, so it was decided to define churn in
another, more universal way.
2. The user is classified as a churned customer
if he does not use the services provided by
the company for 25 consecutive days (does
not call anyone).</p>
      <p>To find the optimal interval of days,
after which we could treat the user as
leaving, rather than just taking a break between
calls, a percentage of users returned to the
system after  days of inactivity is
calculated. In the graph shown in Figure 3, the
abscissa axis reflects the number of inactive
days , and the ordinate axis corresponds to
the number of users (in percent). The blue
bar then shows the percentage of users who
had an -day interval between calls, and the
red bar represents the percentage of users
who returned to the system after  days (call
again). It can be seen from this graph that
almost all users have had a one-day
interval ( = 1) between calls and only about
60% of them have returned to the system
after this interval. Nearly 80% of users have
had a thirty-day interval ( = 30) between
calls, with less than 25% returning to the
system. There is no clear break in the
number of users who have not returned to the
system, but there is a steady decrease in the
number of users who have returned to the
system. It has been decided that 25 days is
a suficient period of inactivity to consider a
user leaving the system.</p>
      <p>The user is monitored from the first day of
registration until churn (25 inactive days in
a row). In this case, it is not the static
variables that are observed, but their change</p>
      <sec id="sec-4-1">
        <title>6. Experiments</title>
        <p>6.1. Evaluating the purchase of a plan
Users classification is performed by dividing
customers into these two groups:
over time. 25th day after the purchase of the first p lan. An
There are times when after a long break (af- attempt is then made to assign the user to one of
ter so called churn) the user returns to the the classes (predicted to remain in the system or
system and starts using the services again. leave). The characteristics describing user activity
For such cases, the algorithm is designed so are presented in Table 1.
that withdrawn customer is still monitored, Customers are divided into model training (70%
and when he returns to the system (calls data) and testing (30% data) sets. Five
diferagain), he is treated as a newly logged-in ent methods are used for classification: k-Nearest
user. Neighbors, Support Vectors Machine, Decision
Tree, Random Forest and Na¨ıve Bayes classifier.</p>
        <p>The values of the confusion matrix elements
evaluating the accuracy of the listed classification
methods and the percentage accuracy of all models
is given in Table 2. It can be seen that the SVM
model achieves the best accuracy.
• An active user is one who re-orders a plan
within 35 days after the first order of the
plan.
• Withdrawn - did not order another plan</p>
        <p>within 35 days after ordering the first plan.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>The selected forecast period is 10 days. User data is tracked for 25 days from the first plan purchase. Based on these data, the characteristics describing the user’s behavior are calculated on the</title>
      <p>5000 users were randomly selected from the full list
of users. For each of them, the variables in Table
1 are calculated on every day from the time the
user registers until he leaves or until the end of the
entire data range. This results a data set in which
each user is described not by one row, but by as
many rows as number of days the user has been in
the system for. Having this data set it is possible
to track how user activity has changed over time.</p>
      <p>Any user who leaves the system for more than 25
days (does not call anyone for 25 days) and then
returns to it (calls again) is treated as a new user
(assigned a new identification n umber). As a result,
the creation of such a user data set increases the
number of users to a total of 15435. A training set
(70 % of these users) is used to create the model.</p>
      <p>To select the most appropriate Cox regression
model, three diferent c ombinations o f variables
were created and three models were constructed.</p>
      <p>Table 3 shows the variables for all three models. In
each case, only non-correlated, statistically
significant variables are included in the model.</p>
      <p>The accuracy of these three models was assessed
using a test set (30% of users). Four dates in the an- In further research the possibility of combining
alyzed period were selected for model testing. Ac- these two methods to predict the likelihood of
custive users are selected on a specific date and it is tomer churn may be considered.
predicted that after 10 they will still be active or
churned. The same is repeated with four
diferent dates. In this way, the real performance of the References
model is verified, when both short-term customers
and long-term customers are evaluated. Some users [1] B. Huang, M. T. Kechadi, B. Buckley,
Cusmay have been analyzed several times at diferent tomer churn prediction in telecommunications,
times. In total, the model evaluated customers Expert Systems with Applications 39 (2012)
2741 times, of which 1976 users did not quit and 1414–1425.
765 when users left the system. The accuracy of [2] H. REN, Y. ZHENG, Y. rong WU, Clustering
the models is presented in Table 4. It can be seen analysis of telecommunication customers, The
that M1 and M3 models have achieved equal ac- Journal of China Universities of Posts and
curacy and are more suitable for predicting churn
than M2 model.</p>
      <sec id="sec-5-1">
        <title>7. Conclusion</title>
        <p>Experiments with the telecommunication customer
data set show that:
1. After assessing the specifics of the available
data, it was decided to define user activity
in two ways: according to the plans to be
purchased and according to the frequency of
calls made.
2. In the case where the customer is
considered active as long as he regularly
purchases the call plans ofered by the
supplier, the following classification algorithms
were used to segment the users: k-Nearest
Neighbors, Support Vector Machine,
Decision Tree, Random Forest, and Na¨ıve Bayes
classifier. Based on experimental studies,
it can be stated that the classification of
Support Vector Machine outperforms other
methods.
3. In the case where user activity is defined by a
faster-changing indicator – the frequency of
calls, it was decided to use the Cox regresion
model with time varying covariates to divide
users into groups. This model is superior
to classical classification methods in that it
can take into account not only static user
parameters but also their change over time.
Telecommunications 16 (2009) 114 – 128.
URL: http://www.sciencedirect.com/science/
article/pii/S1005888508602149. doi:https:
//doi.org/10.1016/S1005-8885(08)60214-9.
[3] J. Pamina, B. Raja, S. SathyaBama, M. Sruthi,
A. VJ, et al., An efective classifier for
predicting churn in telecommunication, Jour of Adv
Research in Dynamical &amp; Control Systems 11
(2019).
[4] A. K. Ahmad, A. Jafar, K. Aljoumaa,
Customer churn prediction in telecom using
machine learning in big data platform, Journal of
Big Data 6 (2019) 28.
[5] I. Braˆndu¸soiu, G. Toderean, H. Beleiu,
Methods for churn prediction in the pre-paid
mobile telecommunications industry, in: 2016
International Conference on Communications
(COMM), 2016, pp. 97–100. doi:10.1109/
ICComm.2016.7528311.
[6] I. M. Mitkees, S. M. Badr, A. I. B. ElSeddawy,
Customer churn prediction model using data
mining techniques, in: 2017 13th International
Computer Engineering Conference (ICENCO),
IEEE, 2017, pp. 262–268.
[7] J. Han, J. Pei, M. Kamber, Data mining:
concepts and techniques, Elsevier, 2011.
[8] G. Norkevičius, G. Raškinis, Lietuvių kalbos
garsų trukmės modeliavimas klasifikavimo ir
regresijos medžiais, naudojant didelės apimties
garsyną, Informacinės technologijos 2007:
konferencijos pranešimų medžiaga, Kauno
technologijos universitetas, 2007 m. sausio 31
d.vasario 1 d. Kaunas: Technologija, 2007 (2007).
[9] G. Biau, E. Scornet, A random forest guided
tour, Test 25 (2016) 197–227.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>