<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Applying Machine Learning Methods to Forecasting Customer Churn for a Telecommunications Company*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sholom-Aleichem Priamursky State University</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shirokaya street</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Birobidzhan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Russian Federation r-i-bazhenov@yandex.ru</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Irkutsk National Research Technical University</institution>
          ,
          <addr-line>83, Lermontova street, Irkutsk, 664074, Russian Federation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Maritime State University named after G.I.Nevelskoy</institution>
          ,
          <addr-line>50A, Verkhneportovaya street, Vladivostok, 690059, Russian Federation</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Sumy State University</institution>
          ,
          <addr-line>2, Rymskogo-Korsakova street, Sumy, 40007</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The paper presents a brief overview of existing approaches to predicting customer churn using the example of a telecommunications company. The authors provide research on churn forecasting using 11 different machine learning methods. For training, data containing 20 different information parameters about clients were used. The quality of education was assessed using the traditional Area Under the Curve characteristic. The paper also provides research results confirming that the use of ensembles of machine learning methods increases the quality of predicting customer churn.</p>
      </abstract>
      <kwd-group>
        <kwd>Machine Learning</kwd>
        <kwd>Customer Churn</kwd>
        <kwd>Telecommunications Company</kwd>
        <kwd>Area Under The Curve</kwd>
        <kwd>An Ensemble of Techniques</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The number of clients for any company is undoubtedly an important parameter, and
the more clients, the higher the company's profit. There is no monopoly on the
provision of mobile services to customers, high-speed Internet access or cable
television, therefore, the so-called outflow of customers is possible for one reason or
another, most often associated with the transition to another telecommunications
company offering more favorable conditions from the point of view of customers ...
To ensure the stable operation of a telecommunications company, an analysis of the
customer base is necessary, for example, to ensure targeted promotions, as well as an
analysis of the reasons for customer churn, including with the ability to predict the
number of customers who have left and the reasons for switching to the services of
another telecommunications company. According to experts, attracting one new client
costs companies on average seven times more (in some cases, from 5 to 20 times
more) than retaining an existing one [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1-3</xref>
        ]. Reducing customer churn by 5% increases
the company's profits from 25% to 85% [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Therefore, understanding how exactly to
maintain customer engagement is a natural foundation for developing strategies for
customer retention.
      </p>
      <p>
        In this regard, there is a number of scientific studies aimed both at predicting
customer churn and predicting their preferences in any specific services. For example,
in work [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the authors try to understand the behavior of customers when paying
bills, using such machine learning methods as logistic regression, one rule and
support vector machines, they proposed an analytical approach to studying and
predicting the payment behavior of customers. As input for the analytical system, they
use: customer ID (a hashed unique number indicating each customer), action type
(such as SMS, IVR, phone calls, service cut related and legal actions), the date of the
action's execution, stage changer fl ag (such as payment through the banking system,
unpaid invoice occurrence, or action timer fl ag). The feasibility of using machine
learning methods is ensured by their ability to detect patterns in data of various
natures and universes [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        In [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], on the basis of Call Data Records data, various types of customers are
determined, which subsequently affects the analysis of the possibility of their outflow
and the impact on other customers of the company. Authors include these types of
customers: follower, standard, leader, core and important customer. The authors note
that customers who interact with the Leader and Important categories are more likely
to churn after an influential member who also left the company. For clustering, the
authors used neural networks such as linear perceptron, multilayer perceptron and
networks with radial basis activation function. In the work of the authors [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], using
the random forest classifier, decision tree classifier, gradient-boosted tree classifier
and multiplexer perceptron, clients were segmented according to the
“time-frequencymonetary” approach, while “Time” characterizes the total of calls duration and
Internet sessions in a certain period of time, the "Frequency" parameter characterizes
the frequency of using services frequently within a certain period, and the "Monetary"
parameter is determined by the amount of money spent during a certain period.
      </p>
      <p>
        It should be noted that there is often a situation where the assessment of the
forecast of customer churn is carried out according to completely different criteria.
For example, in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], forecasting is based on the following information: region where
the customer lives; time since the customer joined the operator (in months); average
revenue; how long did the customer not pay the bills; amount that the customer is
overdue; number of times the service was disconnected, etc. In the study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the
following parameters are used to predict customer churn: state the US state, in which,
the customer resides, the remaining seven-digit phone number, total number of calling
minutes used during the day, the billed cost of daytime calls, total number of calling
minutes used during the nighttime, and a few others. Moreover, some of these criteria
may not be available to telcos, making it difficult to determine how, in practice, a
telco will forecast these models. It is also worth mentioning several studies [
        <xref ref-type="bibr" rid="ref11 ref12">11-12</xref>
        ] on
the use of customer analysis technologies in the banking sector, for example, to
predict the likelihood of outflow based on information related to the
sociodemographic parameters of customers, their activity in obtaining banking information,
information about their salaries, and etc.
      </p>
      <p>There is no unambiguous option for the machine learning methods used, there are a
lot of them, it is almost impossible to choose an empirically suitable method, it is
necessary to conduct training, in a practical way, selecting the optimal architecture
and parameters of the machine learning method.</p>
      <p>This paper presents the results of predicting the churn of customers of a
telecommunications company using machine learning methods, since the latter are
able to use statistical data, to determine dependencies between data of different
nature.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Preparation of training sample</title>
      <p>
        The problem of predicting customer churn can be formulated as a binary classification
problem [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The solution to the problem is to classify customers with the
corresponding characteristics (parameters) x ∈ X to one of the two classes Y = {no
churn, churn}.
      </p>
      <p>To solve such a problem, in this work, we used a dataset consisting of 21 columns
and 7043 rows made publicly available by IBM and containing information about
customers of a telecommunications company (available at
https://www.kaggle.com/blastchar/telco- customer-churn). Information about each
customer includes the following characteristics, which are identifying signs for
predicting churn (they are presented in Tables 1-3), and, of course, each customer is
assigned a unique customer identifier (customerID).</p>
      <p>Also, data on the period of use of the services of a telecommunications company
were used as input information, the parameter is calculated in months (tenure); the
size of the client's monthly fee (MonthlyCharges) and the final amount of payments
for the entire period of work with the client (TotalCharges). For each set of input
information, consisting of 20 indicators, there is a predetermined output - Churn
characterizing whether the client left the telecommunications company with these
parameters or not.</p>
      <p>No.
1
2
3
4
5
6
No.
1
2
3
4
5
6
7
8
9</p>
      <p>Also, data on the period of use of the services of a telecommunications company
were used as input information, the parameter is calculated in months (tenure); the
size of the client's monthly fee (MonthlyCharges) and the final amount of payments
for the entire period of work with the client (TotalCharges). For each set of input
information, consisting of 20 indicators, there is a predetermined output - Churn
characterizing whether the client left the telecommunications company with these
parameters or not.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Modeling machine learning methods</title>
      <p>The following machine learning methods have been selected for training, which have
proven themselves well in solving various classification problems: AdaBoost
adaptive boosting; Decision Tree - decisive trees; Extra Tree Classifier - random
trees; Gradient Boosting - gradient boosting; KNeighbors - k-nearest neighbors
method; Logistic Regression - logistic regression; Naive Bayes - naive Bayesian
classifier; Neural Network - neural networks; Random Forest - random forest method;
SVM - support vector machine; XGB - Gradient boosting on trees.</p>
      <p>
        To assess the effectiveness of the models, the Area Under the Curve (AUC)
characteristic was chosen, a statistical indicator that is often used in machine learning
methods that determines the area bounded by a certain curve and an abscissa axis,
called the ROC curve (from receiver operating characteristic). The use of AUC for
binary classification problems is popular because of its simplicity, intuitive
interpretation [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and also its use in the case of unbalanced datasets [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. For a
random classifier, the AUC is 0.5, and for an ideal classifier, the AUC is 1 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Since
there are only two options for a conclusion in the problem being solved (whether the
client leaves or not), it is the ROC analysis, which is a graphical method for assessing
the quality of the work of a binary classifier, that is most promising for assessing the
proposed method for predicting the churn of customers of a telecommunications
company. To construct the ROC curve, a pair of the following values is used:
sensitivity and specificity. The value "sensitivity" characterizes the share of
truepositive classifications in the total number of positive observations and is marked
along the vertical axis of the ROC-curve graph, and the value "specificity"
characterizes the proportion of true-negative classifications in the total number of
negative observations and is marked along the horizontal axis of the ROC-curve
graph. It should be noted that the higher the “sensitivity” value, the more reliably the
classifier recognizes positive examples, and the higher the “specificity” value, the
more reliable the classifier recognizes negative observations. Thus, the ROC curve
reflects the relationship between the probability of false alarms (proportion of
falsepositive classifications) and the probability of “correct detection” (proportion of
truepositive classifications). With an increase in sensitivity, the reliability of recognition
of positive observations increases (the probability of "missing a target" decreases), but
at the same time the probability of a false alarm increases. Table 4 shows the learning
outcomes of individual machine learning methods and their characteristics according
to the AUC metric.
Thus, as a result of the work carried out, the authors investigated 11 different machine
learning methods to solve the problem of predicting the churn of customers of a
telecommunications company. The best machine learning method was
LogisticRegression, which showed an AUC of 80.38%. The use of Logistic
Regression in conjunction with other machine learning methods within the ensemble
of machine learning methods allowed us to increase the forecasting quality to 81.37.
Further research by the authors will be devoted to identifying those identification
parameters that significantly affect the process of predicting customer churn in order
to use a smaller number of identification features with constant values of the forecast
quality.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Verbeke</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          et al:
          <article-title>Building comprehensible customer churn prediction models with advanced rule induction techniques</article-title>
          .
          <source>Expert systems with applications</source>
          ,
          <volume>38</volume>
          (
          <issue>3</issue>
          ),
          <fpage>2354</fpage>
          -
          <lpage>2364</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Qureshi</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          et al:
          <article-title>Telecommunication subscribers' churn prediction model using machine learning</article-title>
          .
          <source>In Eighth International Conference on Digital Information Management (ICDIM</source>
          <year>2013</year>
          ),
          <fpage>131</fpage>
          -
          <lpage>136</lpage>
          . IEEE,
          <string-name>
            <surname>Islamabad</surname>
          </string-name>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Balle</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al.:
          <article-title>The architecture of a churn prediction system based on stream mining</article-title>
          .
          <source>Artificial Intelligence Research and Development</source>
          ,
          <volume>256</volume>
          ,
          <fpage>157</fpage>
          -
          <lpage>166</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khunteta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Churn Prediction in Telecommunication using Logistic Regression and Logit Boost</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>167</volume>
          ,
          <fpage>101</fpage>
          -
          <lpage>112</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Bahrami</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          , Bozkaya,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Balcisoy</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Using Behavioral Analytics to Predict Customer Invoice Payment</article-title>
          .
          <source>Big Data</source>
          ,
          <volume>8</volume>
          (
          <issue>1</issue>
          ),
          <fpage>25</fpage>
          -
          <lpage>37</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Xie</surname>
          </string-name>
          , S.-M.:
          <article-title>Comparative models in customer base analysis: parametric model and observation-driven model</article-title>
          .
          <source>Journal of Business Economics and Management</source>
          ,
          <volume>21</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1731</fpage>
          -
          <lpage>1751</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Kostić</surname>
            ,
            <given-names>S. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simić</surname>
            ,
            <given-names>M. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kostić</surname>
            ,
            <given-names>M. V.</given-names>
          </string-name>
          :
          <article-title>Social Network Analysis and Churn Prediction in Telecommunications Using Graph Theory</article-title>
          . Entropy,
          <volume>22</volume>
          (
          <issue>7</issue>
          ),
          <volume>753</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Wassouf</surname>
            ,
            <given-names>W.N.</given-names>
          </string-name>
          et al:
          <article-title>Predictive analytics using big data for increased customer loyalty: Syriatel Telecom Company case study</article-title>
          .
          <source>Journal of Big Data</source>
          ,
          <volume>7</volume>
          ,
          <issue>29</issue>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Baesens</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al:
          <article-title>Profit Driven Decision Trees for Churn Prediction</article-title>
          .
          <source>European journal of operational research</source>
          ,
          <volume>284</volume>
          (
          <issue>3</issue>
          ),
          <fpage>920</fpage>
          -
          <lpage>933</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Faris</surname>
            <given-names>H.</given-names>
          </string-name>
          :
          <article-title>A Hybrid Swarm Intelligent Neural Network Model for Customer Churn Prediction and Identifying the Influencing Factors</article-title>
          . Information,
          <volume>9</volume>
          (
          <issue>11</issue>
          ),
          <volume>288</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Alexandru</surname>
            <given-names>C.</given-names>
          </string-name>
          et al:
          <article-title>Propensity to Churn in Banking: What Makes Customers Close the Relationship with a Bank? Economic Computation</article-title>
          &amp;
          <source>Economic Cybernetics Studies &amp; Research</source>
          ,
          <volume>54</volume>
          (
          <issue>2</issue>
          ),
          <fpage>77</fpage>
          -
          <lpage>94</lpage>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Schaeffer</surname>
            <given-names>S. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sánchez</surname>
            <given-names>S. V. R.</given-names>
          </string-name>
          :
          <article-title>Forecasting client retention - a machine-learning approach</article-title>
          .
          <source>Journal of Retailing and Consumer Services</source>
          ,
          <volume>52</volume>
          ,
          <issue>101918</issue>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jafar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aljoumaa</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Customer churn prediction in telecom using machine learning in big data platform</article-title>
          .
          <source>Journal of Big Data</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ),
          <volume>28</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al:
          <article-title>Benchmarking sampling techniques for imbalance learning in churn prediction</article-title>
          .
          <source>Journal of the Operational Research Society</source>
          ,
          <volume>69</volume>
          (
          <issue>1</issue>
          ),
          <fpage>49</fpage>
          -
          <lpage>65</lpage>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>