<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Harnessing the Complexity of Mobile Network Data with Smart Monitoring*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandr Suleykin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Panfilov</string-name>
          <email>ppanfilov@hse.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University - Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>V. A. Trapeznikov Institute of Control Sciences, Russian Academy of Sciences</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Mobile (cellular) networks represent a fast evolving research field that take advantages of recent technological advances such as Big Data and distributed computing to provide extensive network monitoring for network operation planning and management purposes. Challenges related to making use of the large volume of streaming data generated by mobile networks include extracting relevant elements within massive amounts of signals possibly spread across different sources (data bases), reducing dimensionality, summarizing dynamic information in a comprehensible way and displaying it for interpretation purposes. The adequate network modeling provides both statistical data of network performance and important networking QoS related insights. Comprehensible mobile network information is needed that uncovers the role of each attribute and variable. To harness the complexity of mobile network data and to extract relevant information a dedicated distributed computing platforms and Big Data frameworks are needed, able to discover and deal with the inherent properties and complexities of these datasets. Smart network monitoring service plays a central role here.</p>
      </abstract>
      <kwd-group>
        <kwd>Smart Monitoring System</kwd>
        <kwd>Cellular Networks</kwd>
        <kwd>Distributed Computing</kwd>
        <kwd>Big Data</kwd>
        <kwd>Dynamic Data-Driven Application System</kwd>
        <kwd>Lambda Architecture</kwd>
        <kwd>DDSM</kwd>
        <kwd>Roaming Users Detection and Monitoring</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The number of smartphone users has already reached 4.61 billion users in 2016,
and upward trend is forecast for the market with 5.07 billion users by the end of 2019
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The rapid growth of a mobile user count immutably leads to the proportional
increase of data being generated by mobile subscribers, user equipment, cellular
network nodes and mobile networks as a whole. This is becoming more challenging
for mobile operators to handle constantly increasing data volumes for many different
protocols and mobile network interfaces using traditional solutions with standalone
systems, relational databases, many different formats of data storage and transmission.
To address these new challenges a new approaches such as Big Data, Internet of
Things, Machine-to-Machine Communications, distributed computing platforms and
paradigms find its application to cellular network data storage, aggregation,
transformation, enhancement and transfer.
      </p>
      <p>
        Real cellular networks exploit multiple different protocols for data transmission
and corresponding interfaces. Each node of the mobile network is communicating
with other node(s) and external environment according to the standard protocols such
as 3GPP [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and ITU [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Every protocol has its own parameters, which are different
2
from one network element to another. The complexity and variety of data protocols,
very large volumes of data being transferred and importance of data have led to the
need of new approaches to the cellular network data monitoring and analysis, using
latest technology achievements such as Big Data and Machine Learning.
      </p>
      <p>The smart monitoring system (SMS) vision relies on the use of ICT to efficiently
manage and maximize the utility of mobile network infrastructures in order to
improve the quality of service and network performance. Many aspects of SMS
projects are dynamic data driven application systems where data from sensors
monitoring the system state are used to drive computations that in turn can
dynamically adapt the monitoring process as the complex system evolves. In this
context, a research and development of a distributed Big Data driven framework for
cellular network monitoring data entails the ability to dynamically incorporate more
accurate information for mobile network controlling purposes through obtaining
realtime measurements from the network meters, base stations, and other sensors.</p>
      <p>
        Traditional network monitoring services, exemplified by Netboss XT [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] a
sophisticated multipurpose telecommunication networks management system are
designed to monitor networks, detect failures and provide network maintenance and
performance analysis in an multi-vendor, multi-protocol, multi-service heterogeneous
environments. The major advantage of such a system is its flexibility – the system
comprises of hundreds of software intelligent agents that permit to manage the
equipment of almost any known vendor and interface with major OSSs. It also
provides an integrated development environment and programming language to create
new software agents or enhance existing ones.
      </p>
      <p>However, traditional approaches face problems of handling increasing complexity
of network data because they do not scale well enough and are much slower when
compared to the modern high-performance data analytics (HPDA) and distributed
computing enabled solutions.</p>
      <p>
        We propose a Distributed Big Data Driven Framework (DBDF) for cellular
network monitoring data on the basis of the Dynamic Data-Driven Application
System (DDDAS) paradigm [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and a core concept of Lambda architecture [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6-8</xref>
        ],
specifically targeted at scalable and secure real-time Big Data application systems, the
comparison of different Big Data methods, techniques, and available tools. The
DBDF consists of different components that realize SMS paradigm for the cellular
network data in distributed Big Data driven fashion as it is presented in figure 1.
      </p>
      <p>Real System</p>
      <p>External
environment</p>
      <p>Additional
income</p>
      <p>Vendor specific
Cellular Network</p>
      <p>Probes</p>
      <p>Data
collection
enriDcahtmaent</p>
      <p>Geopositioning
Decision makers</p>
      <p>Network
Operation
Network
Planning
Network
Engineering
Other data
consumers</p>
      <p>Fast circuit (real-time processing)</p>
      <p>Big Data
realtime data parsing</p>
      <p>Data
parsing
application
Universal
output data
format</p>
      <p>Big Data
realtime data
collection
Middleware
Protocols
structure</p>
      <p>Realm-toimdeelsdata
Roaming users
detection</p>
      <p>Network
Troubleshooting issues
Case-specific</p>
      <p>models
Slow circuit (batch processing)
Offline data models Big Data storage
reGpeoorts quDeartiaes</p>
      <p>Offline
Recomendations AnDaaltyasis
Advertisment vHeyrpifoicthateisoins</p>
      <p>Other data sources</p>
      <p>Base Station
characteristics
Data Lake
Case-specific
data</p>
      <p>The new concept of DBDF for handling cellular network management and
operation problems is targeted at network management representatives and decision
makers in Mobile Engineering, Operating and Planning departments to help them in
managing complexity of the cellular network data in real-time on the basis of smart
network measurements, simulation and optimization models.</p>
    </sec>
    <sec id="sec-2">
      <title>Distributed Computing and Big Data Issues</title>
      <p>In fact, cellular network data is a streaming data coming from interfaces, base
stations, billing system and other sources and represent highly-loaded systems with
Gigabytes and Terabytes of data per second, even though this is a compressed data.
As such, the CN monitoring data need to be parsed and processed on a cluster of
computers with application of Big Data methods for processing of very large volumes
of data. Distributed data processing methods and techniques play an important role in
the whole CN data processing pipeline. All available and needed application systems
and methods should be adapted for high performance and fully distributed computing
environment and support scalable, reliable, and secure data processing. This is a key
requirement for dealing with complexity in modern CN infrastructures.</p>
      <p>Distributed computing techniques have been widely used by data scientists before
the advent of Big Data concept. Thus, standard and time-consuming algorithms were
successfully replaced by their distributed versions with the aim of agilizing the
learning process. For many of current issues of Big Data, in particularly, in cellular
network operation characterized with large volumes of streaming data, a distributed
approach is becoming immutable nowadays.</p>
      <p>The first distributed computing framework that enabled the processing of big
volumes of data was the MapReduce paradigm. This tool was aimed at easily handling
huge datasets in an automatic and distributed way using Map and Reduce task
concepts, that enable user with building a distributed and scalable applications while
hiding technical details as data partitioning, failure recovery or job communication.
However, MapReduce concept is not designed to scale well when dealing with
iterative and online processes, and usually deal with batch data tasks with relatively
huge latency comparing with online data processing. In our suggested architecture this
paradigm would be recommended for implementation in Batch layer for Big Data
storage system component.</p>
      <p>Another group of methods is based on distributed in-memory computing,
microbatch and real-time techniques. These methods are usually used for online Big Data
processing with millions of tuples per second per node performance. In our proposed
solution of the Smart Monitoring System architecture these methods are used in
realtime data transformation system component, in real-time data parsing component and
also might be used in a real-time data modeling component.</p>
      <p>Message-driven applications are applications, that allow processing of future
messages that will arrive after subscription. The main advantage of such systems is
that many consumers can access the same data in independent way. Thus, a
messagedriven approach is suggested to be used in Big Data real-time data collection to enable
many data consumers easily access cellular network data.
2.1</p>
      <sec id="sec-2-1">
        <title>Lambda Architecture for Big Data</title>
        <p>Building a reliable and efficient distributed Big Data application that satisfies a
variety of end-user requirements is a challenging task. Lambda Architecture (LA)
represents a useful framework for designing such applications. The appearance of the
LA concept was inspired by the following motivations:
 the need for a robust and fault-tolerant system, both against human mistakes
and hardware failures;
 the system should be linearly scalable scaling out rather than up;
 to serve a wide range of workloads and use cases, where low-latency reads
and updates are required, with support of ad-hoc queries;
 the system should be extensible and features should be added easily.
Essentially, the Lambda Architecture (LA) comprises of 3 main parts:
 Batch layer. This layer has two functions: manage the main append-only raw
data streams and pre-compute arbitrary query functions calling batch views.
In our DBDF architecture this is a so called “slow circuit” layer, where
cellular network data is coming from real-time data collection component
using batches.</p>
        <p>Speed layer. There are different fast and incremental algorithms which are
used with low latency. The speed layer deals only with the most recent data.
In the DBDF, these are key components that represent a “fast circuit”, and all
components inside this layer deal with online (real-time) data. Latency is a
major concern there.</p>
        <p>
          Serving layer. This layer indexes the batch views in Batch layer, and data can
be queried ad hoc with low latency. This layer is not included in current
DBDF implementation and will be considered as a future work aimed at
building DBDF Serving data layer [
          <xref ref-type="bibr" rid="ref6 ref7 ref8">6-8</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Apache Spark and Lambda-enabled DBDF Architecture</title>
      <p>Our implementation of the DBDF concept is based on Apache Spark open-source
platform which supports all ranges of Big Data formats like batch data, text data,
realtime streaming data, graphical data, etc. Apache Spark has already proved its huge
potential in the Big Data industry because of its in-memory data processing that
makes it high-speed data processing engine compare to Hadoop MapReduce.</p>
      <p>Apache Spark provides a powerful data processing engine with development APIs
to allow data scientists to execute streaming conveniently. Streaming is unstructured
data that is generated continuously by thousands of data sources, including log files
generated by customers using mobile or web applications, in-game player activity,
information from social networks, financial trading and telemetry from connected
devices or instrumentation in data centers, IoT sensors, etc.</p>
      <p>With Spark running on Apache Hadoop YARN, developers can now create
applications to consume and transform complex data streaming from Apache Kafka.
Complex transformations like exactly-once event-time aggregation can be expressed
using this API and the results can be output to a variety of systems. Apache Spark has
the inbuilt support of over 80 high-level operators and lets programmers write
applications using different programming languages as Python, Clojure, Scala or Java.</p>
      <p>The low-latency in-memory data processing capability of Spark allows for
handling many data analytics challenges using its machine learning libraries and
graph analytics algorithms. With the help of Spark the CN providers can enable
themselves to analyze data coming from various kinds of data sources, because Spark
can easily process continuous streams of low-latency data. Thus, CN providers can
create real-time dashboards and explore data in real-time to monitor and optimize the
CN operation.</p>
      <p>Figure 2 presents a high-level view of the Spark and Lambda-based
implementation of the DBDF for the smart processing and analysis of streaming data.</p>
      <p>Protocols
streaming</p>
      <p>MAP
DIAMETER</p>
      <p>SIP
Fig. 2. Overview of the DBDF component architecture – an implementation
Other data sources
In-Memory DBs</p>
      <p>NoSQL DBs</p>
      <p>SQL DBs
3.1</p>
      <sec id="sec-3-1">
        <title>Real Cellular Network</title>
        <p>Real CN consists of multiple network elements such as RNCs, BSs, NodeBs,
eNodeBs, MMEs, GGSNs, and other nodes receiving and sending data via different
cellular interfaces. Every new generation of mobile networks offers its own protocols
for data exchange and interfaces which adds to the overall complexity of the real CN.
The situation is aggravated by a variety of mobile device interfaces to cellular
networks. Typically, a mobile device has both an application processor and a cellular
processor, which each operate in different power states and data states. On most
mobile devices, applications use multiple forms of connectivity with multiple
transitions from one to the other. To manage connections to CN, the device itself
usually has multiple cellular data modes with different data rates and different power
usage. The cellular processor communicates with the CN in two different planes:
the control plane, which manages the communication protocol, and the data plane,
which carries the actual data packets representing different kinds of data – voice,
SMS, MMS, data itself.</p>
        <p>To provide good quality of service to the mobile user, CN providers attempt to
deploy a CN that will never have its capacity exceeded. They can also be sensitive to
the behavior of mobile applications which possess features that, intentionally or not,
push the boundaries of the networks. As a result, carriers often need to restrict
application behavior to keep them within the limitations of the networks. For
example, a CN provider can limit the number of connections that an application
triggers, especially in the control plane.</p>
        <p>In addition, network measurements might be accompanied with the base stations
characteristics, environmental (weather) condition data, and a mobile user’s demands
and/or complaints. These represent different attributes and properties of the real CN.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Cellular Network Events</title>
        <p>Event-Driven network data are data records that appeared after specific network
event. It means that different data items should be created and transferred after some
event. Thus, these events might be presented as data communications between nodes
(transactions), data on the base station parameters, user complaints, etc.</p>
        <p>Base station data records are the characteristics of base station such as their
geopositioning, supportive technology, LAC code, Cells, address, vendor etc. User
complaints are complaint data records that are collected in connection to the particular
geographical location. Different transactions are data records that represent a part of a
subscriber session. The data in transactions are being transferred according to the
3GPP and ITU standards in predefined format – CN data exchange protocols.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Network Data Parsing</title>
        <p>Big Data driven programming environment is a special purpose application capable of
processing huge volumes and streams of data (gigabit per second, Gbps) that parses
all highly compressed protocol data (depends on vendor format) to the one unified
data format, or processes data records received from real CN (BS parameters,
complaints). The data parsing is implemented via programming language, the choice
of which is highly important because different frameworks offer diverse functionality
and more adopted for some particular cases.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Message Oriented Middleware</title>
        <p>Message-Oriented Middleware is the Continuous Streaming data storage with
predefined data structure according to different protocols. This is the middleware that
provides opportunity for many data consumers to get data in real-time as well as send
data for storage using batches. The issue here is that there might be many data
consumers, including internal mobile provider departments and external partners.
Actually, it hardly depends on the use cases – the more data use cases, the more data
6
consumers are. So the most important requirements for this layer is the capability of
handling extremely fast data streams, high reliability, scalability and support to many
different data consumers.</p>
        <p>
          To meet requirements, it is proposed to use Apache Kafka messaging system as a
distributed streaming platform. In a large distributed system, there are usually a lot of
services that generate different events: logs, monitoring data, noticed attempts to
access secret resources, etc. On the other hand, there are services that need these data.
Kafka helps here as an interface between data producers and consumers: it collects
data from former, then stores it in a distributed store using topics and distributes it at
the latter via subscription. In other terms, Kafka is a hybrid of a distributed database
and message queue [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
3.5
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Offline Data Models</title>
        <p>
          It is proposed to use Apache Storm or Apache Spark applications for data
transformations. Apache Storm is a free and open source distributed real-time
computing system that performs processing of data streams. It has many use cases:
continuous computation, ETL processes, real-time analytics, online machine learning
and others. Storm is fast: a benchmark tests show that it can process at over a million
tuples per second per node. It is fault-tolerant, scalable and simple to set up and
operate [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Apache Spark is a fast and general-purpose cluster computing system. It
has high-level APIs in Scala, Python, Java and R, optimized engine with support of
general execution graphs, higher-level tools such as Spark SQL for SQL and
structured data processing, GraphX for graph processing, Spark Streaming, MLlib for
machine learning [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
        <p>
          Also, for offline data processing mode we propose to use Python programming
language for implementing machine learning and data mining algorithms. It has a
clear syntax, is well developed and documented [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
        </p>
        <p>
          For offline mode we also propose to use Apache Airflow for scheduling. It is a
library (or a set of libraries) for the development, planning and monitoring of work
processes. With the Airflow we can use Python for coding work processes. Hence,
there are advantages for organizing project and application development efforts.
Implementation is fairly simple. One can use, for example, PyCharm plus Git [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
3.6
        </p>
      </sec>
      <sec id="sec-3-6">
        <title>Other Data Sources</title>
        <p>
          These are data sources within a CN, which are mainly used in the network simulation.
Usually, these data are collected and stored in data bases that are maintained by
different departments of the CN provider. It can be SQL-, NoSQL-based, in-memory
or other data structures, but the main concern is that data should be readily accesible
by real-time data models. Other data sources component in the DBDF is considered an
additional on-demand data source(s) for network simulation. In our previous work [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ],
we used data from Base stations data base, user complaints data base and population
density data for the simulation and visualization of the so called “Problem Zones” in
CN. In general, the particular use case and different requirements dictate the selection
of the data bases and simulation tools.
3.7
        </p>
      </sec>
      <sec id="sec-3-7">
        <title>Big Data Storage and Queries</title>
        <p>This Big Data-driven layer in the general DBDF framework represents a component
responsible for protocols data storage, report generation, validation of hypotheses and
protocols data analysis. It can be done by using SQL queries to some database. Offline
storage is a component that realizes reliable and scalable data storage of cellular
network protocol data to be used in offline analysis, hypothesis validation and
different kinds of on-demand reports.</p>
        <p>
          For data storage solution we propose to use Hadoop NoSQL data base. It is
opensource software for scalable, reliable, distributed computing, the framework that
allows the distributed processing of large data sets across clusters of computers using
7
simple programming models [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. It is highly scalable solution specially designed to
scale up from the single servers to thousands of machines, where each server is
offering storage and computation. It is created to detect and come up with failures at
the application layer.
        </p>
        <p>
          We propose to use Apache Hive for SQL queries to Hadoop. It supports analysis of
large data sets stored in Hadoop in HDFS and compatible file systems, such as the
Amazon S3 file system. Hive provides an SQL-like language called HiveQL with read
schemes and with transparent mapping of MapReduce, Apache Tez and Spark jobs.
All three execution engines run on Hadoop YARN. To speed up queries, it provides
indexes, including bitmap indexes [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
3.8
        </p>
      </sec>
      <sec id="sec-3-8">
        <title>Decision Layer</title>
        <p>
          Decision support layer composes the output from Big Data models and reports. It
might be a streaming data in special predefined format, regular reports or triggered
data events after filtering. The decision layer is the environment specially designed for
Mobile Network Engineers, Planners, Operators, Managers and other data consumers.
It includes all applications needed to aid decision process on the basis of model output
data generated at previous stages of data pipeline. Decisions are made on the basis of
this data: detection and prioritization of problem zones in the network [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], planning
the network capacity and performance, base stations construction and deployment,
managing Radio network elements and other.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments and Results</title>
      <p>As a proof of DBDF concept, a near real-time Big Data based prototype application
for roaming users detection and monitoring was developed and tested using the real
cellular network data of one of the largest mobile providers in Russia.</p>
      <p>Mobile providers are keen to keep track of their subscribers who go abroad or
migrate internally between regions. In fact, the significant percentage of all telecom
operator’s income is attributed to roaming users. Usually data usage, voice calls and
SMS delivery are more expensive when traveling in other countries and regions, and
telecom companies want to exploit solutions for detecting in real-time those users
migrating abroad or in other regions in order to offer them special services.</p>
      <p>Thus, telecom companies usually require:
 the ability to obtain data about the geolocation of the subscriber in real time
for communication with the subscriber in Real-Time Marketing system,
 the ability to proactively detect the presence of blocking inconsistencies on
the side of billing systems and HLRs (Home Location Registered).</p>
      <p>The prototype system was developed using Apache Spark application and the
Python programming language. The comparative analysis of the DBDF-based
prototype system with traditional standalone monitoring services in cellular network
was conducted to demonstrate benefits of the proposed framework such as scalability,
reliability, speed and performance, as well as it’s applicability to other similar use
cases and possibility to check new hypotheses. A prototype application of roaming
user detection and monitoring service demonstrated that Apache Spark streaming
provides less delay, less processing time and as a result more time for decision makers
in comparison to the traditional batch processing.
4.1</p>
      <sec id="sec-4-1">
        <title>Experimental Setup</title>
        <p>
          The prototype system was developed and executed on the basis of Apache Spark
application, which is the largest open source project in data processing [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Since its
release, Apache Spark, as the unified analytics engine, has seen rapid adoption by
enterprises across a wide range of industries. Internet powerhouses such as Netflix,
8
Yahoo, and eBay have deployed Spark at massive scale, collectively processing
multiple petabytes of data on clusters of over 8,000 nodes. It has quickly become the
largest open source community in Big Data, with over 1000 contributors from 250+
organizations [
          <xref ref-type="bibr" rid="ref15 ref16">15-16</xref>
          ].
        </p>
        <p>For our experiments, the YARN was chosen as a resource manager for Apache
Spark. This manager is actually managing all cluster resources available for Spark
jobs, which means that the capacity and performance of application are limited by
resource manager YARN.</p>
        <p>
          For messaging system, an Apache Kafka application was selected as a system with
strong performance that can handle more than 100 000 events per second [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
        </p>
        <p>All installations of Spark, YARN and other support applications were done by
Hortonworks (HDP version is 2.6.3.0-235).</p>
        <p>Common configuration parameters are:
 Java version is 1.8.0_77 (Oracle Corporation);
 Scala version is 2.11.8;
 Operational system is Linux;
 Operational system version is 3.10.0-514.21.1.el7.x86_64.







</p>
        <p>The experiments were conducted on a powerful cluster. The characteristics of
cluster and YARN resources available for all Spark jobs that run on cluster are as
follows:


3 nodes for YARN allocated;
Memory allocated for all YARN containers on a node is 306Gb. Total
memory is 918 GB;
Minimum Container Size (Memory) is 2048Mb;
Maximum Container Size (Memory) is 100Gb;
Number of virtual cores is 32;
Percentage of physical CPU allocated for all containers on a node is 80%;
Minimum Container Size (VCores) is 1;
Maximum Container Size (VCores) is 32. VCores total is 96;
1 second interval between jobs;
2 Spark Executors selected on default.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Prototype Application Data and Code</title>
        <p>Following to the specifications 3GPP and ITU, to analyze the subscribers relocation to
other countries and/or regions inside a country we have looked at MAP (Mobile
Application Part) protocol for 2G and 3G technologies (if user is currently using 2G
or 3G network), or Diameter for 4G generation. The following attributes of the data
would have been organized in messaging system entities (table 1):</p>
        <p>The data streaming of MAP and Diameter protocols was organized using one of the
largest telecom company in Russia, and data were received in Kafka application in
thrift data format, during the job implementation converted in Json data format and
finally sent to another Kafka messaging system to topics according to event types
described above (table 1).</p>
        <p>The average size of the streaming dataset is really challenging with an average
number of 37300 records per second for MAP protocol and an average number of
24200 records per second for Diameter protocol, for a data stream of 61500 records
per second in total.</p>
        <p>To run the application it is necessary to have 3 files in the folder:
 startup script (run.sh is in the project in the ./src/main folder);
 main code implementation file
(stream-1.0-SNAPSHOT-jar-withdependencies.jar);
 folder with configuration file.</p>
        <p>The fragment of the main code of a prototype application is shown in Figure 4.
import scala.util.parsing.json_
import scala.collection.JavaConversions._
object Main extends Serializable {
def main(args: Array[String]): Unit = {
val sparkConf = new SparkConf()
sparkConf.set(“spark.streaming.stopGracefullyOnShutdown”,”tru
e”)
val sc = new SparkContext(sparkConf)
val prop = new StreamingContext(sc, Seconds(1))
val topicsSetGSM = prop.gsmmapTopicsSet
val topicsSetDia = prop.diameterTopicsSet
val kafkaParamsInGsm = prop.kafkaParamsIn.toMap + (
“key.deserializer” -&gt; classOf[StringDeserializer],
“value.deserializer” -&gt; classOf[GsmMapDecoder]
)
val kafkaParamsInDia = prop.kafkaParamsIn.toMap + (
“key.deserializer” -&gt; classOf[StringDeserializer],
“value.deserializer” -&gt; classOf[DiameterDecoder]
)
val metricsAccumulator =</p>
        <p>MetricsAccumulator.getInstance(prop.elasticHosts,
prop.elasticIndexName, prop.elasticDateFormat)
val kS = KafkaUtils.createDirectStream[String,</p>
        <p>TGsmMapRecord](
ssc,
PreferConsistent,
Subscribe[String, TGsmMapRecord](
topicsSetGsm, kafkaParamsInGsm)).
mapPartitions(part =&gt; {</p>
        <p>part.map(_.value(()
Fig. 4. The fragment of Spark code.</p>
        <p>As a result of Spark job implementation, the following results of YARN resources
are allocated for this particular Spark job:
 Allocated memory is 6144Mb;
 Allocated CPU VCores is 3;
 3 containers running;
 average Scheduling Delay is 14 ms;
 average Processing Time is 464 ms;
 total Delay is 478 ms.</p>
        <p>The data can be viewed using Kafka UI tool (Fig. 5). It shows the messages in
topic. The result of the application is data streaming with fields as it is shown in
figure. Notice that IMSI and MSISDN fields that represent “private” information are
cut from the image, but the other informative fields are here.</p>
        <p>The approach was validated by comparison of amount of filtered messages for
particular period. For this purpose, we were using code in Python to connect to Kafka
application and to take two regarded data streams – Map and Diameter protocols. We
analyzed the same period of time of messages in Kafka and filtered them using the
same rules that were used for Spark application. The amount of filtered messages
using Python code and Spark was equal.
4.3</p>
      </sec>
      <sec id="sec-4-3">
        <title>Experimental Results</title>
        <p>The comparison with traditional cellular network monitoring system and batch
processing of CN data shows the advantages of the prototype system based on Apache
Spark application.</p>
        <p>The Spark streaming has many advantages comparing with usual batch streaming:
 average scheduling delay (ASD) for batch processing is much longer than
Spark delay. Apache Spark streaming runs its jobs with only 0.015 seconds
delay, while traditional batch processing has 0.5 seconds delay in average;
 average processing time (APT) shows that the same amount of data might be
processed in 45 seconds intervals, while Spark streaming process data in
0.464 time intervals. It is achieved because Spark jobs runs each second, and
the data processing is really fast, in-memory and efficient;
 average interval time (AIT) between Spark jobs is 1 second, while interval
between batch jobs is usually 1 minute. Batch processing cannot run faster
because of overheads before job start. Each start of job takes some additional
resources and needs some time to start job itself. For batch processing it is
larger than for streaming;

11
average available time (AAT) for decision makers to trigger roaming users. It
shows that time to make a decision about some action against “caught” users
is larger with Spark due to its faster computation comparing to the traditional
batch processing. Usually telecom providers are interested in users with no
more than 15 minutes delay when the event occurred that is the user crossed
the country border and this event has been caught by system. After we have
this event in messaging system, all the rest is depending on us – how fast we
process data and filter it for triggering and sending notifications. Thus, if we
consider average time between appearance of event in messaging system and
this event filtered - Spark shows only 1 second delay on average, while
traditional batch requires more than 1 minute and 15 seconds. It means that
decision makers can have more time to understand this user, his behavior and
to decide on sending any notifications.</p>
        <p>The results of comparative study of the DBDF-based Spark streaming monitoring
versus traditional batch-based monitoring service can be described as follows (Fig. 6):
Traditional system monitoring, batch processing
Apache Spark
000 401
,05 ,00
0
0
0
0
,
5
4
0
4
6
4
,
0
0
0
0
0
,
0
6
0
0
0
0
,
1
50 900
,4413 ,145
A S D , S E C S</p>
        <p>A P T , S E C S</p>
        <p>A I T , S E C S</p>
        <p>A A T , M I N S</p>
        <p>Thus, all DBDF components are scalable, and with adding more compute nodes to
a cluster more performance gains can be obtained:
 storage and processing memory are scalable for all DBDF components that
represents a significant advantage over traditional standalone monitoring
systems. The parameter is important because of constant traffic growth
worldwide;
 reliability of the solution is supported by the fact that all data are replicated
in a cluster that makes framework reliable. In case of node failures the data is
not lost;
 speed and performance shows the huge difference. Because of the cluster
mode and in-memory computations, DBDF is processing data very fast,
while traditional standalone systems usually perform much slower;
 amount of possible use cases is not limited with DBDF – all monitoring data
are collected and stored, and many new use cases can be created and
discovered. In traditional system usually one system is covering one use
case, or one department. With DBDF, new use cases can be easily
implemented with all company departments based on processing rules
(online streaming) or new hypotheses validation (offline streaming);
 data access is usually strict in traditional monitoring systems, while with
DBDF all departments can have access to all monitoring data and achieve
synergy effect all together. It means that departments can work together for
new use cases adaptation and verification;
 new hypotheses checks are almost not possible with traditional systems
because of not all monitoring protocols are presented in place. In contrast,
DBDF open up new horizons with petabytes of data exploration.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>To manage the complexity of the cellular network’s large volumes of streaming data it
is extremely important to deploy a powerful framework for data processing,
aggregation, enhancement, enrichment and storage. The adaptation of the distributed
Big Data driven framework for the smart monitoring services in telecom provider
environment and deployment of proposed architectural components will help to
achieve effective, reliable, scalable, high-speed and secure processing of the cellular
network data. The proposed DBDF framework is fully capable of dealing with
highloaded CN data streams and can be considered a foundation for future models
creation, making sure that all data are reliably saved and not lost.</p>
      <p>The proof of concept was achieved by creating a near real-time Big Data based
prototype application for roaming users detection with processing performance of
above 60 000 events per second. The prototype monitoring system has been created
using Apache Spark application and the adequacy of the model was checked by the
test application in Python programming language. The created prototype system has
revealed that Apache Spark is capable of handling thousands and even more events
per second and may be considered a foundation for real-time Big Data hub creation.</p>
      <p>The comparisons of the prototype DBDF-based smart monitoring system with
traditional standalone monitoring systems in cellular network demonstrates many
benefits of the proposed framework such as its scalability, reliability, speed and
performance, possibility to check new hypotheses. Apache Spark streaming facility in
sample application for roaming user detection demonstrated less delay, less
processing time and more time for decision makers comparing with traditional batch
processing.</p>
      <p>The implementation of the DBDF framework can be repeated for any telecom
operator using the same protocols, provided the cluster has the same performance to
maintain stable work of application. Or, the data stream can be also proportionally
decreased along with the amount of nodes in a cluster and their capacity.</p>
      <p>As future developments of the DBDF for the cellular network data processing and
analysis we plan to develop, integrate and test a Serving layer of the Lambda-enabled
implementation of the smart monitoring system for the CN data.</p>
      <sec id="sec-5-1">
        <title>Available at:</title>
        <p>Accessed on:</p>
      </sec>
      <sec id="sec-5-2">
        <title>Kafka 1.0</title>
      </sec>
      <sec id="sec-5-3">
        <title>Documentation.</title>
      </sec>
      <sec id="sec-5-4">
        <title>Kafka 1.0</title>
      </sec>
      <sec id="sec-5-5">
        <title>Documentation.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. https://www.statista.com/statistics/274774/forecast-of
          <article-title>-mobile-phone-users-</article-title>
          <string-name>
            <surname>worldwide</surname>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Number of mobile phone users worldwide from 2013 to 2019 (in billions)</article-title>
          ,
          <source>Accessed on: 2017-10-11.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. http://www.3gpp.
          <source>org/about-3gpp. Accessed on: 2018-01-31.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. https://www.itu.int/en/itutelecom/Pages/default.aspx. Accessed on:
          <fpage>2018</fpage>
          -01-31.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. http://www.pictogramdesign.com/websites/nti_overview/docs/ServiceAssurance_OV_
          <fpage>0826</fpage>
          <lpage>10</lpage>
          .
          <article-title>pdf NetBoss Technologies Integrated Service Assurance</article-title>
          .
          <source>Pictogram Digital Design. Accessed on: 2019-03-17.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Darema</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Dynamic Data Driven Applications Systems: A New Paradigm for Application Simulations and Measurements</article-title>
          . International Conference on Computational Science. M.
          <string-name>
            <surname>Bubak</surname>
          </string-name>
          et al. (Eds.):
          <source>ICCS</source>
          <year>2004</year>
          , LNCS 3038, pp.
          <fpage>662</fpage>
          -
          <lpage>669</lpage>
          ,
          <year>2004</year>
          . © SpringerVerlag Berlin Heidelberg
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Marz</surname>
          </string-name>
          , Nathan and Warren,
          <string-name>
            <surname>James</surname>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Big Data: Principles and best practices of scalable realtime data systems</article-title>
          , 1st ed..
          <source>Manning Publication Co.</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>7. The Lambda architecture: principles for architecting realtime Big Data systems, blog post by James Kinle</article-title>
          . Available at: http://jameskinley.tumblr.com/post/37398560534/thelambda-architecture
          <string-name>
            <surname>-</surname>
          </string-name>
          principles-for Accessed on:
          <fpage>2018</fpage>
          -02-05.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Lambda</given-names>
            <surname>Architecture</surname>
          </string-name>
          :
          <article-title>A state-of-the-art, post by Pere Ferrera</article-title>
          . http://www.datasalt.com/
          <year>2014</year>
          /01/lambda
          <article-title>-architecture-a-state-of-the-art/</article-title>
          <year>2018</year>
          -02-05.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>A.</given-names>
            <surname>Suleykin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Panfilov</surname>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>The Simulation-Based Smart Management Approach for Cellular Network Operation and Planning, in: Annals for DAAAM for 2017</article-title>
          &amp;
          <article-title>Proceedings</article-title>
          , DAAAM International, Viena,
          <year>2017</year>
          , pp.
          <fpage>0423</fpage>
          -
          <lpage>0432</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. http://window.edu.ru/catalog/pdf2txt/503/80503/60870, p.
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          . Accessed on:
          <fpage>2018</fpage>
          -03-21.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11. https://kafka.apache.org/documentation.html#introduction. Accessed on:
          <fpage>2018</fpage>
          -02-04.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. http://spark.apache.org.
          <source>Apache Spark. Accessed on: 2018-02-04.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. https://databricks.com/spark/about/. Accessed on:
          <fpage>2018</fpage>
          -03-08.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. https://kafka.apache.org/documentation.html#introduction. Accessed on:
          <fpage>2018</fpage>
          -02-04.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15. http://storm.apache.org.
          <source>Apache Storm. Accessed on: 2018-02-04.</source>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. http://spark.apache.org.
          <source>Apache Spark. Accessed on: 2018-02-04.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. https://www.python.org/. Accessed on:
          <fpage>2018</fpage>
          -03-08.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18. http://airflow.apache.org/.
          <source>Accessed on: 2018-03-08</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. http://hadoop.apache.org. What Is Apache Hadoop? Accessed on:
          <fpage>2018</fpage>
          -02-04.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. http://libreportal.net/data-warehousing/apache-hive.
          <source>html. Accessed on: 2018-03-07.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>