<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DAZIO: Detecting Activity Zones based on Input/Output call and SMS activity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Miguel Nu n˜ez-del-Prado-Cortez</string-name>
          <email>m.nunezdelpradoc@up.edu.pe</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Luna</string-name>
          <email>ae.lunaa@up.edu.pe</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Romain Gauthier</string-name>
          <email>romain.gauthier@intersec.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Instersec Labs</institution>
          ,
          <addr-line>Par ́ıs -</addr-line>
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universidad del Pac ́ıfico</institution>
          ,
          <addr-line>Av. Salaverry 2020, Lima - Per u ́</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidad del Pac ́ıfico</institution>
          ,
          <addr-line>Av. Salaverry 2020, Lima -</addr-line>
          <country country="PE">Peru ́</country>
        </aff>
      </contrib-group>
      <fpage>130</fpage>
      <lpage>139</lpage>
      <abstract>
        <p>Mobile telecoms operators possess an enormous quantity of data, which could be used to reduce the cost of installing new infrastructure, to provide a better QoS or to plan their infrastructure. Thus, they are concerned to model, understand and predict SMS and calls activity levels in their infrastructures. Besides, SMS and call activities analysis can open new business opportunities for geomarketing as well as trade area analysis. In the present effort, we detected activity zones with a difference of only 0.5 km from the reference activity areas extracted from Geo-tweets. We also used Markov chains to represent and predict SMS and call activity levels, achieving a prediction success rate between 80% and 90%.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 INTRODUCTION</title>
      <p>
        Telecoms data is a rich information source for
many purposes, ranging from urban planning
        <xref ref-type="bibr" rid="ref19">(Toole et al., 2012)</xref>
        , human mobility patterns
        <xref ref-type="bibr" rid="ref4 ref6">(Ficek and Kencl, 2012; Gambs et al., 2011)</xref>
        ,
points of interest detection
        <xref ref-type="bibr" rid="ref20">(Vieira et al., 2010)</xref>
        ,
epidemic spread modeling
        <xref ref-type="bibr" rid="ref12">(Lima et al., 2013)</xref>
        ,
community detection
        <xref ref-type="bibr" rid="ref13">(Morales et al., 2013)</xref>
        ,
disaster planning
        <xref ref-type="bibr" rid="ref16">(Pulse, 2013)</xref>
        and social interactions
        <xref ref-type="bibr" rid="ref2">(Eagle et al., 2013)</xref>
        .
      </p>
      <p>One common effort for these applications is to
determine dense areas where many users stay for a
significant amount of time, namely activity zones.
Another task is to identify contiguous zones
relaying activity zones (i.e., transit zones). Thus, in our
context, the detection of activity and transit zones
as well as the interaction between identified
activity zones are crucial tasks for reducing the cost
of installing new infrastructure, to provide a better
QoS or to plan their infrastructure.</p>
      <p>Therefore, in the present paper, we will identify
activity and transit zones to monitor and to predict
the activity levels in the telecoms operators
network. These monitoring and prediction are based
on the SMS and calls input/output activity levels
issued from the Telecoms Italia Big Data
Challenge1. The results of the present study is directly
applied for: (1) targeting advertisement to
activity zones; (2) proposing a suitable place to open a
new store in a city or (3) planning where to add
cell towers to improve QoS.</p>
      <p>In the present effort, we describe a methodology to
detect activity and transit zones. More precisely,
the contribution of this work is twofold. On one
hand, we present an Activity Markov chain model
to represent activity levels. On the other hand, we
predict future activity levels using the
aforementioned model. The rest of the paper is organized
as follows. First, Section 2 describes the related
works on activity zones detection. Then, Section
3 presents the datasets we use for experiments.
Next, Section 4 introduces our technique to detect
and model activity zones as well as the approach
to forecast activity levels. Section 5 shows
correlation measurements of activity levels versus
pollution and weather conditions. Finally, Section 6
concludes the paper and depicts some future
directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        Dense areas detection has been studied from a
human mobility point of view, using fine grain
and coarse-grained location data. As an example
of fine-grained location, the work of
        <xref ref-type="bibr" rid="ref6">(Gambs et
al., 2011)</xref>
        use mobility traces of 172 Yellow Cabs
1Telecoms Italia Big Data Challenge
website: www.telecomitalia.com/tit/en/
bigdatachallenge.html
Taxis, issued from GPS, in San Francisco Bay
        <xref ref-type="bibr" rid="ref15">(Piorkowski et al., 2009)</xref>
        to detect taxi’s point of
interests (POI). These POIs are equivalent to activity
zones, which tend to be zones with high
pedestrian presence. The authors rely on the
beginend heuristic
        <xref ref-type="bibr" rid="ref5">(Gambs et al., 2010)</xref>
        and
clustering algorithms, such as Density Joinable
        <xref ref-type="bibr" rid="ref21">(Zhou et
al., 2004)</xref>
        , Density Time
        <xref ref-type="bibr" rid="ref8">(Hariharan and Toyama,
2004)</xref>
        and Time Density clustering
        <xref ref-type="bibr" rid="ref5">(Gambs et al.,
2010)</xref>
        to detect POIs in San Francisco city.
Another fine-grain data used to detect activity areas
are issued from Geo-social networks.
      </p>
      <p>
        The work of
        <xref ref-type="bibr" rid="ref17 ref2">(Qu and Zhang, 2013)</xref>
        uses
Foursquare’s check-ins from 446 users during ten
months for identifying trade areas. They rely on
four different techniques like Center of Mass
location, the most commonly checked-in location,
the place with the highest check-in density and
the center of mass of the most frequently visited
location cluster. The algorithm use for
clustering is DBSCAN
        <xref ref-type="bibr" rid="ref3">(Ester et al., 1996)</xref>
        . Once they
have identified the activity centers, they mark the
boundary of the area using drive-time/distance
polygon
        <xref ref-type="bibr" rid="ref11">(Kures and Pinkovitz, 2011)</xref>
        . This
technique consists of computing the decay distance
from a given store to home or work (authors
assume that the two most checked-in places are
home and work). The drawback of this approach
is that the selected users are conditioned to
checkin in the store under study. Thus, the dataset is
biased.
      </p>
      <p>
        Other works use coarse grain location from Call
Data Records (CDR). For instance, the work of
(Isaacman et al., 2011) uses Hartigan’s leader
clustering algorithm
        <xref ref-type="bibr" rid="ref9">(Hartigan, 1975)</xref>
        to identify dense
areas. First, authors sort antennas by the amount
of time that phones contact the antenna. Once data
is sorted, the clustering algorithm takes the first
antenna as the centroid of a cluster. Then, it
verifies if the next antenna is within a distance d from
the centroid. If it is not the case, the antenna
becomes the centroid of a new cluster. In the case
the antenna is within the distance, the algorithm
computes the new centroid as the weighted
average. They repeat the process until all antennas
belong to a cluster. Researchers use CDR locations
of 97 and 71 thousand unique users in Los
Angeles and New York cities collected over 2 and a
half months as well as 19 volunteers as the ground
truth to validate their results. They were able to
estimate dense areas with an error of 3 miles
compared to the ground truth.
      </p>
      <p>
        Another, more refined technique to identify dense
areas respecting natural tessellation is presented
by
        <xref ref-type="bibr" rid="ref20">(Vieira et al., 2010)</xref>
        . Authors use CDR
locations from calls of one million users during four
months over an area of 80 000 km2. They propose
a method composed of three phases: the first step
is the graph construction, which relies on
Delaunay triangulation
        <xref ref-type="bibr" rid="ref1">(Dobkin and Laszlo, 1987)</xref>
        . The
triangulation algorithm makes connexions (edges)
between near antennas (vertex) maximizing the
size of the angles of the triangles. Once the graph
is built, all edges are weighted by the total
activity and by the number of users of both connected
antennas. The second phase is the computation of
dense areas based on a maximum spanning tree
build using the Kruskal algorithm (Kruskal, 1956).
Taking as input the weighted graph G, the idea
behind this algorithm is to find a subgraph of G,
which maximizes the density and does not
contain any cycle. At last, the post-processing phase
uses
        <xref ref-type="bibr" rid="ref18">(Shiloach and Vishkin, 1982)</xref>
        algorithm to
establish groups of antennas representing dense
areas. Thus, the algorithm groups adjacent vertex
from previously computed sub-graph to find a set
of close vertex (a set of antennas). The authors
validate their results empirically based on the subway
structure of the region under study. Inspired by
these works, we propose a novel ad-hoc
methodology to find activity (dense) zones as well as to
model and forecast their activity levels using the
data provided by the Telecom Italia Big Data
Challenge (TIM challenge). We describe this dataset in
the next section.
3
      </p>
      <p>DATASET
Datasets provided by the TIM challenge were
collected in the cities of Milan and Trento over
November and December 2013. Our study only
takes into account the dataset gathered from Milan
city. Nevertheless, all the analysis and the
methodology could be generalized for any city. In the next
subsections, we will describe the datasets provided
for Milan city. These datasets were used to detect
and model the activity zones (primary datasets)
and to find some correlation between activity and
other measures like air quality and weather
conditions (auxiliary datasets). It should be noted that
space was discretized in a grid, and all measures
are normalized to correspond in one square of the
grid.
3.1</p>
      <p>Main datasets</p>
      <sec id="sec-2-1">
        <title>The datasets of Milan city we use to detect and</title>
        <p>model the activity zones are:
Milan Grid is a geographical segmentation over
the city to aggregate the measurements of the
other datasets. The area of each square is
55 225 m2, and it has 10 000 squares in the
form of a point (x, y) and the latitude and
longitude belonging to this x, y position. An
example of the described data as well as the
grid over Milan are introduced in Table 1 and
Figure 1, respectively.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Point</title>
        <p>x1, y1</p>
      </sec>
      <sec id="sec-2-3">
        <title>Latitude 9.011</title>
      </sec>
      <sec id="sec-2-4">
        <title>Longitude 45.568</title>
        <p>Telecommunications (SMS, Call and Internet)
provides information about the activity of a
square concerning received and sent SMS,
incoming and outcoming calls as well as
internet usage. This data is temporal
aggregated in timeslots of ten minutes and
provides the measure of the activity of a
given event as well as the square id (c.f. Table</p>
      </sec>
      <sec id="sec-2-5">
        <title>2). This kind of information is organized in a way that SMS-in and SMS-out activity scale are given in arbitrary units, and their values range from 0 to 1.</title>
        <p>Id
1
SMS-out
0.16</p>
        <p>Time
1383265200</p>
        <p>Call-in
0.108</p>
        <p>Country</p>
        <p>39
Call-out
0.026</p>
        <p>SMS-in
0.24
Data
6.83
number of vehicles that belongs to Milan, the
number of vehicles that is not from Milan,
the number vehicles with engine ignition
systems, in movement and stopped (c.f.
Table 3).</p>
        <p>Id
60
Avg speed</p>
        <p>24
Non-Mi plates
62</p>
        <p>Time
17/12/13 18:00</p>
        <p>Std speed</p>
        <p>95
Ignition
2</p>
        <p>Direction
WEST
Mi plates</p>
        <p>21
Mov/Stopped
2/0</p>
      </sec>
      <sec id="sec-2-6">
        <title>We also used additional datasets to analyze activity zones, dynamics, and correlations, as we show in the following paragraphs.</title>
        <p>Telecommunications - MI to MI provides
information regarding the directional interaction
strength, between the city of Milan and
different areas based on the calls exchanged
between Telecom Italia Mobile users. More
precisely, this dataset contains the origin and
destination Id squares, the time and the
directional interaction strength i.e., Activity (c.f.
Table 4)</p>
        <p>Id 1
1</p>
        <p>Id 2
3
Precipitation describes the intensity and the
precipitation type over the city of Milan. In
more detail, the dataset uses a coarse
spatial aggregation by dividing Milan city into
four quadrants (northeast, northwest,
southeast and southwest). The intensity value of
the phenomenon is between 0 and 3, the
percent of coverage of a given quadrant and the
precipitation type between 0 and 2, where 0
means absence of precipitation, 1 is rain and</p>
      </sec>
      <sec id="sec-2-7">
        <title>3 is snow (c.f. Table 5).</title>
        <p>Type
1
0
0
1
Air Quality describes the air pollution
monitoring system of Milan city obtained by
using various types of sensors located within
the city limits. This environmental dataset
measures a different kind of contamination
agents, such as Ammonia, Nitrogen Dioxide,
Total Nitrogen, Particulate Matter 2.5 µm
(PM2.5), Particulate Matter 10 µm (PM10),
Benzene, Sulphur Dioxide, Black Carbon,
Carbon Monoxide and Ozone. An
example of pollution measure is given in Table
6, where the characteristics of this particular
sensor are in Table 7.
Social Pulse contains data derived from an
analysis of geolocalized tweets originated in
Milan. This dataset provides a user id,
DBPedia entity, tweets language, municipality,
time, timestamp and location (c.f. Table 8).</p>
        <p>User
5fa4b1cc71
Municipality</p>
        <p>Milan</p>
        <p>Entities
Halloween
Timestamp
1383260474</p>
        <p>Language</p>
        <p>En
Lat, Lon
9.21, 45.49
Based on the aforementioned datasets, we have
implemented our experiments using the main and
the auxiliary datasets. These experiments are
detailed in sections 4 and 5, respectively.
4</p>
        <p>EXPERIMENTS
In the present section, we describe our
methodology to discover, model and predict the behavior
of a zone. We distinguish two different areas, the
activity zone, where people stay on a regular
basis for a significant amount of time and the transit
zone which is the area used by individuals to go
from one activity zone to another. In the next
subsections, we describe how to recognize an activity
zone from a transit zone (Subsection 4.1), how to
model activity levels (Subsection 4.2) and finally
how to predict them (Subsection 4.3).
4.1</p>
        <p>
          Detecting activity levels
The basic idea behind this method is to have a
good representation of the activity variation levels
over the time. Activity levels could be classified
in three different degrees, low, medium and high.
Cumulative distribution of incoming/outcoming
SMS and call activity levels illustrated by a Heat
map over Milan city, as shown in Figure 2, were
used to analyze data. The objective is to,
empirically, find a suitable threshold to distinguish a
square with high activity level from a square with
medium or low activity levels represented as green
and red in Figure 3, respectively.
Figure 3 depicts the cumulative distribution of
the aggregated incoming and outcoming SMS and
call activity of the telecommunications dataset (c.f.
Subsection 3.1). The Heat map is built for an
activity threshold of 25 units. Based on this
visualization technique, that amount of units seems to
be a good trade-off between compact and
wellseparated activity zones. Heat maps were used
to represent represent tourist activity as shown by
Olteanu et al.
          <xref ref-type="bibr" rid="ref14">(Olteanu et al., 2011)</xref>
          .
        </p>
        <p>
          In order to detect groups of squares representing
an activity zone, we can use a high activity
threshold (c.f. Subsection 4.2). In addition, we study, in
detail, the activity over work hours to analyze the
difference between busy and idle squares. Figure
4 shows the difference in activity levels between
activity and transit zones. From 8 AM to 8 PM the
activity is considerably high, that is why that area
is composed of a vast number of squares during
the day. On the other hand, transit zones display a
much lower activity level and a higher fluctuation
throughout the day.
Taking into account the elements as mentioned
earlier, we use the Heat map and activity threshold
presented in Figure 3 for detecting high activity
zones. Nevertheless, we need to define the borders
of these activity zones. From Figure 4, we can
infer that this irregularity of transit zone represents
movement. Thus, the Cobra dataset gathers the
information about the movement of vehicles and
we depicted this information in a Heat map over
Milan city in Figure 5. The speed combined with
the activity level allowed us to detect the activity
zones, as well as their borders. As the result of
the combination of this two variables, we obtained
28 activity areas and the centroids of these activity
zones which are shown in Figure 6 in blue color.
Since we do not have the ground truth, we used the
geolocalized tweets dataset to verify the accuracy
of our methodology. Applying DBSCAN
          <xref ref-type="bibr" rid="ref3">(Ester
et al., 1996)</xref>
          clustering algorithm with at least 5
points per cluster within a radius of 3 km over
8 282 users (i.e.,109 762 geolocalized tweets). We
obtained 24 clusters depicted as green points in
Up to this point, we are able to identify activity
zones. Thus, the next tasks are to model the
behavior activity levels in the detected regions as well as
to predict the activity levels in the identified
activity zones.
4.2
        </p>
        <sec id="sec-2-7-1">
          <title>Modeling activity levels</title>
          <p>Relying on the thresholds mentioned above, we
can build an Activity Markov chains model to
represent the change of activity levels over time. An
Activity Markov chain model is a stochastic
process where the changes of states are related to a
probability associated with various state changes
(transition probabilities). In our case, an
Activity Markov chain is a probabilistic automaton (PA)
model that represents, in a compact way, the
occupation (activity) of a square or activity zone. The
nodes symbolize the state (low, medium or high)
of the squares or zones ([activity level] [zone
code], ex: L A) and edges, weighted with a
probability, represent the transition from one state to
another over time windows. This model could be
expressed in the form of a graph or the form of a
transition matrix (c.f. Figure 9).</p>
          <p>The process for building an Activity Markov
model is divided into two stages. The first one
is basically to order the events in a chronological
way. Then we classify them as low, if they have
less than 15 activity units, as medium if activity
units are between 15 and 25 and as high if there
are more than 25 activity units. Then the
transition matrix is built by counting the variations from
one level to another, taking care to avoid loops.
When events are not recorded anymore, the matrix
is normalized to obtain the transition probabilities.
As shown in Figure 8, we have divided the time
into 4 different windows depending on the range
of time studied, each one has 6 hours and starts
at 6:00 am; for both weekdays and weekends;
giving a total number of eight windows. Furthermore,
in each time window interactions between
different levels of activity are also modeled. Moreover,
we matched, after a model processing, an activity
zone of a time window with another (blue arrows).
We finally conclude, from the stationary vector of
Markov chains, that activity areas are occupied in
only 11% and free in 71%. It is important to point
out that, for improving the accuracy of the
analysis, it would be better to divide the time slots in
less intervals instead of 6 hours.</p>
          <p>Until this point we showed that we are able to
identify high activity squares, activity and transit
zones. In the next subsections, we detail how to
predict, in an unprecedented way and with a very
acceptable rate of success, not only activity levels
but also their possible changes.
4.3</p>
        </sec>
        <sec id="sec-2-7-2">
          <title>Prediction of activity levels</title>
          <p>
            Anticipating high activity levels within an
“activity area” allows Telecoms operators to plan or
avoid unnecessary investment in infrastructure, as
well as ensure the QoS or to start a new
entrepreneurship. In this paper, our prediction,
inspired from the work of Gambs et al.
            <xref ref-type="bibr" rid="ref7">(Gambs et
al., 2012)</xref>
            , is performed using the transition matrix
and allows us to obtain changes of activities
between temporary windows and within them. Then,
predictions of activity changes within the same
window and estimations can be made. For
example, we wonder what is the probability of passing
from a low activity state to a high one or what is
the likelihood for a given activity to remain in the
same state when we consider the next window and
the same range of time. To answer these questions,
we rely on the algorithm presented in Algorithm 1.
          </p>
          <p>Algorithm 1: Prediction algorithm</p>
          <p>Data: TransitionMatrix, indexRow,</p>
          <p>inTimeWindows</p>
          <p>Result: indexColumn
1 if inTimeWindows then
2 //predict next state in the time windows
3 i=maxOutgoingProbOnWin(indexRow)
4 indexColumn = i
5 else
6
//predict the next state when the time
window change
i=maxOutgoingProbOnNextWin(indexRow)
indexColumn = i
More precisely, Algorithm 1 takes as input a
transition matrix (transitionMatrix). Where the index
of the row in the transition matrix corresponds to
the actual state of the system (indexRow) and a
boolean value is used to indicate whether the
prediction is local (inTimeWindows). Based on these
inputs, the algorithm returns the maximal
outgoing probability from the transition matrix taking
into account only columns corresponding to the
same time windows of the index row (local
transition from line 1 to 4 of the Algorithm 1). For
instance, in Figure 8, if the actual state is medium on
the time windows from 18:00h to 0:00h on
weekdays, the prediction algorithm will give an output
in the high level in the same time window. Another
kind of prediction is to take into account others
columns instead of those that belong to the same
time window (inter time windows transition from
line 6 to 8 of the Algorithm 1). Given the low in
the time windows from 18h to 0h on weekdays in
Figure 8, the algorithm will output the low state on
time windows from 0h to 6h on weekends. In the
case of the output we have the same probability,
ties are break randomly.</p>
          <p>To validate the accuracy of the predictions, we
used data from the whole month of November as
training set and the first 16 days of December as
testing set (we did not take into account New Year
celebration to avoid special dates that separate our
study from normal behavior). The results are
depicted in Figure 10, where the success rate
(Equation 1) is the ratio between the correct prediction
and the total number of predictions and the
number of overall forecast. Note that the number of
predictions is indicated in parentheses at the
bottom part of each bar.</p>
          <p>success rate =
#goodprediction
#predictions
(1)</p>
          <p>We observe, from Table 9, the success rate for
both kinds of predictions, namely (1) within a
time window (local) and (2) in different time
windows (trans.). It is important to note that there
are two distinct scenarios; the first one considers
both local and trans. predictions based on Activity
Markov chains models which were built from
activity levels of squares; while the second scenario
takes into account activity levels from detected
activity zones to forecast future values. We did not
performed k-fold cross validation since the
training set of a month is representative of the
mobility pattern. Thus, adding more mobility traces to
the training test does not contribute to increase the
success rate.</p>
          <p>So far, we are able to model, identify and predict
activity levels in activity zones. In the next section,
we are going to use the auxiliary datasets to
analyze activity zones interaction and to study
possible correlation with the levels or evaluate how the
weather impact the utilization of the Telecoms
service.
5</p>
          <p>PLAYING WITH OTHER DATASETS
In the present section, we will study the
interaction between detected activity zones (Subsection
5.1); the correlation between activity levels versus
pollution measures (Subsection 5.2) and the
influence of the weather on the activity levels in the
Telecoms operator infrastructure (Subsection 5.3).
5.1</p>
          <p>Interaction between activity zones
Using the directional interaction activity dataset
between zones in the area of Milan, we plotted a
graph to visualize the communication exchange,
as well as various activity levels as we can
appreciate in Figure 11, where the width of the edges
accounts for the logarithm of the aggregated
activity for the whole month of November. To extend
the semantic of this graph, we modulated the size
of the nodes according to the amount of tweets
emitted from the corresponding zone taking into
account global pulses dataset (c.f. Subsection 3.2).</p>
          <p>We observe that Milan city has a star topology,
where there is a central node that communicates
with the other peripheral nodes. Another
interesting fact is that small nodes tend to communicate
to the central node. Nevertheless, there are a few
exchanges between small contiguous nodes.
In this subsection, we study the correlation
between the activity level presented on the
telecommunication dataset and air quality measurements
(both described in Subsection 3.2) to forecast the
pollution level of the Milan city based on the
activity of the telecom operator antennas. Figure 12
shows the results of the correlation of the activity
with respect to different polluting gasses as well
as the number of vehicles in movement, ignition
or stopped.</p>
          <p>We found out that the activity has a positive
correlation with PM10 (particulates matter with a
diameter of 10 microns or less) and PM2.5 pollution
measures (fine particles of 2.5 micrometers of
diameter or less). From Figure 13 we can visualize
that the activity levels has a positive correlation
with the radiation measures and a negative
correlation with the relative humidity.
5.3 Influence of weather on the activity
We compare the outgoing SMS and call activity
in the presence of different weather phenomena’s
scenarios, like rain, snow or the absence of both.
For this purpose, we used the Precipitation dataset
(c.f. Subsection 3.2).</p>
          <p>From Figure 14, we can observe that people send
more SMS and give more calls in presence of rainy
weather, even if rain is slight (blue, red, yellow
and green bars) than in normal conditions (dark,
red and light blue bars). Nevertheless, people tend
to call or send less SMS when it is snowing (dark
and light green bars).
6 CONCLUSION AND FUTURE</p>
          <p>DIRECTIONS
Our purpose in this research is to understand the
activity of the telecommunication network by
analyzing several aspects of Milan’s phone traffic
flows. We were interested in the definition and
morphology of the activity and transit zones; the
prediction of activity levels over different regions
with a success rate between 80% and 90%; the
interactions between the different activity zones and
the influence of the weather and the pollution on
that activity. Thus, our results offer a new way of
looking at the telecommunication traffic data by
examining the various connections between
apparently uncorrelated datasets, providing insights to
manage and to optimize the whole network. For
business opportunities, this means (1) new
geomarketing opportunities through a better
understanding of users communication patterns, (2) new
trade area analysis, (3) cheaper network load
balancing as well as (4) improved QoS. In the future,
we would like to study and analyze in detail the
opinions (sentiment analysis) discussed by users
and generators of tweets and identifying the
geolocation of these activity areas. Another line of
investigation would include coarse mobility Call
Data Record (CDR) to take into account as an
additional element for detecting activity and traffic
areas.</p>
          <p>Alexander Varshavsky. 2011. Identifying important
places in peoples lives from cellular network data. In
Pervasive Computing, June.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Dobkin</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Laszlo</surname>
          </string-name>
          .
          <year>1987</year>
          .
          <article-title>Primitives for the manipulation of three-dimensional subdivisions</article-title>
          .
          <source>In Proceedings of the Third Annual Symposium on Computational Geometry</source>
          , pages
          <fpage>86</fpage>
          -
          <lpage>99</lpage>
          , New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Nathan</given-names>
            <surname>Eagle</surname>
          </string-name>
          , Alex S. Pentland, and
          <string-name>
            <given-names>David</given-names>
            <surname>Lazer</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Inferring social network structure using mobile phone data</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          ,
          <volume>106</volume>
          (
          <issue>36</issue>
          ):
          <fpage>15274</fpage>
          -
          <lpage>15278</lpage>
          ., September.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Martin</given-names>
            <surname>Ester</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hans-Peter Kriegel</surname>
            , Joerg Sander, and
            <given-names>Xiaowei</given-names>
          </string-name>
          <string-name>
            <surname>Xu</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>A density-based algorithm for discovering clusters in large spatial databases with noise</article-title>
          .
          <source>In Second International Conference on Knowledge Discovery and Data Mining</source>
          , pages
          <fpage>226</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Ficek</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Kencl</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Inter-call mobility model: A spatio-temporal refinement of call data records using a gaussian mixture model</article-title>
          .
          <source>In Conference on Computer Communications</source>
          , pages
          <fpage>469</fpage>
          -
          <lpage>477</lpage>
          ,
          <year>March</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Gambs</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>O. Nun</surname>
          </string-name>
          <article-title>˜ez Killijian, and</article-title>
          <string-name>
            <given-names>M.. del Prado</given-names>
            <surname>Cortez</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Gepeto: A geoprivacy-enhancing toolkit</article-title>
          .
          <source>In Advanced Information Networking and Applications Workshops (WAINA)</source>
          ,
          <year>2010</year>
          IEEE 24th International Conference on, pages
          <fpage>1071</fpage>
          -
          <lpage>1076</lpage>
          , April.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Se´bastien Gambs</surname>
          </string-name>
          ,
          <string-name>
            <surname>Marc-Olivier Killijian</surname>
          </string-name>
          , and
          <source>Miguel Nu´o`ez del Prado Cortez</source>
          .
          <year>2011</year>
          .
          <article-title>Show me how you move and i will tell you who you are</article-title>
          .
          <source>Transition on Data Privacy</source>
          ,
          <volume>4</volume>
          (
          <issue>2</issue>
          ):
          <fpage>103</fpage>
          -
          <lpage>126</lpage>
          ,
          <year>August</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>Se´bastien Gambs, Marc-Olivier Killijian, and Miguel Nu´n˜ez del Prado Cortez</article-title>
          .
          <year>2012</year>
          .
          <article-title>Next place prediction using mobility markov chains</article-title>
          .
          <source>In Proceedings of the First Workshop on Measurement, Privacy, and Mobility, page 3</source>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Ramaswamy</given-names>
            <surname>Hariharan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kentaro</given-names>
            <surname>Toyama</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Project lachesis: parsing and modeling location histories</article-title>
          .
          <source>In In Geographic Information Science</source>
          , pages
          <fpage>106</fpage>
          -
          <lpage>124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>John A. Hartigan. 1975. Clustering</given-names>
            <surname>Algorithms</surname>
          </string-name>
          . John Wiley &amp; Sons, Inc., New York, NY, USA, 99th edition.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Sibren</surname>
            <given-names>Isaacman</given-names>
          </string-name>
          , Richard Becke, Ramn Cceres, Stephen Kobourov, Margaret Martonosi, James Rowland, and
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Kruskal</surname>
          </string-name>
          .
          <year>1956</year>
          .
          <article-title>On the Shortest Spanning Subtree of a Graph and the Traveling Salesman Problem</article-title>
          .
          <source>In Proceedings of the American Mathematical Society</source>
          ,
          <volume>7</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>M</given-names>
            <surname>Kures</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Pinkovitz</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Downtown and business district market analysis</article-title>
          . http://fyi.uwex.edu/ downtown-market
          <article-title>-analysis/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Antonio</given-names>
            <surname>Lima</surname>
          </string-name>
          , Manlio De Domenico, Veljko Pejovic, and
          <string-name>
            <given-names>Mirco</given-names>
            <surname>Musolesi</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Exploiting cellular data for disease containment and information campaigns strategies in countrywide epidemics</article-title>
          .
          <source>Computing Research Repository, abs/1306</source>
          (4534):
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          , Jun.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Morales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Creixell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Borondo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Losada</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Benito</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Understanding ethnical interactions on ivory coast. In Data for Development (D4D) challenge</article-title>
          , pages
          <fpage>115</fpage>
          -
          <lpage>120</lpage>
          , May.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Ana-Maria</surname>
            <given-names>Olteanu</given-names>
          </string-name>
          , Roberto Trasarti,
          <string-name>
            <given-names>T</given-names>
            <surname>Couronn</surname>
          </string-name>
          , Fosca Giannotti, Mirco Nanni, Zbigniew Smoreda, and
          <string-name>
            <given-names>Cezary</given-names>
            <surname>Ziemlicki</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Gsm data analysis for tourism application</article-title>
          .
          <source>In Proceedings of the 7th International Symposium on Spatial Data Quality (ISSDQ).</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Michal</given-names>
            <surname>Piorkowski</surname>
          </string-name>
          , Natasa Sarafijanovic-Djukic, and
          <string-name>
            <given-names>Matthias</given-names>
            <surname>Grossglauser</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>CRAWDAD data set epfl/mobility (v.</article-title>
          <year>2009</year>
          -
          <volume>02</volume>
          -24). Downloaded from http://crawdad.org/epfl/mobility/, February.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>United</given-names>
            <surname>Nations Global Pulse</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Mobile phone network data for development</article-title>
          . http://www. unglobalpulse.org/research, October.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Yan</given-names>
            <surname>Qu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jun</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Trade area analysis using user generated mobile location data</article-title>
          .
          <source>In International Conference on World Wide Web</source>
          , pages
          <fpage>1053</fpage>
          -
          <lpage>1064</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Yossi</given-names>
            <surname>Shiloach</surname>
          </string-name>
          and
          <string-name>
            <given-names>Uzi</given-names>
            <surname>Vishkin</surname>
          </string-name>
          .
          <year>1982</year>
          .
          <article-title>An o(log n) parallel connectivity algorithm</article-title>
          .
          <source>Journal of Algorithms</source>
          ,
          <volume>3</volume>
          (
          <issue>1</issue>
          ):
          <fpage>57</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Jameson L. Toole</surname>
          </string-name>
          , Michael Ulm, Marta C.
          <article-title>Gonza´lez, and</article-title>
          <string-name>
            <given-names>Dietmar</given-names>
            <surname>Bauer</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Inferring land use from mobile phone activity</article-title>
          .
          <source>In Proceedings of the ACM SIGKDD International Workshop on Urban Computing</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          , New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>M.R. Vieira</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Frias-Martinez</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Oliver</surname>
            , and
            <given-names>E. FriasMartinez.</given-names>
          </string-name>
          <year>2010</year>
          .
          <article-title>Characterizing dense urban areas from mobile phone-call data: Discovery and social dynamics</article-title>
          .
          <source>In Second International Conference on Social Computing</source>
          , pages
          <fpage>241</fpage>
          -
          <lpage>248</lpage>
          ,
          <year>August</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Changqing</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Dan Frankowski, Pamela Ludford, Shashi Shekhar, and
          <string-name>
            <given-names>Loren</given-names>
            <surname>Terveen</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Discovering personal gazetteers: An interactive clustering approach</article-title>
          .
          <source>In International Workshop on Geographic Information Systems</source>
          , pages
          <fpage>266</fpage>
          -
          <lpage>273</lpage>
          , November.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>