<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>{m.nunezdelpradoc, bn.espositoa, ae.lunaa}@up.edu.pe</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Miguel Nunez-del-Prado</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bruno Esposito</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Luna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad del Pac ́ıfico Av.</institution>
          <addr-line>Salaverry 2020 Lima -</addr-line>
          <country country="PE">Peru</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>91</fpage>
      <lpage>100</lpage>
      <abstract>
        <p>Nowadays, there are a mature set of tools and techniques for data analytic, which help Data Scientist to extract knowledge from raw heterogeneous data. Nonetheless, there is still a lack spatio-temporal historical dataset allowing to study everyday life phenomena, such as vehicular congestion, press influence, the effect of politicians comments on stock exchange markets, the relation between food prices evolution and temperatures or rainfall, social structure resilience against extreme climate events, among others. Unfortunately, there are few datasets combining different sources of urban data in order to carry out studies of phenomena occurring in cities (i.e., Urban Analytics). To solve this problem, we have implemented a Web crawler platform for gathering a different kind of available public datasets.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Providing citizens with free access to raw data is
one of the new global trends. These data,
generated on a daily basis, could come from
different sources such as governmental entities, NGOs,
companies or Public Administration Entities,
social networks, newspapers, congestion services,
etc. Therefore, the data format must be a standard
to make easier the access, use, generation of
information and sharing. Thus, it is crucial that
governments and private organizations, which have
valuable data in their systems, servers, and databases
make available these datasets for common benefits
but taking into account citizen’s privacy.
Unfortunately, Latin American countries and in
particular Peru lacks of historical data available
to citizens. In May 2017, the government
published Legislative Decree1 to create the National
Authority for Transparency and Access to Public
Information; and, strengthen the Regime of
Personal Data Protection. This first step would allow
not only greater transparency on the part of
certain institutions but also the possibility of the
citizens of becoming a partner and author of solutions
that could improve the life quality of our society.
Peru is beginning to generate an open data culture
and developing an open data portal at the national
level2. Nonetheless, some phenomena need
finegrained data. For instance, in a research paper
        <xref ref-type="bibr" rid="ref10">(Srivastava, 2017)</xref>
        , the author highlights the need
to collect segregated data of urban poor for
inclusive urban planning. The huge scarcity of
segregated data does not allow to make a
comprehensive understanding of their vulnerabilities.
Segregated data of urban poor are essential for inclusive
planning and to build sustainable cities.
Undoubtedly, there are many benefits of having segregated
data of urban poor in urban planning, not only for
inclusive planning but also to understand the
vulnerability, to know the contribution of urban poor
in urban economy and to prioritize actions.
      </p>
      <p>Our main contribution is presenting an
alternative to gathering daily basis generated data (from
public sources for storing, organizing and
sharing) to perform Urban Big Data Analytics under
a privacy aware framework in a developing
country such as Peru. The information collection has
followed a sanitization process, which assures the
identity safety of citizens and brands located in
Peru. The aim of this platform is to provide an
urban dataset for studying different phenomena in
urban environments such as urban planning,
on1Legislative Decree 1353: ”Decreto legislativo que crea la
autoridad nacional de transparencia y acceso a la informacio´n
pu´blica, fortalece el re´gimen de proteccio´n de datos
personales y la regulacio´n de la gestio´n de intereses”</p>
      <p>2Sistema Nacional de Informacio´n Ambiental: sinia.
minam.gob.pe/
line emergency detection, vulnerability, climate
change, resilience and even poverty.</p>
      <p>The present paper is organized as follows:
Section 2 describes the related works, Sections 3, 4
and 5 detail the framework architecture, the
collected data and data statistics of some datasets,
respectively. Then, Section 6 shows an application
of Urban Analytics. Finally, Section 7 concludes
our work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>Open Data promotes innovation thoughts
societal participation with the use of the data. Such
datasets include measurement data from city-wide
sensor networks on smart cities as well as from
citizen sensors. In the current section, we present
some efforts for data collection to tackle urban
planning problems, to detect emergencies and to
show city insights in real time.</p>
      <p>
        Concerning data collection for urban planning,
        <xref ref-type="bibr" rid="ref8">(Rathore et al., 2016)</xref>
        propose a smart city data
collection platform. This platform gathers
information about floods, water usage, traffic,
vehicular mobility traces, parking lots, pollution, social
networks and weather from smart homes, smart
parking, vehicular networking, water &amp; weather
and environmental pollution monitoring systems.
The authors use the collected information for
urban planning decision-making. Nevertheless, in
        <xref ref-type="bibr" rid="ref9">(Santos et al., 2017)</xref>
        , authors claim a need to give
some context to this kind of measurements. Thus,
they proposed the Human-Aware Sensor Network
Ontology for Smart Cities (HASNetO-SC) to
describe knowledge associated with data collection
from city-wide sensor networks with an
appropriate level of contextual metadata for data
understanding. Therefore, they implemented the
architecture for data collection in an urban metropolitan
area in Brazil. Consequently, the platform opens
the possibility that citizens, who have a little to no
knowledge about the collection environment and
the collected data, to access and process the
information.
      </p>
      <p>
        About emergencies detection, the work of
        <xref ref-type="bibr" rid="ref11">(Xu
et al., 2016)</xref>
        proposes a mechanism to gather
information from Weibo3 about urban emergency
events. The platform discovers What, Where,
When, Who, and Why of a given emergency event
from Weibo comments. Thus, to complete these
pieces of information, the platform relies on Social
3Weibo website: tw.weibo.com
      </p>
      <p>Sensors and Crowdsourcing layers. The former
receives textual data from Weibo users. The latter
extracts the basic elements of an emergency event
(what, when, where, who, and why) to provide
information for rescue services or decision making.</p>
      <p>
        Regarding urban data gathering, which is an
essential element of modern cities, a great challenge
appears, such as data volume, velocity, data
quality, privacy, and security, among others. In the
paper
        <xref ref-type="bibr" rid="ref7">(Panagiotou et al., 2016)</xref>
        , authors describe
the development of a set of techniques that aim at
effective and efficient urban data management in
real settings in Dublin city. The solutions were
integrated into a system that is currently used by
the city. The system can detect multiple types of
incidents, each one focusing on a different input
source. Hence, the solutions can identify events
by analyzing in real-time GPS trajectories, data
coming from sensors installed in junctions, or
textual information coming from social media.
Authors developed analysis modules to forecast load
so that they can manage efficiently the volume and
velocity. Besides, they combine information to
infer events from data anomalies. Noisy data and
erroneous measurements are also dealt. Moreover,
machine learning algorithms are used to identified
relevant tweets avoiding in that way low-quality
data. Another real-time data collection application
can be found in London where the information can
be viewed in a dashboard4 . More details on this
work can be found in the reference
        <xref ref-type="bibr" rid="ref5">(Gray et al.,
2016)</xref>
        , where its main idea is to understand the city
dynamics in a better way. The system gathers data
from different third party entities and open data
platforms given as CSV file, JSON object or an
HTML.
      </p>
      <p>
        At last, there is also a multi-source dataset of
urban life in Milano and Trentino
        <xref ref-type="bibr" rid="ref2">(Barlacchi et al.,
2015)</xref>
        . The authors put together different data sets,
such as spatial grid, social pulse,
telecommunications, precipitations, weather, electricity, news,
and social pulse. The scientists locate spatially
all the data in a grid, which allows comparing
datasets generated in various companies using
different standards. The idea behind this dataset is
to make a testbed for different solutions to
urban problems like energy consumption, mobility
planning, tourism and migrant flows, urban
structures and interactions, event detection, urban
well4London city dashboard: citydashboard.org/
london
being, etc.
      </p>
      <p>In the next section, we detail the framework
proposed in the present effort.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Data gathering framework</title>
      <p>In the present section, we describe the different
components of the architecture of the Web crawler.
Figure 1 depicts the different parts of the
architecture, where each part is responsible for a given task
as follows:
the particular structure settled for a given website.
It is worth noting that websites structures for
extracting data are stored in the database beforehand.
For implementing this part of the crawler, we
relied on the Scrappy library5.</p>
      <p>Extracted data is temporally stored in the
database to generate a Comma Separated Vector
CSV file at the end of a day. Therefore, we store
the datasets in a daily basis to build a historical
repository. Finally, the last part of the framework
is the Control mechanism that verifies the state of
the crawler, scraper and file generation to notify
by email if something goes wrong while gathering
data from the different Websites.</p>
      <p>In the next section, we detail the different
datasets collected by the platform.
Crawler: this artifact is responsible for reading
the Uniform Resource Locators (URLs) from
the database to download the target web
pages from different websites.</p>
      <p>Data base: this data base engine stores a list of
URL provided by the user and the rules for
reading and extracting relevant fields.</p>
      <p>Web Scraper: it receives the downloaded web
pages and the extracting rules to parse the
web pages for gathering the needed fields of
a given web page.</p>
      <p>Storing: generate the Comma Separated Vector
(CSV) files for each treated web site.</p>
      <p>Control: verifies the aforementioned parts (i.e.,
Web Scraper, Crawler and Storing) are alive
and are able to perform their task without
problems.</p>
      <p>The process begins with the manual creation of
a list of target URLs and the rules needed to extract
relevant fields from these. Then, those information
are stored in the Data Base. The target URLs are
chosen by the user. The data extraction starts with
the Crawler reading the (URLs) from the Data
Base to download target Web pages. Then, the
Web Scrapper receives a set of Web pages as
input. Consequently, it generates an index I of all
the Web pages in each registered Website. Next,
the Web Scrapper extract the data from the
different gathered web pages, listed in the index I, using</p>
      <p>
        In this section, we detail the variables of the
collected datasets, which are available6. It is
important to note that a sanitization process was
performed over the datasets. Therefore, we
consider as sensible information people’s and brands’
names, which are pseudonymized and erased,
respectively. Concerning the opinion dataset, the
comment field is sanitized by erasing stop words
and sorting words alphabetivally to reduce the
impact of a De-anonymization re-link attack
        <xref ref-type="bibr" rid="ref4">(Gambs
et al., 2014)</xref>
        . Consequently, these sanitization
processes are carried out to prevent a privacy breach.
Please note that the sanitization process is
performed off-line. Thus, the sanitization does not
limit the extraction of useful information.
      </p>
      <p>It is possible to extract unstructured datasets
from the websites targeted (i.e. news from
newspapers, satellite imagery, etc.). In this work, we
present two non-tabular datasets: newspapers and
social networks. Nonetheless, a tabular structure
has been applied to them for easier readability.</p>
      <p>In the following paragraphs, we described each
of the nine categories of data sources as well as the
datasets in each category.</p>
      <p>Beauty consists of a description, date, price,
category and cosmetics products (c.f. Table 1).
Climate category contains data from monitoring
stations, atmospheric pollutants and radiation
5Scrapy: scrapy.org
6BITMAP Urban Analytics:
urbands.html
bitmap.com.pe/
category
mujer
mujer
Markets category reports maximum and
minimum prices, description of different products
of the first necessity of three different
markets of supermarkets and suppliers in Lima.</p>
      <p>Medicament category comprises prices of
medicines provided by the Ministry of
Health of Peru (c.f., MINSA). Table 6 shows
the registered drug dataset containing the
condition of the drug, address, technical
director, pharmacy name, price, name, country,
and date of manufacture.
Newspapers category contains news from
different print media in Peru. Table 7 describes
newspapers dataset composed of the
publication date, author, and the section of
newspaper.</p>
      <p>content
long text
section
mundo
date author
2014-02-14 journalist
13:33:47 name
location title
mundo/eeuu Edward</p>
      <p>Snowden
Real Estate Market comprehends prices for
houses, apartments, and offices sales and
rental nationwide. Table 8 shows real estate
data grouped in columns detailing whether
the state of the property is a sale or rental,
describes the property, and its address.
Additionally, the longitude and latitude of
the property are provided.
Social Networks category contains
georeferenced comments from people on
different topics and social relations.
Stock Market category contains two datasets for
money exchange rates and stock exchange
markets. The former contains the different
historical exchange rates, from Soles to other
foreign exchange as shown in Table 11.
dates, amount, the amount of the stock,
operations, and price variation. These variables
are detailed in Table 12.</p>
      <p>pre
7.8
currency</p>
      <p>S/
sector
IND
open
7.7
amountNeg
1 325 647
segm
RV1
sale
7.56
mnemonic
ALICORC1
last
7.56
company
Alicorp
noShare
173 677
var
-3.08</p>
      <p>date
2016-08-04
noOper
16
sale
7.7
Transportation contains main avenues traffic
jams and domestic as well as international
departure and arrivals flights at Jorge Chavez
airport in Lima, Peru.</p>
      <p>Tables 13 and 14 detail datasets of both alerts
and congestion, respectively. These datasets
contain information of some points in Lima
city. It details the street, city, date, latitude
and longitude coordinates of the alerts or the
level of congestion. Table 13 also indicates
the level of traffic, the node where the traffic
level is reset as well as the speed of the traffic.</p>
      <p>street
Av. Los Frutales</p>
      <p>latitude
-12.071628</p>
      <p>node
Calatrava</p>
      <p>city
La Molina
longitude
-76.964632
speed
4.719</p>
      <p>date
2016-08-23
traffic level
2.0</p>
      <p>In the case of Table 14, the types and
subtypes of alerts are gathered in addition to the
above-mentioned data.</p>
    </sec>
    <sec id="sec-4">
      <title>5 Dataset Statistics</title>
      <p>In this section, we detail the statistics of the
different dataset in the described categories in Section</p>
      <p>Beauty.- Table 18 shows the most interesting
characteristics of the beauty dataset,
described in Table 1. On the one hand, two of
the four attributes of the table contains
categorical values (discrete values) for which the
ber of attributes and the number of records per
data set. On the other hand, Table 17
synthesizes the characteristics linked to data types,
headers, and temporal space granularity. Concerning
temporal granularity range from 1.02 to 6.84
minutes. With regard to spatial granularity, there are
seven datasets georeferenced with UTM
coordinates (i.e., latitude and longitude).
values (c.f., NAs).</p>
      <p>N variable
0 category
1 date
2 price
3 article
N mode
0 arrugas
1
2 0
3 Essential cutis graso
type
str
date
float
str
min
0
mean
935
max
1400000
median
45
NAs
0
0
0
0
std
35152</p>
      <p>%NAs
0
0
0
0</p>
      <p>Climate.- We describe the characteristics
associated with climate data. Two datasets will be
described: 1) data from meteorological
stations and their measurements (Table 19); and,
2) data on pollutants by districts (Tables 20).
Other datasets are not described due to lack
of space.</p>
      <p>In Table 19, we have a large number of
categorical attributes and represent the
characteristics of the monitoring stations. However,
each of these monitoring stations measures
two meteorological characteristics, which are
represented by numerical values. These two
attributes have characteristics such as the
mean, median, standard deviation, mode, and
minimum and maximum values.</p>
      <p>In contrast, the pollutant data by districts (c.f.,
Table 20) contain a large amount of
numerical data. As already mentioned in the
previous descriptions, some of them are the mean,
median, standard deviation, mode, minimum
and maximum values.
Markets.- Here we describe the statistics of a
market and super market that sell groceries.
This category was described in Table 5.</p>
      <p>Table 21 shows statistics about product.
In the variable title, the most
popular product is potato. The most
mentioned type is red-headed onion. The
minimum price has an average value of
2, 427P EN and varies between 0.57P EN
and 13.00P EN . The maximum price is
on average 3.057P EN and varies between
0.71P EN and 14.00P EN . Finally, the
average price is 2.746P EN and fluctuates
between 0.61P EN and 13.5P EN .</p>
      <p>Newspapers.- Table 22 shows the most frequent
content that is the text1, which is a given news
(we prefer not to share the content due to the
lack of space). Then, the most cited author,
section, and title are Carlos Battle, executive
zone, world / Current and 5 tips for a startup
to survive, respectively. It should be noted
that the percentage of missing values of
content, author, and location are 81%, 82% and
95%.</p>
      <p>Real Estate.- This dataset contains three different
datasets. We only show one table due to lack
of space. Described data are rents and/or real
estate sales, the data are mostly categorical.
In Table 23 we have only two attributes with
numerical values associated with the
location of the property. These data have
statistics such as mean, median, standard
deviation, mode, minimum and maximum
values. On the other hand, we have another
group of characteristics with categorical
values. Among them, the more frequent section
variable
content
date
author
section
location
title
mode
texto1
Carlos
Batalla
zonaejecutiva
mun./act.
5
para
cons.</p>
      <p>que
una startup
sobre.
available.</p>
      <p>N variable
0 title
1 section
2 description
3 location
4 area
5 price
6 longitude
7 latitude
8 date
N mode
0 Alquiler de...
1 alquiler
2 Rento...
3 Ubicacin...
4 100 m
5 US$ 900
6 -77.03
7 -77.03
8
type
str
date
str
str
str
str
min
type
str
str
str
str
str
str
float
float
date
min
-77.76
-77.76
max
mean
-77.01
-77.01
max
-76.13
-76.13</p>
      <p>mean
median
0
0</p>
      <p>NAs
1264992
1275035
1480648
std
%NAs
erties of 100m2 are the most ”offered”, using
the attribute area. Regarding the price, the
most frequent value is $900. It is important
to note that, although we have a large amount
of null data (NAs), these represent a rather
low percentage due to a large amount of data
Transportation.- We describe the statistics of the
datasets related to the terrestrial
transportation mode, As far as the dataset of air
transport is concerned, we do not show them due
to lack of space. For land transport, Table 24
and 25 report the statistics on alerts and
congestion, respectively.</p>
      <p>N variable
0 rue
1 city
2 date
3 latitude
4 longitude
5 subtype
6 type
N mode
0 Av. Javier Prado Este
1 San Isidro
2
3 -12.091415
4 -77.003755
5 JAM HEAVY TRAFFIC
6 JAM
type
str
str
date
float
float
str
str
min
-12.1
-77.072
district with an average traffic level of three
on a scale of one to five. Finally, it shows the
average speed of 2.59 Km/h.
Tabla 26). We have the half of variables
categorical (currency and date) and the another
half numerical (purchase and sell). It is worth
noting that the mode of the currency attribute
is Swiss franc. Please note that there are some
null values (no data), which are represented
by the NAs and the percentage is given by
%NAs.</p>
      <p>In the present section, we have not described the
statistical metrics of two data categories, which are
medicaments and social networks because of lack
of space. In the next section, we propose an
example of Urban Analytics to show the potential of
our datasets.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Traffic congestion application</title>
      <p>To show a possible application of our datasets,
we take the Congestion (c.f., Table 13) and Alerts
(c.f., Table 14) datasets of the Transport category
to analyze traffic jams in a given district of Lima.
Consequently, we filter the congestion reports and
alerts of Lince district. Then, we select all records
produced in this district. Finally, we extract a CSV
file containing traffic data to analyze it.</p>
      <p>In the present example, we make a visual
analysis of congestion to show the enormous
potentiality of our dataset. Subsequently, we rely on Qlik7
to depict Figure 2.</p>
      <p>Figure 2 shows traffic level in red, minimal and
maximal speed in blue and yellow, respectively.
This analysis was done for five different avenues
and streets. As we can see, the lower the speed, the
higher the traffic level. Another interesting fact is
the reported maximal and minimal speed in Paseo
de la Republica Avenue, which has the same value
(i.e., 1.6Km) meaning a high congestion in this
avenue.</p>
      <p>7Qlik: qlikid.qlik.com</p>
      <p>Figure 4 shows the distribution of the 17700
alerts gathered with our platform. We note that
Jam heavy traffic and Jam stand still traffic are the
most reported alerts with 5500 alerts each one.</p>
      <p>Finally, using Fusion Tables8, we can draw a
geolocated heatmap of reported alerts and jams
in Figure 4A. As we can see, the most reported
and congested part is the road exchange between
Javier Prado Avenue and Paseo de la Repu´blica
highway at the bottom right the figure. We have
analyzed this segment of the city to study the
traffic level between this important road exchange and
the location of our university (i.e., Universidad del
Pac´ıfico) in the upper left part of the figure.</p>
      <p>A simple analysis is already interesting but
8Fusion tables
fusiontablestalks
sites.google.com/site/
crossing datasets could reveal useful insights.
Therefore, we use datasets from real estate and
business locations. The former is a dataset
described in Table 23. The latter is a private dataset
containing the code, latitude, and longitude of
businesses. Relying on these datasets, we
measured the influence of real estate and business
locations over congestion. Figure 5 a heat-map
(where light and strong blue mean short and large
distance, respectively) of the distance among the
datasets locations. It is possible to see that alerts
are influenced by real estate and business location
and jams, only by real estate locations.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>
        In the present work, we have described the
architecture of a Web crawler platform to gather
information about nine different categories of datasets
to make urban analytics. The main contribution
of this work is the provision of information to
the scientific community and policy makers for
analyzing and studying social behavior and
urban phenomena in a developing country such as
Peru. We have collected data following a
privacyaware structure. Sensible information about the
citizens or brand-names has been sanitized. These
datasets have enabled us to implement a
Knowledge Tier Platform for Graph Mining
        <xref ref-type="bibr" rid="ref6">(Nunezdel Prado et al., 2016)</xref>
        and perform urban
analytics
        <xref ref-type="bibr" rid="ref3">(Di Clemente et al., 2017)</xref>
        or study urban
resilience
        <xref ref-type="bibr" rid="ref1">(Abbar et al., 2016)</xref>
        .
      </p>
      <p>In the future, we plan to extend the crawler to
collect more information from new public
available Web sites. We also plan to make a privacy
risk analysis of the described datasets.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Sofiane</given-names>
            <surname>Abbar</surname>
          </string-name>
          , Tahar Zanouda, and Javier BorgeHolthoefer.
          <year>2016</year>
          .
          <article-title>Robustness and resilience of cities around the world</article-title>
          .
          <source>arXiv preprint arXiv:1608</source>
          .
          <fpage>01709</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Gianni</given-names>
            <surname>Barlacchi</surname>
          </string-name>
          , Marco De Nadai, Roberto Larcher, Antonio Casella, Cristiana Chitic, Giovanni Torrisi, Fabrizio Antonelli, Alessandro Vespignani, Alex Pentland, and
          <string-name>
            <given-names>Bruno</given-names>
            <surname>Lepri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A multi-source dataset of urban life in the city of milan and the province of trentino</article-title>
          .
          <source>Scientific data 2.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Riccardo</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Clemente</surname>
          </string-name>
          , Miguel Luengo-Oroz, Matias Travizano, Bapu Vaitla, and Marta C Gonzalez.
          <year>2017</year>
          .
          <article-title>Sequence of purchases in credit card data reveal life styles in urban populations</article-title>
          .
          <source>arXiv preprint arXiv:1703</source>
          .
          <fpage>00409</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>Se´bastien Gambs, Marc-Olivier Killijian, and Miguel Nu´n˜ez del Prado Cortez</article-title>
          .
          <year>2014</year>
          .
          <article-title>Deanonymization attack on geolocated data</article-title>
          .
          <source>Journal of Computer and System Sciences</source>
          <volume>80</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1597</fpage>
          -
          <lpage>1614</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Steven</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <surname>Oliver O'Brien</surname>
          </string-name>
          , and Stephan Hu¨gel.
          <year>2016</year>
          .
          <article-title>Collecting and visualizing real-time urban data through city dashboards</article-title>
          .
          <source>Built Environment</source>
          <volume>42</volume>
          (
          <issue>3</issue>
          ):
          <fpage>498</fpage>
          -
          <lpage>509</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Miguel</given-names>
            <surname>Nunez-del Prado</surname>
          </string-name>
          , Edgardo Bravo, Miguel Sierra, Miguel Canchay, and
          <string-name>
            <given-names>Isaias</given-names>
            <surname>Hoyos</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Knowledge tier platform for graph mining in (smart) cities</article-title>
          .
          <source>In Proceedings of Symposium on Information Management and Big Data.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Nikolaos</given-names>
            <surname>Panagiotou</surname>
          </string-name>
          , Nikolas Zygouras, Ioannis Katakis, Dimitrios Gunopulos, Nikos Zacheilas, Ioannis Boutsis, Vana Kalogeraki, Stephen Lynch, and Brendan OBrien.
          <year>2016</year>
          .
          <article-title>Intelligent urban data monitoring for smart cities</article-title>
          .
          <source>In Joint European Conference on Machine Learning and Knowledge Discovery in Databases</source>
          . Springer, pages
          <fpage>177</fpage>
          -
          <lpage>192</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>M Mazhar Rathore</surname>
            , Awais Ahmad, Anand Paul, and
            <given-names>Seungmin</given-names>
          </string-name>
          <string-name>
            <surname>Rho</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Urban planning and building smart cities based on the internet of things using big data analytics</article-title>
          .
          <source>Computer Networks</source>
          <volume>101</volume>
          :
          <fpage>63</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Henrique</given-names>
            <surname>Santos</surname>
          </string-name>
          , Vasco Furtado, Paulo Pinheiro, and
          <string-name>
            <surname>Deborah L McGuinness</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Contextual data collection for smart cities</article-title>
          .
          <source>arXiv preprint arXiv:1704</source>
          .
          <year>01802</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Ambey</given-names>
            <surname>Kumar Srivastava</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Segregated data of urban poor for inclusive urban planning in india: Needs and challenges</article-title>
          .
          <source>SAGE Open</source>
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>2158244016689377</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Zheng</given-names>
            <surname>Xu</surname>
          </string-name>
          , Yunhuai Liu, Neil Yen,
          <string-name>
            <surname>Lin</surname>
            <given-names>Mei</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiangfeng Luo</surname>
            , Xiao Wei, and
            <given-names>Chuanping</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Crowdsourcing based description of urban emergency events using social media big data</article-title>
          .
          <source>IEEE Transactions on Cloud Computing .</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>