<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Myths and Challenges in Knowledge Extraction and Big Data Analysis on Human-Generated Content from Web and Social Media Sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Politecnico di Milano</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dip. Elettronica</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Informazione e Bioingegneria (DEIB). Via Ponzio</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milano</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italy marco.brambilla@polimi.it</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Whatever people produce on digital media can be a relevant source of knowledge and behavioural analysis. This is the subject of interest of a wide part of the new discipline known as Web Science. However, special care must be exercised when setting up studies on this kind of sources. Indeed, these studies rarely satisfy the established scienti c method guidelines, because of the nature and size of the data, as well as because of the bias and scarce generalizability of results. This paper identi es some of the most crucial challenges that need to be addressed when tackling knowledge extraction and data analysis out of observational studies on human-generated content.</p>
      </abstract>
      <kwd-group>
        <kwd>Social Media</kwd>
        <kwd>Big Data</kwd>
        <kwd>Data Analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The Web and the social media are today the environments where people post
their content, opinions, activities and resources. Therefore, a considerable amount
of user-generated content is generated every day for a wide variety of purposes
and related to diverse topics and contexts, with high frequency and speed [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Exploring and studying this content o ers a great opportunity to understand
the ever evolving modern society, in terms of topics of interest, events, and
people relations and behaviour. Several approaches based on the so called Big Data
paradigm have analyzed social media [
        <xref ref-type="bibr" rid="ref1 ref19 ref2 ref23 ref7">7, 23, 1, 19, 2</xref>
        ], mobile phone usage data
[
        <xref ref-type="bibr" rid="ref11 ref15 ref24 ref9">11, 9, 15, 24</xref>
        ] and many other sources, with the purpose describe and/or predict
behaviour, density, issues, and topics or locations of interests [
        <xref ref-type="bibr" rid="ref13 ref16">16, 13</xref>
        ]. The
potential impact of the analyses can be even larger when information from multiple,
diverse sources is merged together [
        <xref ref-type="bibr" rid="ref12 ref14 ref18 ref20 ref5">12, 18, 20, 5, 14</xref>
        ]. This also include diversity
in the format of the data: greater value can be extracted when both textual and
quantitative information is analyzed, and even more so when also multimedia
content is considered (especially because of the more and more prominent role
that photos, video and audio is acquiring in human-to-human communication)
[
        <xref ref-type="bibr" rid="ref17 ref22 ref25 ref3 ref4">22, 25, 17, 4, 3</xref>
        ].
      </p>
      <p>While the data available on this sources is typically created by individual
entities, i.e. people or businesses, and pertain their own personal or social sphere,
analyses are able to study individuals and also to capture the integrated,
highlevel view of a more complex ecosystem or phenomenon, such as a city, a country,
a community, or an event or social trend.</p>
      <p>
        Therefore, we can say that the aim is to use appropriate tools for
understanding complex societal systems and dynamics. Such a tool was conceptualized back
in 1976 by De Rosnay [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] who proposed the term macroscope as a tool for
capturing complex systems. Time is now mature to put this tool at work, by using big
data technology and statistical analysis over the huge mass of human-produced
content.
      </p>
      <p>In the remaining part of the paper I will present the ideal target of this kind
of studies and some risks that are implied by the context and the data collected.
2</p>
      <p>The Myth: Comprehensive, Coherent, and Cohesive
Analyses
The Big Data paradigm is pushing toward a scenario where people perceive
the broad availability of large-scale storage, massive computational power, and
exible and easy to use statistical and machine learning tools. This re ects in
the expectation (both in individuals and businesses) to become able to obtain
always and quickly any kind of insights they may need out of the existing
data.</p>
      <p>This implies that people expect that the executed analyses will always
produce relevant insights for their needs. In particular, results are expected to be:
{ comprehensive, i.e., always providing a clear picture of the reality with all
the relevant aspects, at all levels of details, from very ne-grained results to
resilient and solid integrated results;
{ coherent, i.e., self-consistent in all their versions and dimensions, no matter
what is the data source, the granularity of the analysis and the time/space
considered;
{ cohesive, i.e., being able to convey a united and insightful understanding
of reality, which can lead to informed and appropriate decision-making
processes.</p>
      <p>Is this the case? Or, better, is this an automatic outcome that we shall expect
from any kind of analysis? Unfortunately, it's quite the opposite, as discussed
in the next section.
3</p>
      <p>The Challenges of Knowledge Extraction, Web Science,
and Content Analytics
In the majority of the analyses, if work is not conducted properly, risks and biases
are not considered, and statistics is not applied correctly, results will su er
of so many critical aspects that they will be hardly instrumental for
any objective. This section has the purpose to introduce some awareness of
the risk of big data and knowledge extraction. I don't aim at completeness, but
just to report some of the most common challenges that a data scientist need
to face when trying to extract knowledge and aggregated analytics out of
usergenerated content. These challenges include the well known 4 V s of Big Data,
i.e., Volume, Velocity, Variety, and Veracity, and some more. I report them in
details below, together with some references to practical experiences.
3.1</p>
      <sec id="sec-1-1">
        <title>Complexity of knowledge</title>
        <p>
          Independently of the technique used for collecting and elaborating the data, and
actually also independently of the data itself, we need to keep into account that
reality is complex and varies in time, space and along many other dimensions,
including societal and economic variables. Aiming at capturing this complexity
in its entirety is a goal too ambitious for a single analysis. On the other side,
single aspects can be tackled, such as studying the continuously evolving nature
of knowledge [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Indeed, knowledge in the world continuously evolves, and
existing knowledge bases are largely incomplete especially with respect to
lowfrequency data, belonging to the so-called long tail. Appropriate means can
be applied to exploiting content generated on social network platforms, which
are excellent sources for discovering emerging knowledge, as they immediately
re ects the evolving information and emerging concepts.
3.2
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Complexity of data (aka., variety or heterogeneity)</title>
        <p>On a more practical level, the data sources that are investigated are
complex themselves, as they feature unstructured, multi-format, heterogeneous data.
Tackling this variety is challenging per se, as each medium and type of content
may need di erent collection, cleaning and analysis techniques. Furthermore,
connecting together the extracted insights is challenging too. Finally, an
additional complexity aspect is related to the structure of the data, which can feature
record formats with data organized in deep hierarchies and relations.</p>
        <p>In general, complexity may lead to harder understandability of the data.
3.3</p>
      </sec>
      <sec id="sec-1-3">
        <title>Cognitive bias (i.e., of the observer)</title>
        <p>As soon as the data is collected and analysed, an implicit level of bias is
introduced. This is the bias of the observer, who is always in uencing the way the
data is perceived, studied and processed, based on his preliminary knowledge,
context, or visibility over the reality. This is also known as the street lamp e ect :
if you look at a completely dark town, with only one lamp on, as an observer
you tend to bias your understanding based only on what you see in the lit area.
3.4</p>
      </sec>
      <sec id="sec-1-4">
        <title>Sampling bias (i.e., of the content)</title>
        <p>The collected data is usually also biased with respect to the question that the
analysis aims to respond. Established scienti c method de nes the way in which
controlled experiments should be run and conclusions should be drawn. One
crucial aspect that should be taken care of is to always apply an appropriate
sampling method, granting that the sample is not biased with respect to the
population the experiment want to assess. However, this is not so trivial to
obtain. For instance, 1 shows two heatmaps representing the density of
geolocated photos taken in Milano and shared on two di erent web sharing platforms,
namely Instagram and Flickr: they both feature heavy bias with respect to the
total population of the city, due to the mean used for capturing and sharing
the photos (mobile app and web site respectively), the audience of the di erent
platforms (casual smartphone users and photographers respectively): the two
representations of the city are dramatically di erent, but there is no \correct"
or \wrong" one: it depends on the kind of questions one wants to answer and on
the kind of population one is interested in.</p>
        <p>Furthermore, in data analysis and extraction from web sources, in many
cases controlled experiments may not be implemented. Data scientists frequently
need to resort to observational studies, which try to draw conclusions out of
observations of the uncontrolled reality of facts and happenings.
3.5</p>
      </sec>
      <sec id="sec-1-5">
        <title>Data Availability</title>
        <p>A preliminary problem that many analyses may face is the di culty in nding
and acquiring the data. In many cases, integrated data analysis requires
availability of information that is now owned by the analyser himself. Therefore, the
analyser need to rely on third parties to release and actually deliver the data.
Such third parties may opt for releasing only part of the data, or only on a
limited amount per time basis (throttling), or they can simply deliver limited or
aggregated views on the data. Some strategies purposefully apply sampling or
data limitations in a non-transparent way, so that the data consumer is not able
to extract generalizable knowledge out of the analysis work.
3.6</p>
      </sec>
      <sec id="sec-1-6">
        <title>Data Quality</title>
        <p>Even when it is actually possible to collect the desired data, a further problem
is about the quality of such data. An entire discipline is focused on capturing
and solving data quality problems, which can span cases like dummy values,
missing values, cryptic data, contradicting data, violation of business rules, data
duplication, non-unique identi ers, and many more.
3.7</p>
      </sec>
      <sec id="sec-1-7">
        <title>Data Granularity</title>
        <p>Granularity of data refers to the fact that the collected information may refer
to di erent units of analysis on the time scale, space scale, and also along other
dimensions. This is a critical issue, especially in the case of analysis on integrated
data sources, where each source is based on its own granularity. In this case,
techniques based on discretization and uniformation must be applied.</p>
        <p>
          For instance, along the geographic positioning dimension di erent granularity
may mean that some data sources expose data about speci c venues/locations,
while others expose data about a geographical area de ned based on
administrative or political borders (e.g., cities, states, provinces, or countries), or on
xed geometrical shapes (squares or Voronoi tassellation); in this case, the
solution is usually to rely on a common grid structure. Analogously, along the time
dimensions some datasets may provide punctual information, while others may
describe di erent time periods, and thus the solution is to align the data on
common periods [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
3.8
        </p>
      </sec>
      <sec id="sec-1-8">
        <title>Data Volume</title>
        <p>One (possibly obvious) concern regarding data is its volume: in web sources,
the amount of data tend to grow enormously. This implies that any analysis
needs to cope with the bandwidth, storage, and computation needs of big data.
On the other hand, the challenge sometimes is just the opposite: the amount of
collectable data is just not su cient to run sensible analyses.
3.9</p>
      </sec>
      <sec id="sec-1-9">
        <title>Data Coherency</title>
        <p>One of the biggest issues with dealing with multiple data sources is that
typically the collected data is not coherent: this implies that data cleaning and
comparisons procedures must be run, to assess correlation and coherency of the
sources.
3.10</p>
      </sec>
      <sec id="sec-1-10">
        <title>Data Velocity (i.e., of the content)</title>
        <p>The speed with which the content is produced may also be a concern, as it
may become challenging to keep the pace with it in terms of data ingestion,
storage and analysis speed. This problem is typically addressed by adopting
data streaming solutions when possible.
3.11</p>
      </sec>
      <sec id="sec-1-11">
        <title>Result Velocity (i.e., of the output)</title>
        <p>Finally, the speed with which the output is produced is also relevant. Practical
experiences show that in most analysis scenarios a strict real-time publishing is
not requested. Careful curation of o -line, post-hoc analyses of the phenomenon
at large is preferred instead.
4</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Conclusion</title>
      <p>In this paper I summarized the main expectations and actual challenges in
modern data science approaches applied to web and user-generated content. The
fact that only observational studies can be applied on web and social network
contents is a further complexity factor in these studies. However, if proper data
collection, aggregation, cleansing (wrangling) and analysis techniques are
applied, these studies can extremely valuable results. Further examples, real cases
and discussions can be found in the presentation of this speech, available online.1
1 https://marco-brambilla.com/2017/09/22/myths-and-challenges-in-knowledge
-extraction-and-big-data-analysis/</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouhorma</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmed</surname>
            ,
            <given-names>M.B.</given-names>
          </string-name>
          :
          <article-title>Smart citizen sensing: A proposed computational system with visual sentiment analysis and big data architecture</article-title>
          .
          <source>International Journal of Computer Applications</source>
          <volume>152</volume>
          (
          <issue>6</issue>
          ) (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Arnaboldi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brambilla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cassottana</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciuccarelli</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vantini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Urbanscope: A lens to observe language mix in cities</article-title>
          .
          <source>American Behavioral Scientist</source>
          <volume>61</volume>
          (
          <issue>7</issue>
          ),
          <volume>774</volume>
          {
          <fpage>793</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bakhshi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shamma</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gilbert</surname>
          </string-name>
          , E.:
          <article-title>Faces engage us: Photos with faces attract more likes and comments on instagram</article-title>
          .
          <source>In: Proceedings of the 32Nd Annual ACM Conference on Human Factors in Computing Systems</source>
          . pp.
          <volume>965</volume>
          {
          <fpage>974</fpage>
          . CHI '14,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2014</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/2556288.2557403
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bakhshi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shamma</surname>
            ,
            <given-names>D.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kennedy</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gilbert</surname>
          </string-name>
          , E.:
          <article-title>Why we lter our photos and how it impacts engagement</article-title>
          .
          <source>In: ICWSM</source>
          . pp.
          <volume>12</volume>
          {
          <issue>21</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Balduini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Della</given-names>
            <surname>Valle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Azzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Larcher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Antonelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Ciuccarelli</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          : Citysensing:
          <article-title>Fusing city data for visual storytelling</article-title>
          .
          <source>IEEE MultiMedia 22(3)</source>
          ,
          <volume>44</volume>
          {
          <fpage>53</fpage>
          (
          <year>July 2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Balduini</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Della</given-names>
            <surname>Valle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>DellAglio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Tsytsarau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Palpanas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Confalonieri</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          :
          <article-title>Social listening of city scale events using the streaming linked data framework</article-title>
          .
          <source>In: International Semantic Web Conference</source>
          . pp.
          <volume>1</volume>
          {
          <fpage>16</fpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naaman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gravano</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Identifying content for planned events across social media sites</article-title>
          .
          <source>In: Proceedings of the Fifth ACM International Conference on Web Search and Data Mining</source>
          . pp.
          <volume>533</volume>
          {
          <fpage>542</fpage>
          . WSDM '12,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2012</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/2124295.2124360
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naaman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gravano</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Learning similarity metrics for event identi - cation in social media</article-title>
          .
          <source>In: Proceedings of the Third ACM International Conference on Web Search and Data Mining</source>
          . pp.
          <volume>291</volume>
          {
          <fpage>300</fpage>
          . WSDM '10,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2010</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1718487.1718524
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Caceres</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loh</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urbanek</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varshavsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volinsky</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A tale of one city: Using cellular network data for urban planning</article-title>
          .
          <source>IEEE Pervasive Computing</source>
          <volume>10</volume>
          (
          <issue>4</issue>
          ),
          <volume>18</volume>
          {
          <fpage>26</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Brambilla</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ceri</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valle</surname>
            ,
            <given-names>E.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Volonterio</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salazar</surname>
            ,
            <given-names>F.X.A.</given-names>
          </string-name>
          :
          <article-title>Extracting emerging knowledge from social media</article-title>
          .
          <source>In: 26th International Conference on World Wide Web, WWW</source>
          <year>2017</year>
          , Perth, Australia, April 3-
          <issue>7</issue>
          ,
          <year>2017</year>
          . pp.
          <volume>795</volume>
          {
          <issue>804</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Caceres</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rowl</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Urbanek</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Exploring the use of urban greenspace through cellular network activity (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Calabrese</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colonna</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lovisolo</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parata</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ratti</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Real-time urban monitoring using cell phones: A case study in rome</article-title>
          .
          <source>IEEE Transactions on Intelligent Transportation Systems</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ),
          <volume>141</volume>
          {
          <fpage>151</fpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Candia</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schoenharl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madey</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Uncovering individual and collective human dynamics from mobile phone records</article-title>
          .
          <source>Journal of physics A: mathematical and theoretical</source>
          <volume>41</volume>
          (
          <issue>22</issue>
          ),
          <volume>224015</volume>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myers</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leskovec</surname>
          </string-name>
          , J.:
          <article-title>Friendship and mobility: User movement in location-based social networks</article-title>
          .
          <source>In: Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>1082</volume>
          {
          <fpage>1090</fpage>
          . KDD '11,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>De Nadai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staiano</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Larcher</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sebe</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quercia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lepri</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>The death and life of great italian cities: a mobile phone data perspective</article-title>
          .
          <source>In: Proceedings of the 25th International Conference on World Wide Web</source>
          . pp.
          <volume>413</volume>
          {
          <fpage>423</fpage>
          .
          <string-name>
            <surname>International World Wide Web Conferences Steering Committee</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hidalgo</surname>
            ,
            <given-names>C.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Understanding individual human mobility patterns</article-title>
          .
          <source>arXiv preprint arXiv:0806.1256</source>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. Jaakonmaki, R., Muller,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Brocke</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.v.</surname>
          </string-name>
          :
          <article-title>The impact of content, context, and creator on user engagement in social media marketing</article-title>
          .
          <source>In: 50th Hawaii International Conference on System Sciences, HICSS</source>
          <year>2017</year>
          ,
          <article-title>Hilton Waikoloa Village</article-title>
          , Hawaii, USA, January 4-
          <issue>7</issue>
          ,
          <year>2017</year>
          (
          <year>2017</year>
          ), http://aisel.aisnet.org/hicss-50/da/data_ text_web_mining/6
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Krings</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calabrese</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ratti</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blondel</surname>
          </string-name>
          , V.D.:
          <article-title>Urban gravity: a model for inter-city telecommunication ows</article-title>
          .
          <source>Journal of Statistical Mechanics: Theory and Experiment</source>
          <year>2009</year>
          (
          <volume>07</volume>
          ),
          <source>L07003</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Psyllidis</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bozzon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bocconi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolivar</surname>
          </string-name>
          , C.T.:
          <article-title>A platform for urban analytics and semantic data integration in city planning</article-title>
          .
          <source>In: International Conference on Computer-Aided Architectural Design Futures</source>
          . pp.
          <volume>21</volume>
          {
          <fpage>36</fpage>
          . Springer (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Quercia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lathia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Calabrese</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di Lorenzo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crowcroft</surname>
          </string-name>
          , J.:
          <article-title>Recommending social events from mobile phone location data</article-title>
          .
          <source>In: Data Mining (ICDM)</source>
          ,
          <year>2010</year>
          IEEE 10th International Conference on. pp.
          <volume>971</volume>
          {
          <fpage>976</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21. de Rosnay, J.:
          <article-title>Le macroscope: vers une version globale</article-title>
          .
          <source>Editions du Seuil</source>
          (
          <year>1975</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Sabate</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Berbegal-Mirabent</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Can~abate,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Lebherz</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.R.:</surname>
          </string-name>
          <article-title>Factors in uencing popularity of branded content in facebook fan pages</article-title>
          .
          <source>European Management Journal</source>
          <volume>32</volume>
          (
          <issue>6</issue>
          ),
          <volume>1001</volume>
          {
          <fpage>1011</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jain</surname>
          </string-name>
          , R.:
          <article-title>Social pixels: genesis and evaluation</article-title>
          .
          <source>In: Proceedings of the 18th ACM international conference on Multimedia</source>
          . pp.
          <volume>481</volume>
          {
          <fpage>490</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Wesolowski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eagle</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noor</surname>
            ,
            <given-names>A.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snow</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buckee</surname>
            ,
            <given-names>C.O.:</given-names>
          </string-name>
          <article-title>The impact of biases in mobile phone ownership on estimates of human mobility</article-title>
          .
          <source>Journal of the Royal Society Interface</source>
          <volume>10</volume>
          (
          <issue>81</issue>
          ),
          <volume>20120986</volume>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Yuheng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lydia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbarao</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>What we instagram: A rst analysis of instagram photo content and user types</article-title>
          , pp.
          <volume>595</volume>
          {
          <fpage>598</fpage>
          . The AAAI Press (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>