<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information Evolution Modeling and Tracking: State-of-Art, Challenges and Opportunities</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ekaterina Shabunina</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabriella Pasi</string-name>
          <email>pasig@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universita degli Studi di Milano-Bicocca, Dipartimento di Informatica Sistemistica e Comunicazione</institution>
          ,
          <addr-line>Viale Sarca 336, 20126 Milan</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the Web 2.0, where everyone is the creator of content, information spreads and evolves rapidly through unpredictable paths of rebounds between news sources and Social Media. In this context, modeling, analyzing and tracking the information evolution through time o ers unprecedented opportunities to diverse research elds, including Information Retrieval. In this paper we propose a synthetic analysis of the state-of-art on Information Evolution on the Web, and we summarize the interesting opportunities it o ers to Information Retrieval.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The emergence of the Web 2.0 has granted user the freedom to interact with
other users and to contribute contents to the World Wide Web. Consequently, it
has motivated novel research directions, and it has also provided new
perspectives within existing ones. The most common way to propagate opinions and
ideas on the Web is constituted by Social Media, which facilitates the creation
and sharing of User Generated Content (UGC). Thus, Social Media provides the
possibility to analyze the content generated by a vast number of users from
different countries and social backgrounds to the aim of excerpting cultural trends
and ideas that spread in real time. This enables unprecedented opportunities
such as, for example, an early detection of social crisis, disasters and
emergencies [
        <xref ref-type="bibr" rid="ref10 ref2">2,10</xref>
        ]. The identi cation of social phenomena in UGC can bring insight on
the behavior of users in Social Networks, the patterns of their interactions, and
the structure of the information spread depending on the phenomenon driving
it. Thus, the study of the evolution of information on Social Media allows to
track how information related to speci c topics or events changes over time, and
it makes possible to monitor the evolution of cultural, political and social ideas.
Evidently, modeling, analyzing and tracking the evolution of information in time
on the Web and, particularly, on Social Media promises a myriad of
unprecedented opportunities and applications to several elds, among which Information
Retrieval and Social Media Analytics.
      </p>
      <p>In the following Sections we will present a synthetic analysis of the
state-ofart in modeling, analyzing and tracking the evolution of information on the Web
with respect to its main challenges (Section 2) and the opportunities it o ers to
Information Retrieval and related elds (Section 3).</p>
    </sec>
    <sec id="sec-2">
      <title>State-of-Art and Challenges</title>
      <p>
        Information Propagation (IP) aims to analyze the spread of information on the
Web through time. This issue has been explored by a number of works in the
literature. The majority of the proposed approaches has considered IP as a
network-centered problem [
        <xref ref-type="bibr" rid="ref13 ref6">6,13</xref>
        ] The main focus of this line of research is to
study how information spreads in a user network. Simultaneously, another line
of works on IP aims to study how the content of a piece of information evolves
in time [
        <xref ref-type="bibr" rid="ref3 ref5 ref8">3,5,8</xref>
        ]. The scope of the present paper is this latter approach, the
objective of which is the quantitative and qualitative evaluation of the evolution of a
stream of information.
      </p>
      <p>
        In content-centered IP the common approach is to primarily identify the
core units of information in a stream of Social Media posts. In the
literature these units of information have been frequently referred to as \memes"
[
        <xref ref-type="bibr" rid="ref1 ref11 ref12 ref3 ref4 ref8 ref9">1,3,4,8,9,11,12</xref>
        ], a notion coined by R. Dawkins in 1976 to refer to a unit of
human cultural evolution, analogous to a gene in genetics [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Subsequently, the
analysis of information evolution is performed on the identi ed core units of
information of the studied information stream.
      </p>
      <p>
        There are two main challenges in content-centered IP. The rst core aspect
is the formal representation of textual information. In the majority of works in
the literature a unit of information is identi ed as a short, frequently quoted
phrase and its slight variations [
        <xref ref-type="bibr" rid="ref1 ref11 ref12 ref8">1,8,11,12</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref12 ref8">8,12</xref>
        ] it is formally de ned as a
phrase graph, with nodes representing the phrases and the edges representing the
edit distance between the phrases. In [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] a unit of information is assumed to be
represented by several objects such as \hashtags" and \mentions" on Twitter,
URLs and the preprocessed text of the tweet itself. Similarly, in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] di erent
types of displays of memes from the Yahoo! Meme platform are considered:
short snippets of text, photos, audio, or video, tokens in URLs, etc, which are
represented as bag-of-words.
      </p>
      <p>
        The work in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is among the ones that pioneered the issue of the identi cation
and formal representation of an information granule in a Social Media stream;
such information granule is de ned by the authors as an ememe or \electronic
meme", and is formally represented as a micro ontology generated by posts on
the blogosphere by means of an OWL schema.
      </p>
      <p>
        The study in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] presents a semi-supervised attempt to represent memes in a
set of documents as semantic networks, by extracting n-grams that co-occur in
many documents and, subsequently, by constructing the semantic network where
nouns and adjectives, from the extracted n-grams, are the nodes and verbs are
the edges between them.
      </p>
      <p>
        The second challenge in the content-centered study of information di usion
concerns the methods for measuring, evaluating and analyzing the information
evolution in time. In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] a set of operators is proposed in the context of semantic
web ontologies aimed at measuring some useful properties of ememes such as
delity (the degree to which a meme is accurately reproduced, computed as the
fuzzy matching between the given blog post and the original ememe
description), mutation (the di erence between the maximum and the minimum delity
among the instances of the ememe), spread (reproductive activity of a meme,
calculated as the number of instances of the ememe in the searched source), and
longevity (the time duration of the ememe's life span, which is the di erence
between the dates of the most recent and the oldest posts that contain the ememe
instance). Similarly, in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] three meme metrics are proposed in the context of
semantic networks: longevity (alike longevity in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]), fecundity (alike spread in
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]) and copy- delity (alike delity in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]). One of the most recent and large-scale
studies on memes, presented in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], analyzes the mutation and replication rates
in memes evolution with the Yule process. The work in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] presents a study
on the changes introduced in quoted texts as they di use through time; the
authors examine properties of the quoted texts variants and uncover patterns in
the rate of appearance of new variants, their length, the types of changes
introduced, their popularity and the type of sites that are replicating them. The
temporal patterns of variations in quoted phrases are studied in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], by
extracting the temporal threads of all blogs and news media sites that mention the
meme phrase, identifying the patterns and time lags of quoting between them,
as well as analyzing their change in time in the whole thread.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Opportunities</title>
      <p>As previously outlined, the possibility to track through time the evolution of
information on the Web can bring large and unprecedented bene ts and
opportunities to many research areas such as Information Retrieval, Social Network
Analysis and others. Here we emphasize some promising directions for the
exploitation of content-centered IP in the context of IR.</p>
      <p>
        User pro ling for personalized search is commonly performed by tracking the
user's activities on the Web to infer a representation of the user's interests. More
recently, users' pro les have been de ned based on the content generated by users
in Social Media [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Commonly, user interests are dynamic and they evolve in
time. Thus, user pro ling presents a natural scenario for an application of the
automatic analysis of the evolution of an information stream. Independently
from the means by which the user's topical interests are gathered, either by
query logs or as the content generated by the user on Social Media, the methods
for tracking the evolution of information in time can be successfully applied to
update the formal representation of the user model.
      </p>
      <p>Additionally, the exploitation of the evolution in time identi ed in the users
topical interests, can help in dealing with the \ lter bubble" problem in
personalized search and personalized recommendation by introducing a diversi cation
in the retrieved results through the natural and non-evident change in the
information.</p>
      <p>Another interesting application of tracking the evolution of textual
information is the analysis of queries formulated by the user over a given time interval.
This could bring insights on how users' interests change in time.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Adamic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Lento</surname>
          </string-name>
          , E. Adar, and
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Ng</surname>
          </string-name>
          .
          <article-title>Information evolution in social networks</article-title>
          .
          <source>In Proceedings of the Ninth ACM International Conference on Web Search and Data Mining, WSDM '16</source>
          , pages
          <fpage>473</fpage>
          {
          <fpage>482</fpage>
          , New York, NY, USA,
          <year>2016</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Avvenuti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cresci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Marchetti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Meletti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tesconi</surname>
          </string-name>
          .
          <article-title>Ears (earthquake alert and report system): A real time decision support system for earthquake crisis management</article-title>
          .
          <source>In Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '14</source>
          , pages
          <fpage>1749</fpage>
          {
          <fpage>1758</fpage>
          , New York, NY, USA,
          <year>2014</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>H.</given-names>
            <surname>Beck-Fernandez</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. F.</given-names>
            <surname>Nettleton</surname>
          </string-name>
          .
          <article-title>Identi cation and extraction of memes represented as semantic networks from free text online forums</article-title>
          .
          <source>In MDAI 2013 - Modeling Decisions for Arti cial Intelligence</source>
          , Barcelona,
          <volume>20</volume>
          /11/
          <year>2013</year>
          2013.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>F.</given-names>
            <surname>Bonchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Castillo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Ienco</surname>
          </string-name>
          .
          <article-title>Meme ranking to maximize posts virality in microblogging platforms</article-title>
          .
          <source>Journal of Intelligent Information Systems</source>
          ,
          <volume>40</volume>
          (
          <issue>2</issue>
          ):
          <volume>211</volume>
          {
          <fpage>239</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>G.</given-names>
            <surname>Bordogna</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Pasi</surname>
          </string-name>
          .
          <article-title>A fuzzy approach to the conceptual identi cation of ememes on the blogosphere</article-title>
          .
          <source>In Fuzzy Systems (FUZZ)</source>
          ,
          <source>2013 IEEE International Conference on, pages 1{8</source>
          ,
          <string-name>
            <surname>July</surname>
          </string-name>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. J. Cheng, L. Adamic,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Dow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Kleinberg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          . Can cascades be predicted?
          <source>In Proceedings of the 23rd International Conference on World Wide Web, WWW '14</source>
          , pages
          <fpage>925</fpage>
          {
          <fpage>936</fpage>
          , New York, NY, USA,
          <year>2014</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>R.</given-names>
            <surname>Dawkins</surname>
          </string-name>
          .
          <source>The Sel sh Gene</source>
          . Oxford University Press, Oxford, UK,
          <year>1976</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Backstrom</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Kleinberg</surname>
          </string-name>
          .
          <article-title>Meme-tracking and the dynamics of the news cycle</article-title>
          .
          <source>In Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '09</source>
          , pages
          <fpage>497</fpage>
          {
          <fpage>506</fpage>
          , New York, NY, USA,
          <year>2009</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>J.</given-names>
            <surname>Ratkiewicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Conover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Meiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Goncalves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Patil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Flammini</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Menczer</surname>
          </string-name>
          .
          <article-title>Truthy: Mapping the spread of astroturf in microblog streams</article-title>
          .
          <source>In Proceedings of the 20th International Conference Companion on World Wide Web, WWW '11</source>
          , pages
          <fpage>249</fpage>
          {
          <fpage>252</fpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. T. Sakaki,
          <string-name>
            <given-names>M.</given-names>
            <surname>Okazaki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matsuo</surname>
          </string-name>
          .
          <article-title>Earthquake shakes twitter users: Realtime event detection by social sensors</article-title>
          .
          <source>In Proceedings of the 19th International Conference on World Wide Web, WWW '10</source>
          , pages
          <fpage>851</fpage>
          {
          <fpage>860</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>M. P. Simmons</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          <string-name>
            <surname>Adamic</surname>
            , and
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Adar</surname>
          </string-name>
          .
          <article-title>Memes online: Extracted, subtracted, injected, and recollected</article-title>
          . In L. A.
          <string-name>
            <surname>Adamic</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          <string-name>
            <surname>Baeza-Yates</surname>
          </string-name>
          , and S. Counts, editors,
          <source>ICWSM. The AAAI Press</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>C. Suen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Eksombatchai</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Sosic</surname>
            , and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Leskovec</surname>
          </string-name>
          .
          <article-title>Nifty: A system for large scale information ow tracking and clustering</article-title>
          .
          <source>In Proceedings of the 22Nd International Conference on World Wide Web, WWW '13</source>
          , pages
          <fpage>1237</fpage>
          {
          <fpage>1248</fpage>
          , New York, NY, USA,
          <year>2013</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          .
          <article-title>Modeling information di usion in implicit networks</article-title>
          .
          <source>In Proceedings of the 2010 IEEE International Conference on Data Mining, ICDM '10</source>
          , pages
          <fpage>599</fpage>
          {
          <fpage>608</fpage>
          , Washington, DC, USA,
          <year>2010</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>A.</given-names>
            <surname>Younus</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. O'Riordan</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Pasi</surname>
          </string-name>
          .
          <article-title>A language modeling approach to personalized search based on users' microblog behavior</article-title>
          .
          <source>In Advances in Information Retrieval - 36th European Conference on IR Research</source>
          , ECIR
          <year>2014</year>
          , Amsterdam, The Netherlands,
          <source>April 13-16</source>
          ,
          <year>2014</year>
          . Proceedings, pages
          <volume>727</volume>
          {
          <fpage>732</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>