<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Capturing users' information and communication needs for the press o cers</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giovanni Stilo</string-name>
          <email>stilo@di.uniroma1.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Morbidoni</string-name>
          <email>c.morbidoni@univpm.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Cucchiarelli</string-name>
          <email>a.cucchiarelli@univpm.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paola Velardi</string-name>
          <email>velardi@di.uniroma1.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Sapienza University of Rome</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universita Politecnica delle Marche</institution>
          ,
          <addr-line>Ancona</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>14</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>The way in which people acquire information on events and form their own opinion on them has changed dramatically with the advent of social media. For many readers, the news gathered from online sources becomes an opportunity to share points of view and information within the micro-blogging platforms such as Twitter, mainly aimed at satisfying their communication needs. Furthermore, the need to deepen the aspects related to the news stimulates a demand for information that is often met through online information sources, such as Wikipedia. This behaviour has also in uenced the way in which journalists write their articles, requiring a careful assessment of what actually interests readers. The goal of this article is to de ne a methodology for a recommender system able to suggest to the journalist, for a given event, the aspects still uncovered in news articles in which the readers' interest focuses. The basic idea is to characterize an event according to the echo it had in online news sources and associate it with the corresponding readers' communicative and informative patterns, detected through the analysis of Twitter and Wikipedia respectively. Our methodology temporally aligns the results of this analysis and identi es as recommendations the concepts that emerge as topic of interest from Twitter and Wikipedia, not covered in the published news articles.</p>
      </abstract>
      <kwd-group>
        <kwd>recommender system</kwd>
        <kwd>wikipedia</kwd>
        <kwd>twitter</kwd>
        <kwd>social networks</kwd>
        <kwd>media</kwd>
        <kwd>press agents</kwd>
        <kwd>events detection</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In a recent study on the use of social media sources by journalists [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] the
author concludes that "social media are changing the way news are gathered and
researched". In fact, a growing number of readers, viewers and listeners access
online media for their news [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. When readers fell involved by news stories they
may react by trying to deepen their knowledge on the subject, and/or confronting
their opinions with peers. Stories may then solicit a reader's information and
communication needs. The intensity and nature of both needs can be measured
on the web, by tracking the impact of news on users' search behaviour on
online knowledge bases, and their discussions on popular social platforms. What
is more, on-line public's reaction to news is almost immediate [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and even
anticipated, as for the case, e.g., of planned media events and performances,
or for disasters [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. A related issue is the so called fact-checking problem, i.e.
the need to validate the veracity of news. The growing relevance of this type of
information need is also demonstrated by the recent announcement by Facebook
of the Journalism Project to help ght fake news [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ].
      </p>
      <p>
        Assessing the focus, duration and outcomes of news stories on public
attention is paramount for both public bodies and media in order to determine the
issues around which the public opinion forms, and in framing the issues (i.e.,
how they are being considered) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Futhermore, real-time analysis of public
reaction to news may provide a useful feedback to journalists, such as highlighting
aspects of a story that need to be further addressed, issues that appear to be of
interest for the public but have been ignored, or even to help local newspapers
echoing international press releases.
      </p>
      <p>
        The aim of this paper is to present a methodology to e ectively exploit social
data sources for the purpose of news media recommendation and for summarizing
the outcome of news stories on the public, de ned in terms of the shorter-term
e ects that news can have, such as informing, engaging, and mobilizing audiences
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The purpose of the the recommender is to support journalists in the task of
reshaping and extending their coverage of breaking news, by suggesting topics
to address when following up on such news.
      </p>
      <p>The paper is organized as follows: in section 2 we review related works, in
section 3 we describe our dataset and additional resources used in our methodology,
which is presented in section 4. Finally, section 5 is dedicated to the discussion
of the preliminary results and section 6 contais the concluding remarks and the
future work directions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>To the best of our knowledge, this is the rst system for recommending
journalists what to write, focusing on presenting users' needs that come from di erent
sources while keeping their original motivation (information and
communication).</p>
      <p>
        Many available studies are concerned with the task of predicting the response
of social media to news articles [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] rather than extracting users' interest
related to news articles to help journalists focalize on additional, yet uncovered,
aspects of a reported event. Other works analyze the symmetric problem of
recommending news to social media users. Among these, the authors in [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] are
concerned with the task of recommending articles to readers in a stream-based
scenario, when large user-item matrixes are not available and time constraints
are strict. In their work, they derive a number of statistics extracted from the
PLISTA [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] dataset used during the ACM Recsys News Challenge 2013. They
also compare performances of several existing recommending algorithms
showing that the precision of algorithms depends upon the particular news articles
domain. The study in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] also deal with real time recommenders. As before,
the considered task is to recommend topical news to users. The authors present
Buzzer, a system able to mine real time information from Twitter and RSS feeds
and use overlapping keywords in most recent tweets and feeds as a basis for
recommendation. Evaluation is performed on a small group of 10 participants over
a period of 5 days.
      </p>
      <p>
        A line of study closer to our work is concerned with the task of identifying
social content related to a given event. In [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] the following task is considered:
given a news article, nd the Twitter messages that "implicitly" refer to the same
topic, i.e. messages not including an explicit link to the considered article. They
are interested in discovering utterances that link to a speci c news article rather
than the news event(s) that the article is about. First, the authors analyze the
KL-divergence between the vocabulary of news articles (using the NT Times as a
primary source) and various social media, such as Twitter, Wikipedia, Delicious,
etc. They nd that, unless part of the original article is copied in the message,
which subsumes explicit reference, the vocabularies might be quite di erent.
The method used by the authors is in three steps: they derive multiple query
models from a given article, which are then used to retrieve utterances from
a target social media index, resulting in multiple ranked lists that are nally
merged using data fusion techniques. Evaluation is performed, in line with other
scholars, using messages with explicit mention to an article, and then removing
the mention. However, as observed by the same authors, evidence suggests that
these messages often copy part of the article, an eventuality that could boost
performances. In [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] the objective is to combine news articles and tweets to
identify not only relevant events but also the opinions expressed by social media
users on the very same event. Like us, they use the news article as the query,
and tweets as the document collection. They use a latent topic model to nd the
most relevant tweets wrt a given news topic. Besides topic similarity, they use
additional features such as recency, follower count etc, which are then combined
using logistic regression or Adaboost. Relevance judgement for evaluating the
system have been collected from 11 computer science students.
      </p>
      <p>
        Only two papers aim to help journalists nd relevant content in social
media, as we do. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] the authors present a tool to help journalists at identifying
eyewitnesses in the context of an event. In [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] a system is described to assist
journalists in the use of social media. The authors use SVM to identify
newsworthy messages on Twitter based on a manually annotated dataset. Their work
however is concerned more with the design of a user interface to help
journalists in digging into trending topics than on algorithms to extract such content
automatically.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Datasets and Resources</title>
      <p>To conduct our study, we have created three datasets: Wikipedia PageViews (W),
On-line News (N) and Twitter messages (T). Data has been collected during 4
months from June 1st, 2014 to September 30th, 2014 in the following way:
1. Wikipedia PageViews: we downloaded Wikipedia page views statistics
from the data dumps provided by the WikiMedia foundation3. We
considered only English queries and we retained only those matching a Wikipedia
document, removing redirected requests. Overall, we obtained 27.708.310.008
clicks on about 388 million pages during the considered period. An
example is: en Ravindra Jadeja 12 345 where the number of requests (12 in the
example) refers to a time span of 1 hour.
2. On-line News: We collected news from GoogleNews (GN) 4 and HighBeam
(HB) 5. Due to existing limitations, we extracted at most 100 news per day
from GN, while for HB we downloaded all available news. Each news item
has a title, source, day of publication and an associated snippet, e.g., GN
"8 1 2014","Bleacher Report",6,"India Will Be Left Furious by the Ravindra
Jadeja and James","On Friday, following a six-hour hearing in
Southampton, judicial commissioner Gordon Lewis found both Anderson and Jadeja
not guilty of breaching the ICC ". Overall, we extracted 351,922 news from 88
sources in GN and 1,181,166 from 325 sources in HB during the considered
period. Snipptes were about 25 words long in average.
3. Twitter messages: we collected 1% of Twitter tra c, the maximum freely
allowed tra c stream using the standard Twitter API 7. Overall, we
collected 235 million tweets, e.g., "James Anderson and Ravindra Jadeja have
both been found not guilty by judicial commissioner Gordon Lewis. #Cricket
#ENGvIND".</p>
      <p>
        Furthermore, in this research we used the following resources:
1. NASARI embedded semantic vectors for Wikipedia pages, generated as
described in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We used the second release8 covering 4.40 million Wikipedia
pages.
2. Dadelion Entity Extraction API (DataTXT)9 and TextRazor10. Both are
commercial tools providing entity recognition REST APIs that, given a
text snippet, identify, disambiguate and link named entities to Wikipedia.
DataTXT is based on previous research [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and has been recently further
developed and engineered [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
3 https://dumps.wikimedia.org/other/pagecounts-raw/
4 https://news.google.com/
5 https://www.highbeam.com
6 http://bleacherreport.com/articles/2148916
7 https://dev.twitter.com/docs/streaming-apis
8 http://lcl.uniroma1.it/nasari/\#two
9 https://dandelion.eu/semantic-text/entity-extraction-demo/
10 https://www.textrazor.com
      </p>
      <sec id="sec-3-1">
        <title>Overview</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Proposed Method</title>
      <p>
        Our methodology is based on the following steps:
1. Events identi cation: rst, we identify breaking news using SAX++, an
enhanced version of the temporal mining algorithm presented in [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] and [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
Terms related to breaking news are extracted from on-line news (N ), Twitter
(T ) and Wikipedia page-views (W ), respectively in representation of media
news providers and of public's communication and information needs. Terms
are grouped into clusters represented as ranked lists of words and named
entities;
2. Intra-source clustering : within each data source (N , T and W ) we attempt
to group terms clusters with the same peak day and related to the same
breaking news, creating meta-clusters.
3. Inter-source alignment : an alignment algorithm explores possible matches
across the three data sources N , T and W . For breaking news ni, we thus
obtain three meta-clusters mirroring respectively the media coverage of the
considered event, and its impact on readers' communication and information
needs.
4. Identi cation of missing information : the nal step is comparing the three
meta-clusters to identify in T and W meta-clusters the most relevant words
and named entities, considering both their impact on users and novelty wrt
to what has already been published in N . These terms can then be used to
recommend journalists additional aspects to cover or deepen when following
up on a news item. At the current stage of our research, we are in the phase
of de ning an e ective method to automatically derive, rank and evaluate
such recommendations. However, in section 5 we discuss preliminary results
in this direction.
4.2
      </p>
      <sec id="sec-4-1">
        <title>Events identi cation</title>
        <p>
          Algorithm In this section we shortly summarize the SAX++ algorithm, a
multi-thresholds version of the SAX algorithm, presented in [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] and [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], with
pluggable domain dependent variants. The phases of the new algorithm, named
SAX++, are the followings:
1. The temporal series associated to terms/wikipages and hashtag are sliced
into sliding windows of length W , normalized and converted in symbolic
strings using Symbolic Aggregate ApproXimation [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. The parameters of
this step are the dimension of the alphabet j j and the number W of
partitions of equal length .
2. Using a set of seed keywords related to known events, we convert their
temporal series into symbolic strings and automatically learn regular expressions
representing common usage patterns. For example, with an alphabet of 3
symbols, we learn the following expression:
        </p>
        <p>
          (a + [bc]?[bc][bc]?a+)?(a + [bc]?[bc][bc]a )?
which captures all the temporal series with one or two peaks and/or plateaus
in the analyzed window. These are common temporal patterns of breaking
news. Only terms/wikipages with frequency higher than a threshold f 0 and
hashtags with frequency higher than a threshold f 00 and matching the learned
regular expressions are considered in the subsequent steps. These are
hereafter denoted as active tokens.
3. Tokens are analyzed in sliding windows Wi and the detected active tokens
are clustered in each Wi using a bottom-up hierarchical clustering algorithm
with complete linkage [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] and similarity threshold .
4. In order to cope with temporal collision (i.e. co-occurring but unrelated
events) an additional cluster splitting step is performed, which considerably
improves clustering results. First, we build a graph G = (V; E) for each
cluster c previously detected by SAX*11 in a window W . A graph G is built
associating each vertex v 2 V with a token ti and adding an edge (ti,tj ) if
token ti and tj :
{ co-occurs in a number of documents greater than a threshold (for social
networks or for news domain);
{ or show "su cient" semantic similarity, as derived by an external
resource (for Wikipedia domain); speci cally, we use NASARI vectors to
compute similarity between two Wikipedia Pages that must be higher
than a threshold nas 12;
Next, we detect connected components in G. Each connected component is a
split of the original cluster. Extracting connected components from a graph
is a well-established problem [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and does not have a heavy impact on the
computational cost of the entire algorithm [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], mainly due to the fact that
the size of the graphs, when pruning non-active tokens, is relatively small.
Entities Enrichment Clusters extracted from News and Twitter are made out
of single words, however, including named entities would be preferable both for
relevance (named entities are pervasive in texts and especially in news) and to
compare results among di erent data sources (N , W and T ). Named Entity
expansion is applied only as post-processing phase over all the produced clusters,
because extracting names on the full Twitter and News stream is
computationally very demanding. Starting from all the tokens included in a cluster, we
retrieved the d best matching documents querying the original data-set (tweets
or news). We performed entity recognition for those documents only. Next, we
consider all the named entities recognized in these d documents, we score them
by occurrence, from 0 to 1, and add the top e entities to the related clusters.
11 for e ciency we maintain just one graph at a time in memory
12 Experiments shows that reeling on the Wikipedia graph doesn't split clusters e
ectively.
        </p>
        <p>Note that for every source (Twitter, News) we adopted the entities tagger
that best performs on the input text (see subsection 4.2).</p>
        <p>Parameter Setting In our experiments we ran SAX++ using di erent
parametrizations for each of the three sources, then we manually evaluated the resulting
clusters considering 10 known events to select the best parameters con guratios,
shown in table 1. The value of the parameter (the time granularity) was set
to 24 hours in all the datasets as this is the mininum granularity in news, where
the exact time of publication is not present.</p>
        <p>As tweets and news are very di erent in nature (length and style), before
performing the Named Entity Enrichment phase we processed some sample
documents with two available systems, DataTXT and Textrazor, and selected the
tool which provided more accurate results: DataTXT for news articles an
TextRazor for tweets. We then experimentally set optimal values for d and e (see
table 1).
Since SAX++ works on sliding windows, the same event is usually captured by
di erent clusters extracted from adjacent windows. In order to obtain a better
characterization of an event, we aggregate similar clusters with the same peak
day, forming meta-clusters which contain the most relevant terms for the event.
When considering clusters of the same events we note that the pivot cluster,
i.e. the cluster whose peak day is closer to the centre of the considered window,
shows a higher precision as compared to those clusters twith a peak day closer
to the extremes of the window.</p>
        <p>First, we select all the pivot clusters P d from the set of clusters Cd with
the same peak day. We then build a similarity graph GJ = (Cd; E ) (where an
edge is built if two clusters have a Jaccard similarity coe cient higher than a
threshold ), comparing every pivot cluster p 2 P d with all the clusters of the
original set, c 2 Cd n P d. Each extracted connected components from the graph
GJ , represents a meta-clusters composed by the identi ed clusters.</p>
        <p>To represent each meta-cluster in a compact and readable way, we create a
scored list of all the terms (words, hash-tags ad named entities) contained in the
meta-cluster. The score of a term is calculated as the normalized ratio between
the sum of the terms scores in all the clusters and the number of clusters. We
refer to such a scored set of terms as the meta-cluster signature.
4.4</p>
      </sec>
      <sec id="sec-4-2">
        <title>Inter-source Alignment</title>
        <p>The subsequent phase aligns meta-clusters from the three sources (T, N and W)
corresponding to the same popular event. We use as "seeds" the News
metaclusters, and nd the most similar meta-clusters from Twitter and Wikipedia.
As there might be a slight di erence in peak days in di erent data sources for
the same event, we use a similarity measure T empSym with two components: a
content based component and a time based one. The content based component
is the Jaccard similarity between terms of the meta-cluster's signature, while
the time based component takes into account the distance between the two peak
days: the closer the two, the higher the similarity. Considering two meta-clusters
ma and mb we use the following formula:</p>
        <p>T empSym(ma; mb) = J accard(ma; mb)
(jpeak(ma) peak(mb)j)</p>
        <p>Where alpha is a decay coe cient. The smallest is alpha, the less past clusters
are considered similar.
dataset
News
Twitter
Wikipedia
average size of
meta-clusters
122.46
136.76
6.44</p>
        <p>In table 2 we show some statistics of the obtained results wrt the three data
sources: the total number of clusters extracted by SAX++, the total number
of meta-cluster obtained running the intra-source clustering and the size of the
meta-clusters expressed as the average number of terms in such meta-cluster's
signatures.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion of preliminary results</title>
      <p>We present here an example of the result produced by the intra-source clustering
phase and three examples produced by the inter-source alignment algorithm;
these examples allow us to present a qualitative discussion of our results and to
show the di culty of a systematic evaluation.</p>
      <p>In table 3 we show a sample meta-cluster obtained as output of the
intrasource clustering step, along with its composing clusters and the scored set of
terms derived as discussed above. The meta-cluster clearly refers to a popular
event: the crash of the Malaysia Airlines ight 17 on 17 July 201413.
T 1405296000000 C5 [tragic, tragedi, Airline 1.0, Malaysia Airlines Flight 17 0.97,
Malaysia Airlines 0.70, Malaysia 0.58, Ukraine 0.40, Twitter 0.39, Gaza Strip 0.32, Barack Obama
0.32, Vladimir Putin 0.27, CNN 0.26, Tragedy 0.25, God 0.24, Airliner 0.22, Israel 0.22,
Malaysia Airlines Flight 370 0.22, Netherlands 0.21, ... ]
T 1405123200000 C17 [tragedi, tragic, Malaysia Airlines Flight 17 1.0, Airline 0.89,
Malaysia Airlines 0.62, Malaysia 0.54, Tragedy 0.47, Gaza Strip
0.42, Twitter 0.38, Ukraine 0.38, Hamas 0.33, Barack Obama 0.32, Israel 0.29, Vladimir Putin
0.27, God 0.26, CNN 0.25, Hell 0.25, Airliner 0.23, Malaysia Airlines Flight 37 0.20, ...]
T 1404950400000 C36 [crash, russian, tragedi, tragic, ukrain, Ukraine 1.0, Malaysia Airlines 0.42,
Malaysia 0.400, Russia 0.40, Airline 0.36, Airliner 0.18, Malaysia Airlines Flight 17 0.18,
Aviation accidents and incidents 0.13, Vladimir Putin 0.114, Kuala Lumpur 0.09, Eastern Ukraine
0.09, Boeing 777 0.075, Jet aircraft 0.068, ... ]
T 1405209600000 C5 [tragedi, tragic, Malaysia Airlines Flight 17 1.0, Airline 0.98,
Malaysia Airlines 0.80, Tragedy , Malaysia 0.54, Gaza Strip 0.50, Ukraine 0.48, Hamas 0.408,
Israel 0.38, Barack Obama 0.7, Twitter 0.37, Vladimir Putin 0.36, CNN 0.32, Airliner 0.28,
Malaysia Airlines Flight 370 0.26, Hell 0.252, God 0.25, ...]
Meta-cluster signature
17 Jul 2014 [tragedi 0.22, tragic 0.22, airline 0.20, malaysia airlines ight 17 0.20, ukraine 0.19,
malaysia airlines 0.19, malaysia 0.17, russia 0.129, tragedy 0.12, vladimir putin 0.12, airliner 0.12,
crash 0.12, gaza strip 0.11, barack obama 0.11, aviation accidents and incidents 0.11, cnn 0.106,
malaysia airlines ight 370 0.10, god 0.10, ... ]</p>
      <p>In table 4 we show some examples of the inter-source alignment algorithm.
Three popular events from di erent domains are considered: the celebration
of the USA Independence Day, the FIFA 2014 World Cup nal match and the
Malaysia Airlines crash. For each event we show the news meta-cluster signature,
used as seed in the alignment algorithm, and the most similar meta-cluster's
signatures emerged in Twitter and Wikipedia. In addition, we mark in bold the
novel terms in T and W , which could be the candidate recommended terms. Due
to lack of space, we do not show all the terms, except for the Wikipedia where
the meta-clusters are smaller on average.</p>
      <p>Some emerging terms, especially in Wikipedia cluster, clearly highlight
information needs related to the corresponding events. For example, the emerging of
terms like american revolutionary war and the start-spangled banner suggests a
keen interest to deepen the knowledge of the historical events that led to the
13 BBC page of the event: http://www.bbc.com/news/world-europe-28357880
Independence Day (04 Jul 2014)
News
03 Jul 2014 [united states declaration of independence 0.25, independence day 0.24,
life, liberty and the pursuit of happiness 0.24, natural and legal rights 0.14, continental congress
0.13, all men are created equal 0.12, thomas je erson 0.12, self-evidence 0.12, washington, d.c.
0.12, human events 0.12, reworks 0.11, united states house of representatives 0.10 ... ]
Twitter
03 Jul 2014 [independence day 0.29, textb ourth 0.28, safe 0.22, bbq 0.16, grill 0.16, sparkler
0.15, reworks 0.14, united states 0.14, barbecue 0.13, parad 0.12, co ee 0.10, god 0.10,
pittsburgh steelers 0.10, heinz eld 0.09, canada 0.09, ... ]
Wikipedia
Jul 04 2014 [the star-spangled banner 0.16, independence day 0.16,
american revolutionary war 0.12]
FIFA World Cup 2014 nal match (13 Jul 2014)
News
13 Jul 2014 [germany national football team 0.30, fa world cup 0.27, overtime (sports) 0.27,
argentina 0.26, germany 0.26, argentina national football team 0.24, brazil 0.24, mario goetze 0.24,
rio de janeiro 0.23, maracana stadium 0.22, 2014 fa world cup 0.22, brazil national football team
0.21, lionel messi 0.21 ... ]
Twitter
13 Jul 2014 [shakira 0.33, gervsarg 0.29, kramer 0.29, gerarg 0.29, argvsger 0.29,
germany national football team 0.29, lionel messi 0.28, argentina national football team 0.26,
argentina 0.24, champion 0.22, germany 0.21, fa world cup 0.21, neuer 0.19, ceremoni 0.19 ...
]
Wikipedia
13 Jul 2014 [2018 fa world cup 0.19, 2026 fa world cup 0.19, 2022 fa world cup 0.12]
Malaysia Airlines igth 17 crash (17 Jul 2014)
News
17 Jul 2014 [malaysia 0.38, ukraine 0.33, malaysia airlines 0.33, surface-to-air missile 0.30,
kuala lumpur 0.28, eastern ukraine 0.27, malaysia airlines ight 17 0.27, boeing 78 0.272,
amsterdam 0.27, ... ]
Twitter
17 Jul 2014 [malaysia 0.37, aircraft 0.31, plane 0.29, condol 0.29, malaysian 0.29, airlin 0.29,
missil 0.29, passeng 0.29, ukraine 0.29, malaysia airlines 0.28, ukrain 0.28, ... , tragedi 0.12,
eastern ukraine 0.11, aviation accidents and incidents 0.11, boeing 777 0.11, missile 0.11,
passenger 0.11, kuala lumpur international airport 0.10, interfax 0.09, jet aircraft 0.09, , ... ,
russian language 0.08, ..., tragic 0.06, expens 0.06, ...]
Wikipedia
17 Jul 2014 [siberia airlines ight 0.93, boeing 0.85, malaysia airlines ight 0.73, iran air ight
0.71, korean airlines ight 0.44, pan am ight 0.41, kuala lumpur 0.28,
surface-toair missile 0.26, malaysia airlines 0.26, malaysia 0.24, buk missile system 0.24, ukraine 0.12,
bermuda triangle 0.06, 2014 crimean crisis 0.06]</p>
      <p>US independence and of the US national anthem, respectively. These could be
topics that are worth deepening, e.g., in editorials. Looking at T meta-clusters,
the emerging of popular terms like bbq, grill or parad (stem of parade) in tweets
immediately before Independence Day may simply suggest that most people are
preparing to celebrate, while other terms like pittsburg steelers and heinz eld
refer to co-occurring related sports events and could be reasonably labelled as
noise.</p>
      <p>Looking at the second event, terms like (2018|2022|2026) fa world cup
in the W meta-cluster expose a widespread interest in future editions of the
football World Cup, which, again, could suggest related topics to be deepened.
In the T meta-cluster, the appeareance of the term shakira, referring to the
popular singer, in association with the FIFA football match seems apparently
unrelated. However "googling" the term highlights a strong connection, as the
singer sang the theme song of the 2014 World Cup during the FIFA world cup
closing ceremony; ceremoni is another term in the same cluster, con rming this
interpretation. The terms gerarg and argvsger are popular hashtags used to
comment the match on Twitter; while not novel per-se, nding relevant hashtags
for an event may prove useful in some contexts.</p>
      <p>Around the time of the Malaysia Airlines crash it is not surprising that most
people are encouraged to check Wikipedia about similar incidents in the past,
e.g., Siberia Airlines ight 1812, shot down by the Ukrainian Air Force over
the Black Sea in 2001, and about somehow related topics, e.g. bermuda triangle.
Finding past similar events is a common information need, frequently highlighted
in our data. Terms like condol and tragic mirror a popular mood emerging among
Twitter users, while for other terms it is hard to say if they are noisy or not, e.g.
expens. Finally, the term interfax, apparently unrelated, turned out to be related
to the event, since Interfax is a Moscow-based wire agency which reported that
Ukrainian rebel forces had the airplane black boxes and they had agreed to hand
them over to the Russian-run regional air safety authority; this news sub-topic
captured the attention on Twitter.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and future work</title>
      <p>In this paper we presented a methodology to derive recommendations for
journalists based on the detection and analysis of readers' information needs on
Wikipedia and communication needs on Twitter. Preliminary experiments
suggest that our methodology succeeds in aligning manifestations of interest wrt
to highly popular events in all three di erent information sources: online news,
twitter and wikipedia. No strong conclusions can be drawn from our preliminary
experiments, even if results suggest that many relevant and recurrent behaviours
can be extracted by systematically comparing the three sources.</p>
      <p>
        Future works are focused on de ning an e ective method to extract missing
information that might provide useful insight for press agents, from T and W
aligned meta-cluster of a given News n 2 N . Measuring the quality of the
results in a meaningfull way is still an open issue. Speci cally, we want to identify
terms that expose relevant and novel topics wrt to what have already been
published. Several works [
        <xref ref-type="bibr" rid="ref10 ref14 ref20 ref29 ref5 ref9">5, 9, 10, 14, 20, 29</xref>
        ] try to formally de ne how to evaluate
relevance and novelty, but this task still remain di cult even for humans, since
connections among topics might not be evident and what may seem noise at
rst glance, turn out to be relevant and interesting after an in depth (and
timeconsuming) analysis. On the other hand, trying to reduce noise, e.g. discarding
loosely connected topics, could dramatically a ect novelty.
7
      </p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We thank SpazioDati and TextRazor for granting us the use of their entity
recognition API beyond the non commercial use limit. We would also like to thank
Giacomo Marangoni for his kind support in developing the work ow related to
the Wikipedia source.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Brooker</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schaefer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Public Opinion in the 21st Century: Let the People Speak? New directions in political behavior series</article-title>
          , Houghton Mi in Company (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Camacho-Collados</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pilehvar</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Navigli</surname>
          </string-name>
          , R.:
          <article-title>Nasari: a novel approach to a semantically-aware representation of items</article-title>
          .
          <source>In: Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          . pp.
          <volume>567</volume>
          {
          <fpage>577</fpage>
          . Association for Computational Linguistics, Denver, Colorado (May{June 2015), http://www.aclweb.org/ anthology/N15-1059
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Diakopoulos</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Choudhury</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naaman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Finding and assessing social media information sources in the context of journalis</article-title>
          .
          <source>In: Proceedings of the 2012 ACM annual conference on Human Factors in Computing Systems</source>
          . pp.
          <volume>24151</volume>
          {
          <fpage>2460</fpage>
          .
          <source>CHI</source>
          <year>2012</year>
          , ACM, New York, NY, USA (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scaiella</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          : Tagme:
          <article-title>On-the- y annotation of short text fragments (by wikipedia entities)</article-title>
          .
          <source>In: Proceedings of the 19th ACM International Conference on Information and Knowledge Management</source>
          . pp.
          <volume>1625</volume>
          {
          <fpage>1628</fpage>
          . CIKM '10,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2010</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1871437.1871689
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ge</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delgado-Battenfeld</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jannach</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Beyond accuracy: Evaluating recommender systems by coverage and serendipity</article-title>
          .
          <source>In: Proceedings of the Fourth ACM Conference on Recommender Systems</source>
          . pp.
          <volume>257</volume>
          {
          <fpage>260</fpage>
          . RecSys '10,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2010</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1864708.1864761
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gloviczki</surname>
            ,
            <given-names>P.J.:</given-names>
          </string-name>
          <article-title>Journalism in the Age of Social Media</article-title>
          , pp.
          <volume>1</volume>
          {
          <fpage>23</fpage>
          .
          <string-name>
            <surname>Palgrave Macmillan</surname>
            <given-names>US</given-names>
          </string-name>
          , New York (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hopcroft</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarjan</surname>
          </string-name>
          , R.:
          <article-title>E cient algorithms for graph manipulation</article-title>
          .
          <source>Commun. ACM</source>
          <volume>16</volume>
          (
          <issue>6</issue>
          ),
          <volume>372</volume>
          {378 (Jun
          <year>1973</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/362248.362272
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>A.K.</given-names>
          </string-name>
          :
          <article-title>Data clustering: 50 years beyond k-means</article-title>
          .
          <source>Pattern Recogn. Lett</source>
          .
          <volume>31</volume>
          (
          <issue>8</issue>
          ),
          <volume>651</volume>
          {666 (Jun
          <year>2010</year>
          ), http://dx.doi.org/10.1016/j.patrec.
          <year>2009</year>
          .
          <volume>09</volume>
          .011
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jenders</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lindhauer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krestel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naumann</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A Serendipity Model for News Recommendation</article-title>
          , pp.
          <volume>111</volume>
          {
          <fpage>123</fpage>
          . Springer International Publishing,
          <string-name>
            <surname>Cham</surname>
          </string-name>
          (
          <year>2015</year>
          ), http://dx.doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -24489-
          <issue>1</issue>
          _
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Kaminskas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bridge</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Measuring surprise in recommender systems</article-title>
          . In: Adamopoulos,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Bellogn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Castells</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Cremonesi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Steck</surname>
          </string-name>
          , H. (eds.)
          <source>Procs. of the Workshop on Recommender Systems Evaluation: Dimensions and Design (Workshop Programme of the Eighth ACM Conference on Recommender Systems)</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Kille</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hopfgartner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brodt</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heintz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The plista dataset</article-title>
          .
          <source>In: Proceedings of the 2013 International News Recommender Systems Workshop and Challenge</source>
          . pp.
          <volume>16</volume>
          {
          <fpage>23</fpage>
          . NRS '13,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2013</year>
          ), http://doi.acm.
          <source>org/10</source>
          . 1145/2516641.2516643
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Journalism as usual: The use of social media as a newsgathering tool in the coverage of the iranian elections in 2009</article-title>
          .
          <source>Journal of Media Practice</source>
          <volume>13</volume>
          (
          <issue>1</issue>
          ),
          <volume>61</volume>
          {
          <fpage>74</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. Konig,
          <string-name>
            <given-names>A.C.</given-names>
            ,
            <surname>Gamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Q.</surname>
          </string-name>
          :
          <article-title>Click-through prediction for news queries</article-title>
          .
          <source>In: Proceedings of the 32Nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . pp.
          <volume>347</volume>
          {
          <fpage>354</fpage>
          . SIGIR '09,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2009</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1571941.1572002
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kotkov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veijalainen</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A survey of serendipity in recommender systems</article-title>
          .
          <source>Know.-Based Syst. 111(C)</source>
          ,
          <volume>180</volume>
          {192 (Nov
          <year>2016</year>
          ), http://dx.doi.org/ 10.1016/j.knosys.
          <year>2016</year>
          .
          <volume>08</volume>
          .014
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Krestel</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Werkmeister</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiradarma</surname>
            ,
            <given-names>T.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kasneci</surname>
          </string-name>
          , G.:
          <article-title>Tweet-recommender: Finding relevant tweets for news articles</article-title>
          .
          <source>In: Proceedings of the 24th International Conference on World Wide Web</source>
          . pp.
          <volume>53</volume>
          {
          <fpage>54</fpage>
          . WWW '15 Companion,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lehmann</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goncalves</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramasco</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cattuto</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Dynamical classes of collective attention in twitter pp</article-title>
          .
          <volume>251</volume>
          {
          <issue>260</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Leskovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Backstrom</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleinberg</surname>
          </string-name>
          , J.:
          <article-title>Meme-tracking and the dynamics of the news cycle</article-title>
          .
          <source>In: Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>497</volume>
          {
          <fpage>506</fpage>
          . KDD '09,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keogh</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lonardi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A symbolic representation of time series, with implications for streaming algorithms</article-title>
          .
          <source>In: Proceedings of the 8th ACM SIGMOD Workshop on Research Issues in Data Mining and Knowledge Discovery</source>
          . pp.
          <volume>2</volume>
          {
          <fpage>11</fpage>
          . DMKD '03,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2003</year>
          ), http://doi.acm.
          <source>org/10</source>
          . 1145/882082.882086
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Lommatzsch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albayrak</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Real-time recommendations for user-item streams</article-title>
          .
          <source>In: Proc. of the 30th Symposium On Applied Computing</source>
          ,
          <string-name>
            <surname>SAC</surname>
          </string-name>
          <year>2015</year>
          . pp.
          <volume>1039</volume>
          {
          <fpage>1046</fpage>
          . SAC '15,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Murakami</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mori</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Orihara</surname>
          </string-name>
          , R.:
          <article-title>Metrics for evaluating the serendipity of recommendation lists</article-title>
          .
          <source>In: Proceedings of the 2007 Conference on New Frontiers in Arti cial Intelligence</source>
          . pp.
          <volume>40</volume>
          {
          <fpage>46</fpage>
          . JSAI'
          <volume>07</volume>
          , Springer-Verlag, Berlin, Heidelberg (
          <year>2008</year>
          ), http://dl.acm.org/citation.cfm?id=
          <volume>1788314</volume>
          .
          <fpage>1788320</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Napoli</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          :
          <article-title>Measuring media impact an overview of the eld</article-title>
          . Rutgers University (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Phelan</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCarthy</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Using twitter to recommend real-time topical news</article-title>
          .
          <source>In: Proceedings of the Third ACM Conference on Recommender Systems</source>
          . pp.
          <volume>385</volume>
          {
          <fpage>388</fpage>
          . RecSys '09,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Reingold</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Undirected connectivity in log-space</article-title>
          .
          <source>J. ACM</source>
          <volume>55</volume>
          (
          <issue>4</issue>
          ),
          <volume>17</volume>
          :1{
          <fpage>17</fpage>
          :24 (Sep
          <year>2008</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/1391289.1391291
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Scaiella</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prestia</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Del Tessandoro</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ver</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barbera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Parmesan</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Datatxt at microposts2014 challenge</article-title>
          . vol.
          <volume>1141</volume>
          , pp.
          <volume>66</volume>
          {
          <issue>67</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Stilo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velardi</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>E cient temporal mining of micro-blog texts and its application to event discovery</article-title>
          .
          <source>Data Min. Knowl. Discov</source>
          .
          <volume>30</volume>
          (
          <issue>2</issue>
          ),
          <volume>372</volume>
          {402 (Mar
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Stilo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velardi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Hashtag sense clustering based on temporal similarity</article-title>
          .
          <source>Computational</source>
          Linguistics pp.
          <volume>1</volume>
          {
          <issue>32</issue>
          (
          <year>2017</year>
          /01/17 2016), http://www. mitpressjournals.org/doi/abs/10.1162/COLI_a_
          <fpage>00277</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Tsagkias</surname>
            , E., de Rijke,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weerkamp</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Predicting the volume of comments on online news stories</article-title>
          .
          <source>In: Proceedings of CIKM 09. ACM</source>
          , New York, NY, USA (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Tsagkias</surname>
          </string-name>
          , M.,
          <string-name>
            <surname>de Rijke</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weerkamp</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Linking online news and social media</article-title>
          . In: King,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Nejdl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>H</surname>
          </string-name>
          . (eds.) WSDM. pp.
          <volume>565</volume>
          {
          <fpage>574</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Vargas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castells</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Rank and relevance in novelty and diversity metrics for recommender systems</article-title>
          .
          <source>In: Proceedings of the Fifth ACM Conference on Recommender Systems</source>
          . pp.
          <volume>109</volume>
          {
          <fpage>116</fpage>
          . RecSys '11,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2011</year>
          ), http://doi.acm.
          <source>org/10</source>
          .1145/2043932.2043955
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Zubiaga</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ji</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Knight</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Curating and contextualizing twitter stories to assist with social newsgathering</article-title>
          .
          <source>In: 18th International Conference on Intelligent User Interfaces</source>
          ,
          <source>IUI '13</source>
          ,
          <string-name>
            <surname>Santa</surname>
            <given-names>Monica</given-names>
          </string-name>
          , CA, USA, March
          <volume>19</volume>
          -22,
          <year>2013</year>
          . pp.
          <volume>213</volume>
          {
          <issue>224</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Zuckerberg</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>New facebook project to ght 'fake news'</article-title>
          (
          <year>Jan 2017</year>
          ), https: //www.dawn.com/news/1307895/new-facebook
          <article-title>-project-to-fight-fake-news</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>