<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>BroDyn</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Real-time collection of reliable and representative tweets datasets related to news events</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Beatrice Mazoyer</string-name>
          <email>bmazoyer@ina.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julia Cage</string-name>
          <email>julia.cage@sciencespo.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Celine Hudelot</string-name>
          <email>celine.hudelot@centralesupelec.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marie-Luce Viaud</string-name>
          <email>mlviaud@ina.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CentraleSupelec, Mathematics interacting with computer science laboratory</institution>
          ,
          <addr-line>Gif-sur-Yvette</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institut National de l'Audiovisuel</institution>
          ,
          <addr-line>Bry-sur-Marne</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Sciences Po Paris Paris</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <volume>1</volume>
      <fpage>23</fpage>
      <lpage>34</lpage>
      <abstract>
        <p>This paper is part of a wider work studying the co-in uences of Twitter and the production of information by traditional media. A strong prerequisite of this study is to collect, with the limitations of the Twitter API, tweets linked to media events that are representative of the real Twitter activity. This paper describes two proposed approaches to handle this important task. The rst one, inspired by information retrieval, puts the focus on query formulation. It consists on bridging the vocabulary gap between traditional news articles and tweets, by iteratively modifying the queries sent to the Twitter API depending on the tweets retrieved by previous queries. The second approach consists in streaming a representative sample of all emitted tweets and dynamically clustering them in events. We also discuss approaches to evaluate the collected datasets under the point of view of their representativity of the real activity on Twitter.</p>
      </abstract>
      <kwd-group>
        <kwd>Twitter</kwd>
        <kwd>news</kwd>
        <kwd>tweets retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>How does information propagate online? Does it propagate di erently on news
websites and on social media? What is the role of social networks and in
particular of Twitter in breaking news?</p>
      <p>In an ideal world, to compare news production on social media and on
mainstream media, one would need the universe during a given period of time (e.g.
the year 2017) and a geographical location (e.g. France, the UK or the US) of
documents published on the one hand on social media and on the other hand
on mainstream media. Unfortunately, given the limitation of the Twitter API,
it is not possible for the researcher to capture the universe of the documents (or
tweets) published on Twitter. Does it mean that the research wont be able to
answer the previously formulated questions? No. Because, the researcher can rather
use a random sample of the documents, as long as this sample is representative.
Why do we need representativity?</p>
      <p>Assume that you want to answer the following question: what is the
probability for a news story broken on social media to make it to the mainstream
media? With the entire set of news stories broken on social media, it will be
pretty simple to answer this question; but we dont have this data. Now assume
that we get access to a subsample of the documents published on Twitter, but
that this sample is not representative. For example, assume that this sample of
tweets is such that the tweets characteristics (perhaps because the API provides
the researcher with documents tweeted by users with more followers) are such
that these documents have a higher probability to make it to the mainstream
media. Then using this biased subsample will lead the researcher to overestimate
the probability for a news story broken on social media to appear on mainstream
media.</p>
      <p>The same issue will arise if the researcher wants to tackle the follow-up
question: what are the determinants of the success of a news story initially broken
on social media? Imagine that the researcher is using a selected sample of tweets
that is not representative. Imagine for example that this sample of tweets comes
mainly from journalists working for a given media, e.g. Le Monde, and that, at
the same time, within the set of tweets posted by Le Mondes journalists, only the
successful ones are part of the sample, then the results of the empirical analysis
will be biased in favor of Le Monde. In other words, when the researcher will
study the impact of the company for which the journalist work (independent
variable) on the probability for the news story broken on Twitter to make it to
mainstream media (dependent variable), the coe cient obtained for Le Monde
will ovestimate the real causal impact of the company.</p>
      <p>It seems very di cult to correct for this bias. Hence the necessity to have
a representative sample of tweets, i.e. a sample of tweets such that the tweets
included in our sample do not di er from the tweets that are not included along
all the dimensions that may have a direct impact on the dependent variable of
interest.</p>
      <p>
        In this paper, we aim at collecting, in real-time, tweets linked to French news
that are representative of the total Twitter activity concerning media events.
For now, we only work on tweets in French, which is an advantage in terms of
volume. Indeed, since the Twitter API puts high restrictions on the volume of
tweets returned, we get a largest proportion of emitted tweets by working on a
language that is not very represented on Twitter (French tweets represent 1.8%
of all emitted tweets). Based on the di erent ways of collecting tweets with the
Twitter API, we explore two types of approaches to collect tweets:
1. Vocabulary-constrained collection based on keywords extracted from AFP
dispatches. We use the stream of dispatches from the French news agency
\Agence France Presse" (AFP). This news agency, like Reuters or Associated
Press (AP), has a large network of journalists in many countries, that covers
a wide range of news topics. Previous research has shown that this source
covered 95% of all events covered by French news media in 2013 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
2. Random tweets collection followed by clustering. To get a continuous stream
of random tweets, we studied Twitter Sample API4. Indeed, previous works
show that the tweets from that API are \a representative sample of the true
activity on Twitter"[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, the number of French tweets in this sample
is very small (2400 tweets per hour on average). It is likely that small events
do not generate enough French tweets to be represented in that sample. We
thus describe, in Section 3.3, an approach to increase this sample of French
tweets.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>State of the art</title>
      <p>Collecting representative sets of tweets means performing together the tweets
collection and the evaluation of its representativity. We therefore present in this
part: (1) techniques to build sets of relevant tweets and (2) methods to evaluate
tweets sets. A tool designed to collect representative tweets sets should comprise
both aspects.
2.1</p>
      <sec id="sec-2-1">
        <title>Query building strategies</title>
        <p>
          To our knowledge, few approaches have been proposed in the literature to
perform targeted tweet collection based on news articles. Indeed, the literature
related to the task of linking tweets and news [3{8] mainly used existing datasets,
mostly in English. These articles do not address the issue of real-time collection
of representative tweets. Becker et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] do not directly work on news articles,
but they de ne query building strategies based on the title, description, time
and location of a selected set of expected events to collect related tweets. Tanev
et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] extract 1-grams and 2-grams from the title and rst sentence of news
articles, and weight them depending on their IDF in a one million news articles
collection. They build queries out of these n-grams and explore several
queryexpansion strategies based on co-occurrences in their news articles corpus. Ning
et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] build chains of articles concerning the same event and rst collect
tweets containing the articles' urls, then extract a list of top ten keywords from
those tweets and collect tweets containing these keywords. This method is an
attempt to bridge the vocabularies between tweets and news by using keywords
from the tweets. Castillo et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], that share common objectives with our work,
also use urls to retrieve tweets related to news articles. However, our
observations show that most tweets containing a link to an article share only the title
or rst sentence of that article, without additional vocabulary. This approach
is restrictive and may lead to miss a large part of Twitter activity concerning
news. Our work aims at building queries that contain Twitter-speci c terms
in order to achieve higher recall than previously introduced methods, without
losing precision.
4
https://developer.twitter.com/en/docs/tweets/sample-realtime/api-reference/getstatuses-sample
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Representativity of tweets collection</title>
        <p>
          Few papers address the problem of the representativity of the collected tweets:
we could only nd three studies addressing this issue [
          <xref ref-type="bibr" rid="ref13 ref14 ref2">13, 2, 14</xref>
          ]. Morstatter et al.
and Wang et al. [
          <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
          ] propose methods to evaluate the Twitter Filter API5.
This API returns a stream of tweets that match one or several predicates given
as parameters. Both articles point out some biases in the collected datasets. In
another article [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], Morstatter et al. evaluate the Sample API and conclude to
its representativity of the global Twitter activity. However this study does not
consider methods to collect larger representative datasets in a speci c language
as we do for French in this paper.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        We present here two approaches for tweet collection:
{ A tool designed to build queries in real time based on the stream of AFP
dispatches and expand them using Twitter speci c vocabulary. This approach
is novel with respect to existing works [9{11] since the step of query
expansion using word embeddings, while developed in other contexts such as web
retrieval [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], has, as far as we know, never been used in the context of
tweets retrieval.
{ A method to get a signi cant and representative set of random tweets in
French. Our work here aims at maximizing the volume of our random sample,
while keeping its distribution characteristics the closest to the one of the
5 https://developer.twitter.com/en/docs/tweets/
lter-realtime/api-reference/poststatuses- lter
random sample collected with Twitter Sample API. Events will then be
detected by clustering following the same procedure as [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
3.1
      </p>
      <sec id="sec-3-1">
        <title>Vocabulary-constrained collection</title>
        <p>Our query building tool is composed of several modules designed to (1) extract
relevant keywords from AFP dispatches and build queries with them, (2) expand
those queries using Twitter vocabulary, (3) validate that these queries return
tweets linked to the event. All steps are summarized in Figure 1.</p>
        <p>AFP dispatches comprise a title, a timestamp and a body containing a few
lines of text describing a news event. 900 dispatches are published every day in
average. We de ne a news event very simply by a new AFP dispatch with a title
containing less than k words in common with the titles of previous dispatches
published in a time window T . We do not proceed to more complex dispatch
clustering since our observations show that when AFP journalists publish several
dispatches on a given event, they usually keep the same title and only update
the main text with additional details. In practice we have set k to 3 and T to
24 hours. For each event, we aim at building a set of queries that will retrieve
tweets related to that event.</p>
        <p>
          Keywords extraction For a given dispatch, we extract named entities from
the title and the body. We then build queries using combinations of those entities
following a number of rules: two locations cannot form a query (to avoid queries
like \Alep Syria"), a query must contain at least two entities and less than four
entities. This method is quite similar from what is done by Becker et al. [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
We also take the exact title of the dispatch as a query, if it contains no named
entities. The tweets containing these queries are then collected.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Query expansion</title>
        <p>collected tweets:</p>
        <p>
          We combine several methods to build new queries from
{ Urls extraction: like Ning et al. [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], we extract urls from collected tweets
and extend our query set with these urls.
{ Keywords detection: we build a TF-IDF matrix from all terms of the tweets
collected in the past 7 days. All tweets related to the same event form a
document. We pick terms with a TF-IDF score higher than a de ned threshold
(di erent for 1-grams, 2-grams and 3-grams)
{ Synonyms: We use the CBOW model [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] with negative sampling [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] to
build vector representations of words in Twitter vocabulary and nd
potential \synonyms" (terms used in very similar contexts) to expand our rst
queries. We train the model on a sample of French tweets collected with the
Sample API in the past 7 days to get a vector representation of Twitter
words. We de ne as \synonyms" of a word the words that have a cosine
similarity higher than a certain threshold with that word-vector. For instance,
using 0.85 as a threshold, we get \fh" (abbreviation for \Francois Hollande")
as synonym for \hollande". We then build new queries by replacing terms of
previous queries with their synonyms if they have some.
Query validation The previous steps allow us to combine a large number of
words to build queries that might be related to the given event. However, these
methods can also produce wrong queries, either not precise enough (\Elysee
Macron" for example, since l'\Elysee" is the common name for the French
president's o ce), or clearly pointing to something or someone else than the entities
extracted from the original dispatch. For instance, the synonyms of \sarkozy"
are \sarko" (which is correct), \juppe" and \ llon" (who are politicians from
the same party as Nicolas Sarkozy, but not the same person). We therefore need
a method to lter incorrect queries.
        </p>
        <p>We can make the following hypothesis: if a news event is discussed on Twitter,
the tweets containing words related to that news should form a peak in a time
window of a few hours before or after the time of the event. If the tweets collected
through a certain query do not peak around the event's time, it is either because
the news is not discussed on Twitter, or because the query is not linked to the
event in question. Starting from that assumption we are working on an algorithm
to detect such anomalies on the time series of tweets and thus remove irrelevant
queries.</p>
        <p>Iteration The keywords extraction and query validation steps are repeated to
increase the number of tweets linked to the considered event, until the set of
collected tweet does not provide any new query.
3.2</p>
      </sec>
      <sec id="sec-3-3">
        <title>Evaluation of the tweets collection tool</title>
        <p>To the best of our knowledge, there is no dataset corresponding to our task. We
are currently designing a user interface to manually label tweets. This interface
will allow two types of evaluations:
{ Precision: given the text of an AFP dispatch and the list of tweets collected
by our tool as presumably linked to the dispatch, the user will assess if each
tweet is correctly attributed.
{ Recall: the user has to manually formulate queries to retrieve as many tweets
as possible in relationship to a given AFP dispatch, and assess if the returned
tweets are indeed linked to the dispatch. This method, however, is only a
partial solution to the problem of evaluating the proportion of tweets not
found by our tool, since the user will not necessarily nd all possible terms to
query relevant tweets. Another way of approximating the recall would be to
compare the proportion of tweets in each collected set to the size of clusters
identi ed with our second approach.
3.3</p>
      </sec>
      <sec id="sec-3-4">
        <title>Random tweets collection</title>
        <p>
          Twitter Sample API returns 1% of the world stream randomly selected[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], among
which 1.8% on average are in French. 1% of the French stream represents in
average only 2400 tweets an hour. A previous article [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] has estimated the
proportion of news-related documents in tweets corpora to less than 0.2%. If
this ratio is the same for French tweets, it means that we would receive between
4 and 5 news-related tweets in French per hour by using the Sample API, a
volume much too small for our goals.
        </p>
        <p>Using the Filter API on \neutral" terms with \French" as language
parameter allows us to collect a larger volume of tweets in French, but Twitter does
not communicate about the tweets distribution. Since we need to ensure that
the stream of collected tweets is representative of the tweets emitted on Twitter
at the same moment, we designed some methods to compare it with the stream
from the Sample API.</p>
        <p>
          Joseph et al. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] compare several samples collected with the Filter API on
the same keywords using di erent connection tokens6: they nd that two samples
taken at the same time with the same keywords as inputs are \nearly identical"
(96% of the tweets are the same, and the sets of tweets that are not have a very
similar structure in terms of number of hashtags, urls and mentions per tweet).
It is thus not possible to use several connections with the same keywords to get
a higher number of tweets. However, spreading di erent keywords over several
API connections should return a higher number of tweets. If these tweets are
representative of the real activity on Twitter, the distribution of words within it
should be the same (proportionally) to the one of tweets from the Sample API.
To ensure that it is the case, we worked on optimizing the collect parameters
(number of connections, number of keywords, distribution of keywords over the
di erent tokens) by performing the steps described below:
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Finding the most frequent keywords on Twitter. We collected tweets</title>
        <p>from the Sample API during three months and selected French tweets (retweets
excluded). After removing punctuation, capital letters and non-Latin characters,
we built a list of the 200 most frequent words in the French Twitter vocabulary.
Clustering keywords depending on their co-occurrences. We built a
matrix of co-occurrences of the rst 50, 100 and 200 most frequent words in the
sampled tweets, and used it to cluster keywords in n clusters (n 2 [2; 4]). By
doing so, we aimed at putting together terms that are frequently used together,
to collect sets of tweets with the smallest possible intersection. In total, we
had 9 ways of collecting tweets (50, 100 or 200 words, spread over 2, 3 or 4
clusters). To control that our clustering approach was the best, we also tried
to randomly spread the same number of words (50, 100 or 200) over the same
number of Twitter access tokens (2, 3 or 4). We ran each test during 24 hours
and compared the returned tweets with those obtained using the Sample API
during the same period of 24 hours.
6 To use the Twitter API, a connection token is required. Twitter limits the access
to its data by generating only one connection token per Twitter account. The total
number of tweets that one can get with only one token is limited to 1% of the global
tweets volume at a given moment.
Running tests simultaneously. In order to select the best collection method
we had to run series of tests in parallel: it is not possible compare tests conducted
on di erent periods of times since the di erences between the results could be
due to a di erent tweets distribution between periods. However, we only had
13 Twitter access tokens, and thus could not run all tests simultaneously. Since
a collection method requires 2 to 4 tokens, we could only run 6 to 3 tests in
parallel.
3.4</p>
      </sec>
      <sec id="sec-3-6">
        <title>Evaluation of random tweets collection</title>
        <p>We used Kullback-Leibler divergence to compare the words distributions of the
set of collected tweets to the distribution from the Sample API. The best
collection method should have a words distribution very similar to the distribution
of words collected with the Sample API, and thus the Kullback-Leibler
divergence between the two distributions should be close to zero. We also evaluated
our collection method with regards to the volume of collected tweets: we used
a simple ratio between the volume collected with each method and the volume
collected using the Sample API during the same period.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>Note: we have no results on the 01/22, since the connection to the API was broken.
4.1</p>
      <sec id="sec-4-1">
        <title>Tweets collection tool</title>
        <p>As explained in Section 3.2, we lack a labeled test-set to perform a correct
evaluation of our tweets retrieval method. Our work so far has been directed
towards the implementation of a robust architecture to stream and analyze a high
number of documents in near real-time. The development of our test interface
will allow us to provide concrete results in the near future. We also plan to
compare our strategy with the state of the art, presented in section 2.1.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Random tweets collection</title>
        <p>Our results are displayed in Figures 2 and 3. We get consistent results over
time: the collection method using words clustered by co-occurrences performs
better in terms of proximity to the Sample than other methods using random
distribution of words. The volume of collected tweets greatly di ers from one
collection method to another (from 50 to 75 times higher than the Sample),
but the trend also remains stable over time7. However, these results need to be
7 The drop in KL-divergence and in volume of collected tweets on the 01/21 could be
linked to the fact that it was a Sunday, but we need to reproduce the experiment
over a longer period to con rm that this drop is not a coincidence.
reproduced over a longer period to be con rmed, and we need to conduct wider
tests on all collection methods.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Related work</title>
      <sec id="sec-5-1">
        <title>Detecting events in the Twitter stream</title>
        <p>
          A large number of methods have been proposed for detecting \stories" [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] or
\events" [22{26] in a continuous stream of tweets. However these methods do
not link the detected clusters to news articles. Sankaranarayanan et al. [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ] use
handpicked \seeders" (Twitter accounts that are known to publish news) and
train a naive Bayes classi er to lter news tweets from \junk" tweets. News
tweets are then clustered based on TF-IDF. Liu et al. [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] also adopt a
twosteps approach to detect news clusters in a continuous stream of tweets. First,
they use a noise- ltering algorithm based both on a set of ltering rules and
on information credibility features [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. Then, they perform a clustering step
depending on a similarity metric containing named entities, verbs and common
nouns. These methods do not treat the problem of linking the collected tweets
to news articles, but they are interesting to perform a rst selection of relevant
tweets.
5.2
        </p>
      </sec>
      <sec id="sec-5-2">
        <title>Topic Models</title>
        <p>
          Many methods designed to link tweets and news are based on topic models. A
topic model takes a collection of documents as input and returns a set of topics
and their distribution within the collection (a document is considered as a mix
of several topics). Latent Dirichlet Allocation (LDA) [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ] is currently the most
commonly used topic model. Zhao et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] perform LDA on a collection of
news articles and an LDA-like algorithm on a collection of tweets and match
the two sets of topics based on the similarity of words distribution. Other works
develop algorithms to jointly learn the topics of tweets and news datasets [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ],
[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. Guo et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] design a more speci c latent variable model including features
like hashtags, named entities and temporal relations to link each tweet of their
dataset to the closest news article. The previous works perform well on aligning
tweets to related news articles, but they are based on static datasets and do not
tackle the dynamic nature of news events.
        </p>
        <p>
          In 2006 Blei and La erty proposed Dynamic Topic Model (DTM) [
          <xref ref-type="bibr" rid="ref30">30</xref>
          ] a
model able to discover the evolution of topics over time. Mele et al. adapted
this model to news events [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] and used clustering to link the documents covering
the same event across several news streams (news websites, RSS feed, Twitter
accounts of media outlets).
        </p>
        <p>Overall, several works address the task of linking tweets and news but they
rely either on static datasets or on stream of tweets from manually selected users.
We could not nd any approach questioning the representativity of the collected
tweets. Moreover, most approaches treat tweets and news articles only as text
documents and do not take their multimodal nature (urls, mentions of users,
pictures, videos) into account.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and future work</title>
      <p>In this paper, we present two types of methods to collect tweets related to news
events. The rst one is a query formulation approach, that consists in generating
potential queries related to an event and using the time repartition of collected
tweets to detect and reject wrong queries. The second one is an approach based
on clustering randomly collected events. We detailed the tests used to ensure
that the collected tweets have a close to random distribution and showed from
our rst results that neutral words clustered by co-occurrence tend to perform
better.</p>
      <p>In future works, we plan to develop a user interface that will enable the
labeling of numerous tweets, and give us a way of assessing the performance of
our search tool. We will also work on methods to dynamically link tweets to
news events, with a focus on processing tweets not only as text documents, but
rather as multimodal objects.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cage</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herve</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Viaud</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          :
          <article-title>The production of information in an online world: Is copy right</article-title>
          ? CEPR Discussion Paper #
          <volume>12066</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Morstatter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfe</surname>
            <given-names>er</given-names>
          </string-name>
          , J., Liu, H.:
          <article-title>When is it biased?: assessing the representativeness of twitter's streaming API</article-title>
          . In: WWW, Companion Volume.
          <article-title>(</article-title>
          <year>2014</year>
          )
          <volume>555</volume>
          {
          <fpage>556</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kothari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magdy</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darwish</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mourad</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taei</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Detecting Comments on News Articles in Microblogs</article-title>
          . In: ICWSM. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>W.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lim</surname>
            ,
            <given-names>E.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Comparing Twitter and Traditional Media Using Topic Models</article-title>
          . In: ECIR. (
          <year>2011</year>
          )
          <volume>338</volume>
          {
          <fpage>349</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diab</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Linking Tweets to News: A Framework to Enrich Short Text Data in Social Media</article-title>
          . In: ACL. (
          <year>2013</year>
          )
          <volume>239</volume>
          {
          <fpage>249</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Mele</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bahrainian</surname>
            ,
            <given-names>S.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Linking News across Multiple Streams for Timeliness Analysis</article-title>
          .
          <source>In: CIKM</source>
          . (
          <year>2017</year>
          )
          <volume>767</volume>
          {
          <fpage>776</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>John</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kambhampati</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : ET-LDA:
          <article-title>Joint Topic Modeling for Aligning Events and their Twitter Feedback</article-title>
          . (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hua</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ning</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramakrishnan</surname>
          </string-name>
          , N.:
          <article-title>Topical Analysis of Interactions Between News and Social Media</article-title>
          . In: AAAI. (
          <year>2016</year>
          )
          <volume>2964</volume>
          {
          <fpage>2971</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Becker</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naaman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gravano</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Automatic Identi cation and Presentation of Twitter Content for Planned Events</article-title>
          . In: ICWSM. (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tanev</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ehrmann</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piskorski</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zavarella</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Enhancing Event Descriptions through Twitter Mining</article-title>
          . In: ICWSM. (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ning</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muthiah</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tandon</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramakrishnan</surname>
          </string-name>
          , N.:
          <article-title>Uncovering News-Twitter Reciprocity via Interaction Patterns</article-title>
          . In: ASONAM. (
          <year>2015</year>
          ) 1{
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>El-Haddad</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfe</surname>
            <given-names>er</given-names>
          </string-name>
          , J.,
          <string-name>
            <surname>Stempeck</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Characterizing the life cycle of online news stories using social media reactions</article-title>
          .
          <source>In: CSCW</source>
          . (
          <year>2014</year>
          )
          <volume>211</volume>
          {
          <fpage>223</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Morstatter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfe</surname>
            <given-names>er</given-names>
          </string-name>
          , J., Liu,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Carley</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>M.: Is the Sample Good Enough? Comparing Data from Twitter's Streaming API with Twitter's Firehose</article-title>
          . In: ICWSM. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Callan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Should we use the sample? analyzing datasets sampled from twitter's stream api</article-title>
          .
          <source>ACM Trans. Web</source>
          <volume>9</volume>
          (
          <issue>3</issue>
          ) (
          <year>June 2015</year>
          )
          <volume>13</volume>
          :
          <fpage>1</fpage>
          {
          <fpage>13</fpage>
          :
          <fpage>23</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Diaz</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Craswell</surname>
          </string-name>
          , N.:
          <article-title>Query Expansion with Locally-Trained Word Embeddings</article-title>
          .
          <source>CoRR abs/1605</source>
          .07891 (May
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Roy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paul</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garain</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          :
          <article-title>Using Word Embeddings for Automatic Query Expansion</article-title>
          .
          <source>CoRR abs/1606</source>
          .07608 (
          <year>June 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nourbakhsh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            , R., Thomas,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andersony</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kociubay</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veddery</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pomervilley</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wudali</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martiny</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dupreyy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vachhery</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keenan</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Reuters Tracer: A Large Scale System of Detecting &amp; Verifying Real-Time News Events from Twitter</article-title>
          . In: CIKM. (
          <year>2016</year>
          )
          <volume>207</volume>
          {
          <fpage>216</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient Estimation of Word Representations in Vector Space</article-title>
          .
          <source>CoRR abs/1301</source>
          .3781 (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: NIPS</source>
          . (
          <year>2013</year>
          )
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Joseph</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Landwehr</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>M.: Two 1% s Don't Make a Whole: Comparing Simultaneous Samples from Twitter's Streaming API,</article-title>
          . In: SBP. (
          <year>2014</year>
          )
          <volume>75</volume>
          {
          <fpage>83</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Petrovic</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Osborne</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavrenko</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Streaming rst story detection with application to twitter</article-title>
          . In: HLT-NAACL. (
          <year>2010</year>
          )
          <volume>181</volume>
          {
          <fpage>189</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Atefeh</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khreich</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>A Survey of Techniques for Event Detection in Twitter</article-title>
          .
          <source>Computational Intelligence</source>
          <volume>31</volume>
          (
          <issue>1</issue>
          ) (
          <year>February 2015</year>
          )
          <volume>132</volume>
          {
          <fpage>164</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Aggarwal</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subbian</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Event detection in social streams</article-title>
          .
          <source>In: SIAM</source>
          . (
          <year>2012</year>
          )
          <volume>624</volume>
          {
          <fpage>635</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>B.S.</given-names>
          </string-name>
          :
          <article-title>Event Detection in Twitter</article-title>
          . In: ICWSM. (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yilmaz</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
          </string-name>
          , E.:
          <article-title>Dynamic Clustering of Streaming Short Documents</article-title>
          . In: KDD. (
          <year>2016</year>
          )
          <volume>995</volume>
          {
          <fpage>1004</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lei</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhuang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanratty</surname>
            , T., Han,
            <given-names>J</given-names>
          </string-name>
          .: TrioVecEvent:
          <article-title>Embedding-Based Online Local Event Detection in Geo-Tagged Tweet Streams</article-title>
          . In: KDD. (
          <year>2017</year>
          )
          <volume>595</volume>
          {
          <fpage>604</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Sankaranarayanan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samet</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teitler</surname>
            ,
            <given-names>B.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lieberman</surname>
            ,
            <given-names>M.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sperling</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>TwitterStand: news in tweets</article-title>
          . In: GIS. (
          <year>2009</year>
          )
          <volume>42</volume>
          {
          <fpage>51</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mendoza</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poblete</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Information credibility on twitter</article-title>
          .
          <source>In: WWW</source>
          . (
          <year>2011</year>
          )
          <volume>675</volume>
          {
          <fpage>684</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research 3(Jan)</source>
          (
          <year>2003</year>
          )
          <volume>993</volume>
          {
          <fpage>1022</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , J.D.:
          <article-title>Dynamic topic models</article-title>
          . In: ICML. (
          <year>2006</year>
          )
          <volume>113</volume>
          {
          <fpage>120</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>