<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detecting Filter Bubbles in Ongoing News Stories</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giang Binh Tran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eelco Herder</string-name>
          <email>herder@L3S.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>L3S Research Center</institution>
          ,
          <addr-line>Hannover</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we analyze di erences in perspective between timelines created by news agencies from various countries. By employing methods for date and headline selection, which have been extensively evaluated in previous work, we show several types of bias that exist in the media landscape. As users typically select only a small number of news sources to follow, they necessarily experience at least some 'tunnel vision', which is commonly associated with the so-called 'filter bubble'. By recognizing and emphasizing the peculiarities of the users' self-selected news sources, we can help them to break out of the bubble.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>One of the tasks of a journalist is to monitor, gather, curate and contextualize the
relevant information for the target audience. He needs to go through an enormous amount
of records with information of very diverse degrees of granularity, in order to put
information into context and tell his story from all significant angles, and, at the same time,
he needs to reduce the noise of irrelevant content.</p>
      <p>Many newspapers regularly publish manually created timelines, which allow the
reader to gain a quick overview over events that span a longer time period - such as the
Egypt revolution - and to answer questions such as: how and when did the event start?
What were the main consequences of the initial events? What happened to the main
protagonists in the event?</p>
      <p>
        Though convenient for the reader, creating a timeline is expensive, as it requires
substantial expertise. Moreover, creating a timeline is a very subjective task and
therefore huge di erences can be found in expert-generated timelines, even if they are on
the same topic [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In addition, there exist di erent models of reporting, in the
Western world varying from the Anglo-Saxon tradition of just reporting the facts, via the
Northern-European model of combining news reporting with opinions, to the
‘polarized Mediterranean’ type of reporting [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Moreover, it is a known fact that newspapers
in di erent countries provide di erent perspectives and points of focus, reflecting
differences in national ideologies, priorities and opinions [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        On the Internet, the filter bubble, as coined and popularized by Eli Pariser [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], is a
well-known phenomenon that explains why personalization leads users to mainly
encounter products, news articles and other types of content that match the users’ own
interests and viewpoints. The tunnel vision that the filter bubble is said to create, is not
just a recent online phenomenon, but inherent to journalism: editors need to be selective
and therefore necessarily focus on matters that are of highest interest to their readers.
      </p>
      <p>These are some of the reasons why summaries of events, such as timelines, di er
wildly, depending on the nationality, audience, political viewpoints and other sources of
subjectivity associated with a newspaper. At the same time, readers are known to select
a small number of newspapers - or even just one - that best match their personality,
viewpoints and interests. As argued by Paul Resnick in his keynote at UMAP 20111,
emphasizing di erences in perspective is an important instrument for allowing users to
break out of their ‘filter bubbles’.</p>
      <p>In this paper, we investigate methods for automatically recognizing and
emphasizing di erences between various timelines, and evaluate them by comparing timelines
from well-known news agencies and newspapers. These methods are based on our
previous work on date and headline selection for news event summarization; by employing
these methods and the Shannon Diversity Index, we are able to recognize dates that were
considered important in one timeline but not in others, and di erences in news coverage
on dates that were considered important in all timelines.</p>
      <p>The work is carried out in the context of the EUMSSI2 project, in which
crossmodal analysis techniques are developed for analysing news articles, videos and audio
reports. This will allow journalists - and media consumers - to relate these messages
with one another and to understand the underlying events.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Many studies specific to timeline summarization, such as [
        <xref ref-type="bibr" rid="ref10 ref7 ref8">7, 10, 8</xref>
        ], focus on the
extraction of salient sentences or headlines for generating the textual content of timelines.
They assume either that the dates are given in advance or they use simple measures such
as burstiness for date selection [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Prior approaches dedicated specifically to date selection include work by Tran et al.
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Kessler et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. They use supervised methods that score dates independently
of each other, making use of frequency-based, temporal and topical features that are
extracted from a corpus of event-related newspaper articles. We, however, score dates
jointly, making use of interactions between dates in a graphical model.
      </p>
      <p>
        Additionally, by using the Shannon diversity index [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] on the top of the graphical
model, our approach can highlight dates on which important events happened, but that
are likely to be ignored by many news agencies. This aspect has not been considered so
far in the previous work.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>In order to recognize and emphasize di erences between timelines, it is important to
first find events and corresponding dates that are ‘subjective’, in other words dates that
are not considered important by all news agencies.</p>
      <p>
        Our approach for subjective date detection makes use of a corpus of news articles
(more details in Section 4) and consists of two steps. First, we use a random walk model,
which we proposed in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], to rank the dates based on their importance with regard to
their impact on the future events3. After that, we re-rank the top selected dates based on
1 https://presnick.wordpress.com/2011/07/17/personalized-filters-yes-bubbles-no/
2 http://www.eumssi.eu/
3 A demo of the WikiTimes system for automatic timeline creation is available at http://
wikitimes.l3s.de/
their Shannon Diversity Index (SDI) scores. SDI is widely used in ecology and biology
to measure the diversity of species in a community [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. We use the SDI to express
the rarity and commonness of events, as reported by di erent news agencies. When an
(important) event is commonly reported by many news agencies, it is less subjective
and thus less likely to be filtered out. For this reason, our approach here is to give a high
rank to the date that has a low SDI score.
      </p>
      <p>Formally, in the first step, we build a date reference graph, which is a fully directed
graph G = (V; E), where V is the set of dates mentioned in any text in corpus C,
including publication dates. The edges E = fe(di; d j)g indicate that at least one text
published on di refers to the date d j.</p>
      <p>We represent each link between events as a multi-value tuple</p>
      <p>e(di; d j) = (Mi j; f req(di; d j); Itemporal(di; d j); Itopical(di; d j))
to integrate di erent measures of date importance. The first value, Mi j = N1 expresses
the prior transition probability between 2 dates where N = jVj. The other values express
the strength of the connection between di and d j, modeled by the following aspects:
frequency ( f req), temporal influence (Itemporal) and topical influence (Itopical). Our random
walk model uses these perspectives to rank the collection of dates. The intuition is that
when a date d j is referred to from either a past or future news article (published on di),
it is likely involved in the events that are reported in that article.</p>
      <p>In the second step, we compute the SDI for the top-K ranked dates, based on the
distribution of news articles published on those dates. For example, a date that contains
only news articles from the BBC is considered less diverse than a date that contains
news articles from several large news agencies over the world. The computation of SDI
is sketched as follows:</p>
      <p>S DI(di) =</p>
      <p>R
X pi ln pi
i=1
where R is the number of news agencies from which we collected data and pi is the
proportion of news articles from news agencies pi. That measure quantifies uncertainty in
predicting the agency identity of a news article that is taken randomly from the dataset,
hence, suggesting how subjective (or non-diversified) a date is.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experiment</title>
      <p>We present our preliminary results of two experiments on detecting subjective dates
in the Crisis data4 dataset, which contains ground truth (GT) timelines (written by
professional journalists) as well as a corpus of around 12K news articles that cover
events that happened in the context of four news stories: Egypt Revolution, Syria War,
Libya War and Yemen Crisis. The dataset is suitable for our purpose for the following
reasons: (1) it is a heterogeneous dataset that contains news articles and expert timeline
summaries from 25 well-known news agencies and; (2) it covers long-term stories that
have been happening since 2011, making the date selection problem non-trivial for any
system.
4 The Crisis data is currently available at http://l3s.de/~gtran/timeline/
4.1</p>
      <sec id="sec-4-1">
        <title>Experiment 1: Subjective Dates Selected by One News Agency</title>
        <p>In this experiment, we aim to detect dates that have been included in the timeline of
only one news agency (and not in other timelines). We use a cut-o K = 50 to select
the 50 most important dates, using our random walk model, and rank them by their SDI
score. We then compare these selected dates with those in our ground-truth timelines.
The performance of the subjective date detection process is measured as the proportion
of the top-10 dates (this can be extended to a larger number) that are included in exactly
one timeline.</p>
        <p>
          The result is described in Table 1, which shows that a significant number of
important dates and their events are reported by only one news agency. We took a closer look
into those events (see examples in Table 2) and found evidence that the events that are
included in only one timeline are typically not about the main theme of the ongoing
news stories, but about related aspects, such as business and human rights. As we
discussed in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], timelines often contain substories or side-paths that involve major actors
of the main story. Due to space limitations, journalists can only incorporate a limited
number of substories and the decision which stories to incorporate is often subjective.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Experiment 2: Important Dates that Are Not Covered</title>
        <p>In this experiment, we aim to find out how many important dates are not included in
any of the news agencies’ timelines. Similar to the previous experiment, we picked the
top-50 important dates from the results of our random walk algorithm and calculated
the proportion of dates that are not mentioned in any timeline. The results are shown in
Table 3. The high number of ‘missed dates’ does not mean, however, that the ground
truth timelines are of poor quality: as explained earlier, journalists need to be selective
when creating timelines. The SDI values also confirm that journalists usually do a good
job in selecting the dates to be included: for all stories, except for Libya, the average
SDI of dates included in at least one timeline is higher than the average SDI of dates
that did not make it into any of the timelines. In other words, dates that are included in
at least one timeline are typically covered by more news agencies than dates that do not
appear in any timeline.</p>
        <p>To better understand which events were left out of the timelines, we analyzed several
of those dates with a high SDI score but that were not included. Some example events
are shown in Table 4. Except perhaps for the Libya article, these are events that readers
may be interested in, but that they will not be able to find in the timelines of the news
agencies that they are subscribed to.
In this paper, we have shown that it is possible to automatically find dates and
corresponding events in news stories that are likely to be covered in only one or a subset of
manually generated timelines. Further, we have shown that there are important dates
that are not incorporated in any timeline.</p>
        <p>As we argued earlier, it is unavoidable that timelines are selective and, to a certain
extent, subjective. Still, it would be very desirable to make readers aware of the
peculiarities of the timeline that they currently inspect. This can be achieved by highlighting
dates with large di erences in coverage between news agencies or timelines, by adding
links to other timelines for dates that are not incorporated in the current timeline, or,
alternatively, by constructing an annotated ‘timeline of timelines’ for those who wish to
contrast the bias in their personal selection of news sources with news sources that they
usually do not visit.</p>
        <p>Conversely, journalists will be able to create better balanced - or, alternatively, even
more argumentative - timelines with feedback on their current selection of dates and
events. Even though - as we have shown in previous work - it is possible to
automatically construct timelines by selecting the most relevant dates and headlines, still manual
processing and editing would be needed to enhance the communicative qualities of the
timeline, and to adapt it to the needs of the readers.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>The work was partially funded by the European Commission for the FP7 project EUMSSI
(611057)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>H. L.</given-names>
            <surname>Chieu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y. K.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Query based event extraction along a timeline</article-title>
          .
          <source>In Proceedings of SIGIR'04</source>
          , pages
          <fpage>425</fpage>
          -
          <lpage>432</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>F.</given-names>
            <surname>Esser</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Umbricht</surname>
          </string-name>
          .
          <article-title>Competing models of journalism? political a airs coverage in us, british, german, swiss, french and italian newspapers</article-title>
          .
          <source>Journalism</source>
          ,
          <volume>14</volume>
          (
          <issue>8</issue>
          ):
          <fpage>989</fpage>
          -
          <lpage>1007</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>R.</given-names>
            <surname>Kessler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tannier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hagege</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Moriceau</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Bittar</surname>
          </string-name>
          .
          <article-title>Finding salient dates for building thematic timelines</article-title>
          .
          <source>In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics: Long Papers-Volume</source>
          <volume>1</volume>
          , pages
          <fpage>730</fpage>
          -
          <lpage>739</lpage>
          . Association for Computational Linguistics,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>E. Pariser.</surname>
          </string-name>
          <article-title>The filter bubble: What the Internet is hiding from you</article-title>
          .
          <source>Penguin UK</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>S.</given-names>
            <surname>Seo</surname>
          </string-name>
          .
          <article-title>Hallidayean transitivity analysis: The battle for tripoli in the contrasting headlines of two national newspapers</article-title>
          .
          <source>Discourse &amp; Society</source>
          ,
          <volume>24</volume>
          (
          <issue>6</issue>
          ):
          <fpage>774</fpage>
          -
          <lpage>791</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Shannon</surname>
          </string-name>
          .
          <article-title>A mathematical theory of communication</article-title>
          .
          <source>SIGMOBILE Mob. Comput. Commun. Rev.</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>3</fpage>
          -
          <lpage>55</lpage>
          , Jan.
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Swan</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Allan</surname>
          </string-name>
          .
          <article-title>Timemine: visualizing automatically constructed timelines</article-title>
          .
          <source>In SIGIR, page 393</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alrifai</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. Quoc</given-names>
            <surname>Nguyen</surname>
          </string-name>
          .
          <article-title>Predicting relevant news events for timeline summaries</article-title>
          .
          <source>In WWW</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>G.</given-names>
            <surname>Tran</surname>
          </string-name>
          , E. Herder, and
          <string-name>
            <given-names>K.</given-names>
            <surname>Markert</surname>
          </string-name>
          .
          <article-title>Joint graphical models for date selection in timeline summarization</article-title>
          .
          <source>In Proceedings of ACL</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>R.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Otterbacher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Y. Zhang.</surname>
          </string-name>
          <article-title>Evolutionary timeline summarization: a balanced optimization framework via iterative substitution</article-title>
          .
          <source>In Proceedings of SIGIR'11</source>
          , pages
          <fpage>745</fpage>
          -
          <lpage>754</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>