<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Timeline as Information Retrieval and Ranking Unit in News Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adam Jatowt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Innsbruck</institution>
          ,
          <addr-line>Innrain 15, Innsbruck 6020</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>News articles are one of the most often read online documents, and the amount of news stories generated daily is quite large. News search engines are then often used to retrieve relevant news articles. The understanding of the returned results can be however impaired when the returned articles are about events being parts of complex or long stories. In this paper we discuss the dificulties resulting from the missing context and information complexity that users searching in temporal news collections may face. To support users in their search and learning we propose the concept of timeline as a retrieval and ranking unit of news search. We believe that automatically generating and ranking timelines could become an efective mechanism to facilitate search and browsing in large news collections.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;news collections</kwd>
        <kwd>news search</kwd>
        <kwd>timeline summarization</kwd>
        <kwd>news archives</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>chronic text collections such as collections of web pages
or Wikipedia articles. In this context, we believe that the
News are one of the most commonly read types of online unique temporal characteristics of news archives
necesdocuments nowadays. They matter much to all users sitate novel methods for efective information retrieval.
who want to understand key events in the world or in Professional users typically know what they wish to find
their localities. In recent years, large amounts of news when searching in or interacting with news collections
articles have been published, many of which are about and they also have suficient skills and knowledge of the
complex and long events or stories. Thus specialized collections (or their covered time periods) to efectively
news search engines are available to users helping them retrieve information. On the other hand, general users
to find relevant information. may lack necessary knowledge and expertise to be able</p>
      <p>IR technologies have been adapted to many diverse to perform successful searches in large news collections.
domains including web search, legal document search, Searchers in large news collections face to some extent
code retrieval, social media search, etc. [1]. We think that similar issues as users who search in collections
connovel and efective developments in the field of accessing taining documents from unknown and complex domains
temporal news collections such as news archives are also (e.g., medicine or law). The lack of knowledge causes
difneeded due to large amounts of news generated these ifculties in understanding and interpreting search results,
days and their complex characteristics. News articles updating further queries and continuing search sessions.
are a highly temporal type of documents in which time This is especially a problem when the searched events
ocsignals (whether in the form of timestamps or embedded curred in more distant past. Imagine a user who wishes to
temporal expressions) are of key importance, and news learn about the progress of Iraq War - the event that took
tend to be often understood chronologically or in causal place over 10 years ago. Even if the user managed to find
order. While text indexing or query suggestion meth- all relevant news articles, she or he will have problem
ods have been already studied in the context of temporal to fully understand them, arrange them into
meaningsearch such as search in web archives or in news archives ful groups as well as continue the searching process by
[2, 3, 4], still there is need for research on efective re- refining subsequent queries. First, this is because the
trieval and ranking approaches. Currently, the usual user most likely does not know the context of the times
retrieval approach seems to be basically one based on when the war happened and the relations between its
applying the same access methods as for traditional syn- individual sub-events. Second, the news articles returned
by a typical search engine would be probably ranked just
DESIRES 2021 – 2nd International Conference on Design of by their relevance despite that the intuitive chronological
Experimental Search &amp; Information REtrieval Systems, September ordering is useful and intuitive for the user. The lack of
15–18, 2021, Padua, Italy chronological or causal order would prevent users from
" adam.jatowt@uibk.ac.at (A. Jatowt) understanding the interrelations between events and the
~ h00tt0p0s-:0//0d0s1--i7n2f3o5rm-0a6t6i5k.(uAib.kJa.atco.watt/)(A. Jatowt) overall progress of the underlying story. Next, the user
© 2021 Copyright for this paper by its authors. Use permitted under Creative would also not be able to find important articles as the
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g CCoEmmUoRns LWiceonsrekAstthribouptionP4r.0oIncteerenadtiionnagl s(CC(CBYE4U.0)R.-WS.org) retrieval method probably does not utilize any notion of
event importance in an explicit way. Hence, less impor- conceptualizing or refining queries and let the user better
tant news articles could be ranked highly with some of understand what she or he would like to actually retrieve
them perhaps being returned at the top of the ranked or further explore.
list, efectively preventing the user from understanding
the key events of the war. Note that manually arranging
articles into meaningful sequences would likely not be 3. Ranked Timelines as Output
realistically feasible. This is because of the presumed
lack of knowledge that an average user has about events
in arbitrary time periods, and due to potentially large
numbers of articles that can be returned, especially for
longer events or stories.</p>
      <p>Incorporating the notion of event importance into the
ranking mechanism could let users receive news articles
ordered not just by their relevance but also by the
quantiifed importance levels of the described events. However,
still this would not fully help users obtain large picture of
all the major events related to their query, as the returned
events would not be ordered chronologically.
Furthermore, the set of ranked results could contain articles
describing events that are parts of diferent stories or
that are related to diverse aspects of the same story. This
would likely happen for ambiguous queries such as
location names or names of popular entities that took parts in
diverse events (e.g., a name of a country’s president), or
queries about complex and long-evolving events having
diverse aspects (e.g., war in Syria).</p>
      <p>TimeLine Summarization (TLS) research helps alleviate
the problem of redundancy and complexity inherent in
news article collections, thereby supporting users to
better understand the news landscape. Usually a timeline is
in the form of short descriptions of major events which
are presented in chronological order, and it may also
reflect causal relations. In traditional setting, the TLS
approaches [5] work on a homogeneous type of datasets
(i.e., collections of documents about the same event or
about events of the same story) and produce just a single
timeline as output. For heterogeneous document
collections such as the set of relevant search results obtained in
news search, one would need to apply the generalization
of TLS.</p>
      <p>Recently, Multi-TimeLine Summarization (MTLS) [6]
has extended the typical TLS settings by allowing
heterogeneous document collections as an input and by
generating the output in the form of multiple timelines.</p>
      <p>Outputting multiple timelines as a retrieval result in
news search would be suitable to efectively organize
and present search results, especially when the query is
2. Timeline for Providing Missing ambiguous, relates to complex or diverse stories, or the
related news have multiple diferent aspects. The
timeContext and Structuring Search lines would summarize diferent relevant stories or could
Results reflect diferent aspects of the same story (e.g., timeline of
the military actions, timeline of the economical aspects of
We argue that presenting search results in the form of ten the Iraq war, timeline of the societal responses towards
blue links ordered by the traditional notion of keyword the war, etc.). These timelines could be ranked based
relevance would not sufice in many cases of news search. on their relevance degrees to the user query (e.g.,
comEspecially, in cases when (a) the sought news are parts of puted as the sum of relevance scores of their constituent
longer ongoing stories, (b) the searched events happened events) or on their importance (e.g., computed as the sum
in more distant past, (d) the issued queries are ambiguous, of importance scores of the constituent events, or simply
or (c) when the user lacks necessary knowledge of event’s bound to the number of events, i.e., the timeline length).
context, more efective solutions are required. We think In short, rather than ranking individual news articles, the
that arranging the returned news results in the form of search engine would construct on-the-fly and rank
timetimelines that summarize the key events and reflect their lines to be presented as ranked results to users. These
chronological and/or causal order would be a more user- timelines would provide missing context to individual
friendly way towards efective search in longitudinal events and help structure the entire news landscape that
news article collections. In the above-mentioned example is relevant to user query. Finally, the timelines used as
of the search intent of Iraq War the user could receive retrieval units could be then further expanded based on
automatically generated timelines, instead of the usual user clicks to let the users view the detailed articles. This
ranked list of news articles. Such timelines would be then style of result presentation bears resemblance to search
the first step towards combating the lack of context on result clustering by their temporal aspects and ranking
the user side which prevents successful search experience. in Web search as proposed by Alonso et al. [7].
After checking the timelines the user should be better Regarding the actual algorithm that could be used for
informed of the event landscape (e.g. the progress of MTLS task, Yi et al. [6] have proposed a two stage
Afinthe war represented by the sequence of its major events) ity Propagation approach to automatically generate the
in order to perform further search. This would support set of timelines from a heterogeneous news collection.
An advantage of that approach is that the number of
the generated timelines does not need to be known or
set beforehand, as it is dynamically obtained based on
the underlying document collection. This flexibility is
important in search scenarios as users can input
arbitrary queries that could necessitate varying numbers of
constructed timelines.
search episodes).</p>
    </sec>
    <sec id="sec-2">
      <title>Acknowledgments</title>
      <p>We thank anonymous reviewers for their useful
comments and suggestions.
[1] W. B. Croft, D. Metzler, T. Strohman, Search
engines: Information retrieval in practice, volume 520,
In this position paper we propose considering timeline Addison-Wesley Reading, 2010.
as an atomic retrieval1, ranking and presentation unit [2] K. Berberich, S. Bedathur, T. Neumann, G. Weikum,
for news search. We think that incorporating Multiple A time machine for text search, in: Proceedings of
Timeline Summarization methods as well as adding the the 30th annual international ACM SIGIR
confertimeline ranking component would provide useful alter- ence on Research and development in information
native to users besides the usual retrieval method that retrieval, 2007, pp. 519–526.
centers on fine-grained retrieval units such as individual [3] N. K. Tran, A. Ceroni, N. Kanhabua, C. Niederée,
documents. This novel approach should be especially use- Back to the past: Supporting interpretations of
forful in cases of searching in longitudinal news collections. gotten stories by time-aware re-contextualization,
Users could, for example, benefit from the interwoven in: Proceedings of the Eighth ACM International
combination of providing timelines as search results and Conference on Web Search and Data Mining, 2015,
of outputting the ranked news articles in a traditional pp. 339–348.
way. For example, upon understanding the news land- [4] Y. Zhang, A. Jatowt, S. Bhowmick, K. Tanaka,
Omscape based on the constructed timelines, the users could nia mutantur, nihil interit: Connecting past with
switch to the usual document relevance-based search present by finding corresponding terms across time,
in order to locate particular articles and satisfy detailed in: Proceedings of the 53rd Annual Meeting of the
search needs. Timelines would then not only provide the Association for Computational Linguistics and the
"global picture" useful for conducting search, but could 7th International Joint Conference on Natural
Lanalso let users narrow down the results (e.g., by selecting guage Processing (Volume 1: Long Papers), 2015, pp.
a sub-part of some timeline for subsequent search). 645–655.</p>
      <p>There are many open problems that need to be ap- [5] D. G. Ghalandari, G. Ifrim, Examining the
stateproached for using timelines as retrieval and ranking of-the-art in news timeline summarization, arXiv
units including: (a) efective and eficient generation of preprint arXiv:2005.10107 (2020).
timelines for arbitrary user queries, (b) timeline labelling2 [6] Y. Yu, A. Jatowt, A. Doucet, K. Sugiyama,
and presentation3 for easy and quick result understand- M. Yoshikawa, Multi-timeline summarization (mtls):
ing, as well as (c) the actual ranking mechanism for ar- Improving timeline summarization by generating
ranging the generated timelines. multiple summaries, in: Proceedings of the Joint</p>
      <p>Finally, worth investigation are user interaction mod- Conference of the 59th Annual Meeting of the
Assoels and methods with search results composed of time- ciation for Computational Linguistics and the 11th
lines. This encompasses timeline zooming in and out, International Joint Conference on Natural Language
expanding timeline events to read their detailed articles, Processing, 2021.
or switching between timeline search results and usual [7] O. Alonso, M. Gertz, R. Baeza-Yates, Clustering and
search results. One can also imagine the possibility of exploring search results using timeline constructions,
users updating/constructing their own timelines (e.g., by in: Proceedings of the 18th ACM conference on
Inremoving irrelevant events/articles or adding relevant formation and knowledge management, 2009, pp.
ones, perhaps the ones discovered through subsequent 97–106.</p>
      <p>1Actually, to be correct, in the context of our proposal, the term
"retrieval" should be substituted by "construction" or "generation"</p>
      <p>2For fast overview of the ranked timelines, each timeline could
have its generated label or short description</p>
      <p>3Some of the constructed timelines may also contain shared
events, the explicit indication of which could be perhaps added for
more efective result presentation</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>Research Directions</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>