<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>What Snippets Say About Pages (Abstract)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>T. Demeester</string-name>
          <email>tdmeeste@ugent.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Nguyen</string-name>
          <email>d.nguyen@utwente.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Trieschnigg</string-name>
          <email>d.trieschnigg@utwente.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C. Develder</string-name>
          <email>cdvelder@ugent.be</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Hiemstra</string-name>
          <email>d.hiemstra@utwente.nl</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ghent University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Twente</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We summarize ndings from [1]. What is the likelihood that a Web page is considered relevant to a query, given the relevance assessment of the corresponding snippet? Using a new Federated Web Search test collection that contains search results from over a hundred search engines on the internet, we are able to investigate such research questions from a global perspective. Our test collection covers the main Web search engines like Google, Yahoo!, and Bing, as well as smaller search engines dedicated to multimedia, shopping, etc., and as such re ects a realistic Web environment. Using a large set of relevance assessments, we are able to investigate the connection between snippet quality and page relevance. The dataset is strongly heterogeneous, and care is required when comparing resources. To this end, a number of probabilistic variables, based on snippet and page relevance, are introduced and discussed.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>Finding our way around among the vast quantities of data
on the Web would be unthinkable without the use of Web
search engines. Apart from a limited number of very large
search engines that constantly crawl the Web for publicly
available data, a large amount of smaller and more focused
search engines exist, specialized in speci c information goals
or data types (e.g., online shopping, news, multimedia,
social media). In order to promote research on Federated Web
Search, we created a large dataset containing sampled
results from 108 search engines on the internet, and
containing relevance judgments for the top 10 results (both
snippets and pages) from all of these resources for 50 test topics
(from the TREC 2010 Web Track). The relevance
judgements are particularly interesting for analysis, partly
because they originate from very diverse collections (both in
size and in scope, whereby the relevance judgments are done
in a generic way), and partly because we not only judged the
result pages, but also, independently, the original snippets.
Our analysis deals with ranked result lists from diverse
retrieval algorithms, and with snippets from various snippet
generation strategies, as they are currently in use on the
Web.</p>
      <p>
        This abstract is based on [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which has the following scope.
      </p>
      <p>
        First, after an overview of related work, the relevance
judgments for the new dataset are discussed at length, with
emphasis on the assessors' consistency. Second, a number
of potential di culties in Federated Web Search and
especially in the evaluation of relevance are discussed, related
to the heterogeneous character of the resources. Finally,
a probabilistic analysis of the relationship between the
indicative snippet relevance and the actual page relevance is
presented (where by `page' we denote a result item like a
web page, a video, scienti c paper... as returned by the
included search engines). In a further contribution [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], it is
shown that the information carried by an average snippet
can be used to make a reasonable prediction of the
relevance of the result page itself. Within the limits of this
abstract, we will primarily focus on the question of why the
user's snippet-based prior estimation of the page relevance
is of paramount importance for the overall performance of
the search service. Using the relevance judgments for the
dataset presented in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], the relevant concepts are illustrated
for the speci c case of large general web search engines.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. SNIPPET VS. PAGE RELEVANCE</title>
      <p>The intuition behind this paper is simple: a search engine
can only exploit the full potential of its retrieval algorithm if
the result snippets re ect the relevance of the corresponding
pages as well as possible. This means that a highly relevant
result should be presented to the user by a very promising
snippet, and a less relevant result page by a less interesting
snippet. If there is a mismatch between what the user
estimates from a result snippet and the actual result page, the
overall performance of the system degrades.</p>
      <p>For a more formal analysis, we introduce the snippet
relevance variable S, and the page relevance variable P. As for
the speci c relevance levels, the snippet relevance S ranges
from No, over Unlikely and Maybe, to Sure, indicating how
likely the assessor estimates the result page behind the
snippet to be relevant. The levels for P, the page relevance, are
Non, Rel (containing minimal relevant information), HRel
(highly relevant), Key (worthy of being a top result), and
Nav (for navigational queries). In this paper we will either
indicate the considered relevance level explicitly, such as S
= Sure (i.e., considering only snippets with the label Sure),
or de ne binary relevance levels, such as P HRel
(indicating page relevance levels of HRel, Key, or Nav).</p>
      <p>General Web search
Multimedia
News
Shopping
Encyclopedia/Dict
Books
Blogs
Retrieval systems are typically being evaluated based on the
probability of relevance of the result page, written P(P). If
however the access to that page also depends on the user's
estimate of a snippet, the actual measure to consider should
be P(P;S), the mutual probability of relevance for both
the snippet and the page. Note that it can be written as
P(S)P(PjS), in which P(PjS) is the conditional probability
of the page label, given the snippet label. Studying P(PjS)
is especially instructive, for instance to nd out how often a
relevant page remains hidden behind a non-convincing
snippet.</p>
      <p>For several resource categories, table 1 gives empirical
estimates of such probabilities for binary page relevance P HRel,
based on our relevance judgements. Comparing P(PjS) for
the snippet labels Maybe and Sure shows that a relatively
large amount of HRel pages are behind snippets which were
judged only Maybe, especially for the general search engines.
This shows that often a HRel page's snippet cannot convince
the user that the page is indeed highly relevant. We also
observe that for the snippet label S=Sure, e.g., the News
resources display a relatively high P(PjS), against a very low
P(P;S). In other words, these resources returned only very
few relevant results for our test topics, but if a snippet was
found relevant, 4 out of 10 times it points to one of those
few relevant results.</p>
      <p>As the test topics are best suited for the general Web search
engines, we can explicitly compare the performance of four of
the largest general Web search engines in our collection, i.e.,
Google, Yahoo!, Bing, and Baidu, as well as Mamma.com,
which is actually a metasearch engine. Table 2 presents the
results. It appears that for the snippet label S=Sure and
two page relevance levels (P HRel and P Key), P(P;S) is
consistently lower than P(P), which is actually the averaged
precision@10 of page relevance, and does not take into
account the fact that the snippet is not always as promising as
the page is relevant. The metasearch engine outperforms the
others, as it aggregates results from a number of resources,
such as Google, Yahoo!, and Bing. We want to stress that
the considered test topics are still no representative
collection of, for example, popular Web queries, and therefore we
cannot draw any further conclusions about these search
engines beyond the scope of our test collection. Yet, here is
another example of how the table might be interpreted, with
that in mind. Considering only Key results, we could
compare Yahoo! and Bing. Yahoo! seems to score higher for
all reported parameters, so either Bing's collection contains
a smaller number of relevant results, or Yahoo!'s retrieval
algorithms are better tuned for our topics. The lower value
of P(PjS) for Bing shows that it has a slightly increased
chance that the page for a promising snippet appears less
relevant. However, the ratio of P(P;S) and P(P) is higher
for Bing than for Yahoo!, indicating that for Yahoo!, its own
recall on Key pages will be decreased more due to the
quality of the snippets, than for Bing. In fact, we found that
P(S=SurejP Key) is 79% for Yahoo!, but 91% for Bing.</p>
    </sec>
    <sec id="sec-3">
      <title>3. CONCLUSIONS</title>
      <p>Analyzing the relationship between the relevance of
snippets from a large amount of on-line search engines and the
relevance of the corresponding result pages, clearly shows
that in the evaluation of and comparison between di erent
resources, the snippets cannot be left out.</p>
    </sec>
    <sec id="sec-4">
      <title>4. ACKNOWLEDGMENTS</title>
      <p>This research was partly supported by the Netherlands
Organization for Scienti c Research, NWO, grants 639.022.809
and 640.005.002, and partly by iMinds in Flanders.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Trieschnigg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Develder</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>What Snippets Say about Pages in Federated Web Search</article-title>
          . In AIRS,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Trieschnigg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Develder</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>Snippet-Based Relevance Predictions for Federated Web Search</article-title>
          . In ECIR,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Demeester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Trieschnigg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <article-title>Federated Search in the Wild: the Combined Power of over a Hundred Search Engines</article-title>
          . In CIKM,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>