<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of WebCLEF 2007</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Valentin Jijkoun Maarten de Rijke ISLA, University of Amsterdam</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the WebCLEF 2007 task. The task definition-which goes beyond traditional navigational queries and is concerned with undirected information search goals-combines insights gained at previous editions of WebCLEF and of the WiQA pilot that was run at CLEF 2006. We detail the task, the assessment procedure and the results achieved by the participants.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
      <p>Task model
Our hypothetical user is a knowledgable person, perhaps even an an expert, writing a survey
article on a specicfi topic with a clear goal and audience, for example, a Wikipedia article, or a
state of the art survey, or an article in a scientific journal. She needs to locate items of information
to be included in the article and wants to use an automatic system to help with this. The user
does not have immediate access to offline libraries and only uses online sources.</p>
      <p>The user formulates her information need (the topic) by specifying:
• a short topic title (e.g., the title of the survey article),
• a free text description of the goals and the intended audience of the article,
• a list of languages in which the user is willing to accept the found information,
• an optional list of Google retrieval queries that can be used to locate the relevant information;
the queries may use site restrictions (see examples below) to express the user’s preferences.
Here’s an example of an information need:
• topic title: Significance testing
• description: I want to write a survey (about 10 screen pages) for undergraduate students on
statistical signicfiance testing, with an overview of the ideas, common methods and critiques.</p>
      <p>
        I will assume some basic knowledge of statistics.
• language(s): English
• known source(s): http://en.wikipedia.org/wiki/Statistical_hypothesis_testing ;
http://en.wikipedia.org/wiki/Statistical_significance
• retrieval queries: signicfiance testing; site:mathworld.wolfram.com signicfiance testing;
significance testing pdf; significance testing site:en.wikipedia.org
Defined in this way, the task model corresponds to addressing undirected informational search
goals, that are reported to account for over 23% of web queries [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Each participating team was asked to develop several topics and subsequently assess responses
of all participating systems for the created topics and
1.2</p>
      <p>Data collection
In order to keep the idealized task as close as possible to the real-world scenario (i.e., there are
many relevant documents) but still tractable (i.e., the size of the collection is manageable), our
collection is defined per topic. Specicfially, for each topic, the subcollection for the topic contains
the following set of documents along with their URLs:
• all “known” sources speciefid for the topic;
• the top 1000 (or less, depending at the actual availability) hits from Google for each of the
retrieval queries speciefid in the topic, or for the topic title if the queries are not specified;
• for each online document included in the collection, its URL, the original content retrieved
from the URL and the plain text conversion of the content are provided. The plain text
conversion is only available for HTML, PDF and Postscript documents. For each document,
the subcollection also provides its origin: which query or queries were used to locate it and
at which rank(s) in the Google result list it was found.
1.3</p>
      <p>System response
For each topic description, a response of an automatic system consists of a ranked list of plain
text snippets extracted from the sub-collection of the topic. Each snippet should indicate what
document of the sub-collection it comes from.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Assessment</title>
      <p>
        In order to comply with the task model, the manual assessment of the responses of the systems
was done by the topic creators. The assessment procedure was somewhat similar to assessing
answers to OTHER questions at TREC 2006 Question Answering task [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The assessment was to be blind. For a given topic, all responses of all system were pooled into
anonymized sequence of text segments. To limit the amount of required assessments, for each topic
only first 7,000 characters of each response were included (according to the ranking of the snippets
in the response). The cut-off point 7,000 was chosen so that for at least half of the submitted runs
the length of the responses was at least 7,000 for all topics. For the pool created in this way for
each topic, the assessor was asked to make a list of nuggets, atomic facts, that, according to the
assessor, should be included in the article for the topic. A nugget may be to character spans in
the responses, so that all spans linked to one nugget express this atomic fact. Different character
spans in one snippet in the response may be linked to more than one nugget. The assessors used
a GUI to mark character spans in the responses and link each span to the nugget it expresses (if
any). Assessors could also mark character spans as “known” if they expressed fact relevant for
the topic but alredy present in one of the known sources.</p>
      <p>Figure 1 shows the assessment interface for the topic “iPhone opinions”. Snippets (left bottom
of the figure) are separated by grey horisontal lines. Note that only part of the rfist snippet is
marked as relevant and the second snippet contain two marked spans that are linked to distinct
nuggets. This example illustrates the many-to-many relation between nuggets and character spans
in the response.</p>
      <p>
        Similar to INEX [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and to some tasks at TREC (i.e., the 2006 Expert Finding task [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ])
assessment was carried out by the topic developer, i.e., by the participants themselves.
      </p>
      <p>Table 1 gives the statistics for the 30 test topics and for the assessments of the topics.1
3</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation measures</title>
      <p>The evaluation measures for the task are based on standard precision and recall. For a given
response R (a ranked list of text snippets) of a system S for a topic T we define:
1Full definition of the test topics is available from http://ilps.science.uva.nl/WebCLEF/WebCLEF2007/Topics.
#o a
n p
k s
l p
ta ip
to sn
#o u
n o
k s
s
e
g
a
u
g
r
o
s
s
e
s
s</p>
      <p>T
I
,
PT ,IT ,IT
, T T</p>
      <p>E E
A</p>
      <p>B</p>
      <p>B B
#rak an 16 4 2 1 1
n N ,P ,P ,E ,E ,E ,N ,N ,N ,N ,N ,N ,N
a N N ,E N N N N N N N L L L L L N N N N N N N N N N N N N N
L E N E SE ,E ,E E E E E E N N N N N E E E E E E E E E E E E E E</p>
      <p>S S</p>
      <p>E
A A A</p>
      <p>C C C C C D D D D D E E E E E E E E E E E E E E
r
e
h
t
o
s
e
g
a
u
g
n
s
t e
an R
ic n
l o
app ita
- s
n t s
g u a
ire cd ”G
o o s
f o n ’s
ro w co le
f d v</p>
      <p>n d a
f m
liittceopT ireogagyhnBBT flfliiItrrsezaoaovyunnnpSdubAmm lifttrrrrececeeeoooabuuhhpudBTTm ´iliiftrssscceeceaooaaaaonpdSBmm ´´fiiiiiittttrsssceceeceoaooaadbSnnumm ıfi´iittrssssscceecceeaoaavnpn sagpuunOmM l)(saaooydndBBm lseeeoaaynhnBDm iiifrrceeanphETAm ilseceaaGGm lijttrrsseeceeaaooaaavvkkdnnhnnw lilifttrrssseeeaoaaaoynnnnpuEm iltsgahnnhE lirrsseceeegagoaohnbnpunE iitrssevpuh ijrszoauuYO iiliItttrrrsssseeoaaggaogadnnuuVVH illisseaoaovydndHM iliilffItrrrrsseeeoaooaoovnnpndCmm litreav llifttt()rreeeecaoaaaduudnhTAmmM iiiseooonnpnhP iillttrssececeeoaooanhnnudhdhTN iiiliiittttrssssseagaooaavunnbnunRD rse ii’ltrrrrssssseceegao”a”agvnhbnnPAD ili:-trecceceoaaagadudnnunpnhHmm lilittttrrseeeceeaaaaovndhnbnnPwm iiiilfttttrrrseeeecaooaaonpnnuRDMiltre”aauddpN iii’Itttsee”gaooa”ooaavvknnbndhBN liiliiiftttrrseeaoaagoagydbudDm il’frrssceoooaoakhdbnndnM</p>
      <p>Note that the evaluation measures described above differ slightly from the measures originally
proposed in the task description. 2 The original measures were based on the fact that spans are
linked to nuggets by assessors: as described in section 2, different spans linked to one nugget are
assumed to bear approximately the same factual content. Then, in addition to character-based
measures above, a nugget-based recall can be defined based on the number of nuggets (rather than
lengths of character spans) found by a system. However, an analysis of the assessments showed
that some assessors used nuggets in a way not intended by the assessment guidelines: namely, to
group related rather than synonymous character spans. We believe that this misinterpretation
of the assessment guidelines indicates that the guidelines are overly complicated and need to be
simpliefid in future edition of the task. As a consequence, we did not use nugget-based measures
for evaluation.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Runs</title>
      <p>In total, 12 runs were submitted from 4 research groups. To provide a baseline for the task, we
created an artificial run: for each topic, a response of the baseline was created as the ranked
list of at most 1000 snippets provided by Google in response to retrieval queries from the topic
definition. Note that the Google web search engine was designed for a task very different from
WebCLEF 2007 (namely, for the task of web page nfiding), and therefore the evaluation results of
our baseline can in no way be interpreted as an indication of Google’s performance.</p>
      <p>Table 2 shows the submitted runs with the basic statistics: the average length (the number of
bytes) of the snippets in the run, the averate number of snippets in the response for one topic,
and the average total length of response per topic.</p>
      <p>2See http://ilps.science.uva.nl/WebCLEF/WebCLEF2007/Tasks/\#Evaluation_measures.
We described WebCLEF 2007. This was the rfist year in which a new task was being assessed, one
aimed at undirected information search goals. While the number of participants was limited, we
believe the track was a success, as most submitted runs outperformed the Google-based baseline.
For the best runs, in top 7,000 bytes per topic about 1/5 of the text was found relevant and
important by the assessors.</p>
      <p>The WebCLEF 2007 evaluation also raised several important issues. The task denfiition did
not specify the exact size of a system’s response for a topic, which has make a comparison across
systems problematic. Furthermore, assessor’s guidelines appeared to be overly complicated: not
all assessors used nuggets as was intented by the organizers.</p>
      <p>Our plans for future work include a detailed per-topic analysis of runs and a study of re-usability
of the results of the assessments for improving the systems.
7</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>Valentin Jijkoun was supported by the Netherlands Organisation for Scientific Research (NWO)
under project numbers 220-80-001, 600.065.120 and 612.000.106. Maarten de Rijke was supported</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Balog</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kamps</surname>
          </string-name>
          , and M. de Rijke.
          <article-title>Overview of WebCLEF 2006</article-title>
          .
          <source>In CLEF</source>
          <year>2006</year>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Fuhr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lalmas</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A</surname>
          </string-name>
          . Trotman, editors.
          <source>Comparative Evaluation of XML Information Retrieval Systems: 5th International Workshop of the Initiative for the Evaluation of XML Retrieval</source>
          ,
          <string-name>
            <surname>INEX</surname>
          </string-name>
          <year>2006</year>
          . Springer,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>V.</given-names>
            <surname>Jijkoun and M. de Rijke</surname>
          </string-name>
          .
          <article-title>Overview of the WiQA task at CLEF 2006</article-title>
          .
          <source>In CLEF</source>
          <year>2006</year>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>V.</given-names>
            <surname>Jijkoun</surname>
          </string-name>
          and M. de Rijke.
          <article-title>WiQA: Evaluating Multi-lingual Focused Access to Wikipedia</article-title>
          . In T. Sakai,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanderson</surname>
          </string-name>
          , and D.K. Evans, editors,
          <source>Proceedings EVIA</source>
          <year>2007</year>
          , pages
          <fpage>54</fpage>
          -
          <lpage>61</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.E.</given-names>
            <surname>Rose</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Levinson</surname>
          </string-name>
          .
          <article-title>Understanding user goals in web search</article-title>
          .
          <source>In WWW '04: Proceedings of the 13th intern. conf. on World Wide Web</source>
          , pages
          <fpage>13</fpage>
          -
          <lpage>19</lpage>
          , New York, NY, USA,
          <year>2004</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sigurboj</surname>
          </string-name>
          ¨rnsson, J. Kamps, and M. de Rijke.
          <article-title>Overview of WebCLEF 2005</article-title>
          . In C. Peters,
          <string-name>
            <given-names>F.C.</given-names>
            <surname>Gey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , H. Mu¨ller,
          <string-name>
            <given-names>G.J.F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Magnini</surname>
          </string-name>
          , and M. De Rijke, editors,
          <source>Accessing Multilingual Information Repositories</source>
          , volume
          <volume>4022</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>810</fpage>
          -
          <lpage>824</lpage>
          . Springer,
          <year>September 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I.</given-names>
            <surname>Soboroff</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.P. de Vries</surname>
            , and
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Craswell</surname>
          </string-name>
          .
          <article-title>Overview of the TREC 2006 Enterprise Track</article-title>
          .
          <source>In The Fifteenth Text REtrieval Conference (TREC</source>
          <year>2006</year>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.M.</given-names>
            <surname>Voorhees</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.T.</given-names>
            <surname>Dang</surname>
          </string-name>
          .
          <article-title>Overview of the TREC 2005 question answering track</article-title>
          .
          <source>In The Fourteenth Text REtrieval Conference (TREC</source>
          <year>2005</year>
          ),
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>