<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>being sought are not expressed in such a language</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>document selection. The focus of this</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>writing) vocabulary</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>uency</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>however</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>the query would be posed in a language for which</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2001</year>
      </pub-date>
      <volume>33</volume>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Gloss</title>
      <p>MT
Gloss
MT</p>
    </sec>
    <sec id="sec-2">
      <title>An English topic description, consisting of title, description, and narrative that served as a basis</title>
      <p>for the CLIR system’s query,</p>
    </sec>
    <sec id="sec-3">
      <title>Topic 17, Topic 11</title>
      <p>Topic 11, Topic 17
Topic 17, Topic 11
Topic 11, Topic 17</p>
    </sec>
    <sec id="sec-4">
      <title>Participant Before break After break</title>
      <p>=P )=R
eness
The original French version of each document, and
3
An English translation of each document that was produced using Systran Professional 3.0.
F =
Table 1: iCLEF-2001 experiment design as run. Topics 11 and 13 are broad, Topic 17 and 29 are narrow.
Provide the participants judgments to the iCLEF coordinators in a standard format for scoring.
retained locally.
Conduct data analysis using the scored results and other measurements that were recorded and</p>
    </sec>
    <sec id="sec-5">
      <title>A ranked list of the top 50 documents produced automatically by a CLIR system using an English</title>
      <p>query,</p>
    </sec>
    <sec id="sec-6">
      <title>Have participants make relevance judgments for each topic. Each participant was allowed 20 minutes</title>
      <p>or \unsure." A \not judged" response was also available.
for each topic (including reading the topic description, reading as many documents or document
was asked to select one of four possible judgments: \not relevant," \somewhat relevant," \relevant,"
summaries as time allowed, and making relevance judgments). For each document, the participant</p>
    </sec>
    <sec id="sec-7">
      <title>For each topic, the following resources were provided:</title>
      <p>French test collection contained four search topics for use in the experiment, plus a fth practice topic.</p>
    </sec>
    <sec id="sec-8">
      <title>The iCLEF experiment was designed in a manner similar to that used in the TREC Interactive Track,</title>
      <p>in which a Latin square design is used to block topic and searcher eects so that the system eect can be
performance of related tasks, but time and resource limitations precluded our use of a larger sample.
Control ) and two \narrow" topics that asked about some specic ev ent (e.g., Nobel Prize for Economics
The four topics included two \broad" topics that asked about a general subject (e.g. Conference on Birth
judgments were used only to evaluate the results after the experiment was completed. As might be
expected, it turned out that in every case there were more relevant documents in the top-50 for the broad
this design, every searcher sees all four topics, two with one system and two with the other. The order
fatigue and learning eect on the observ ability of the system eect. W e realized at the outset that four
in 1994 ). Relevance judgments for the top-50 documents for each topic were also known, but those
The task to be performed at each participating site included:
characterized. Table 1 shows the order in which topic-system combinations were presented to users. In
in which topics and systems are presented is varied systematically in order to minimize the impact of
topics than for the narrow ones.
participants was an undesirably small number given the large variability that has been observed in human</p>
    </sec>
    <sec id="sec-9">
      <title>Gloss</title>
      <p>MT
Gloss
MT
and their subjective assessment of the two systems.
Ask each searcher to complete questionnaires regarding their background, each search, each system,
measure for the evaluation:
An unbalanced version of van Rijsbergen’s F measure was selected for use as the ocial eectiv
umd02
umd04
umd03
umd01</p>
    </sec>
    <sec id="sec-10">
      <title>Topic 13, Topic 29</title>
      <p>Topic 29, Topic 13
Topic 29, Topic 13
Topic 13, Topic 29
+ (1
1
was optional, but we choose to use them as our MT system.</p>
      <p>Design and implement two interactive document selection systems. Use of the Systran translations
identify terms in the document that can be translated using the bilingual term list. If no multi-word or
term list in which all source language terms have been If the source-language term was still stemmed.3
because we wanted to focus on a single factor (the translation strategy). As we have before, we chose the
single word match is found, the French word in the document is stemmed and a match with the term
possible translations for some terms. In past work, we have explored display strategies for presenting
retrieval to extend the source-language (French) coverage of the term list. The rst step w as to remove
Bilingual term lists found on the Web often contain an eclectic combination of root and inected forms.
then proceeded in the normal reading order through the text, using greedy longest string matching to
English translation that occurred most often in the Brown Corpus (a balanced corpus of English) when
all punctuation and convert every character to unaccented lower case in both the documents and the
not found in the term list, it was copied unchanged into the translated document. We used the stemmer
term list. This had the eect of minimi zing problems due to character encoding. The translation process
that we had developed for CLEF 2000 for this purpose. Bilingual term lists typically contain several
We therefore applied the same backo translation strategy that w e have previously used for automatic
multiple alternatives, but for our iCLEF experiments we chose only a single translation for each term
list is attempted again. If that fails, the previous step is repeated using a second version of the bilingual
more than one possible translation was present in the term list.
design allows users to make a quick pass through the documents and then go back for a more
(not shown in the Figure 1).
detailed examination if they desire. The submit button is at the bottom of the ranked list page
Simultaneously record the relevance judgments for all documents when a search is completed. This</p>
    </sec>
    <sec id="sec-11">
      <title>Provide topic selection and translation option selection mechanisms.</title>
      <p>3.4 Searcher Characteristics
3.3 User Interface
5
an eye tracker, since multiple summaries are displayed on the same page. On the other hand, since
Javascript timer built in a CGI script. The timer was started when the title link was selected,
judgment in such cases is likely to be quite small.
Record the amount of time spent on judging each document. This was implemented with a
mary since in that case the title link would never be selected. It would be hard to do better without
and stopped when one of the relevance judgment radio buttons was selected. One can easily see
this method fails to record the time correctly if the judgment was based solely on a displayed
sumthe summaries are very short (often only one line on the screen), the time required to render a
topic (see Figure 1). The summary information that we displayed for this experiment is simply the
the topic description) that appeared in a translated summary were detected using string matching
Display a ranked list providing summary information for the top 50 documents for the selected
relevance judgments to be selected, with \not judged" initially selected for all documents.
and highlighted in red and rendered in italics. A set of v e radio buttons under each title allowed
translation of its title, as specied b y the appropriate SGML tag. Query terms (i.e., any term in
this capability was not needed. The system included the following capabilities:
a Web browser, and their relevance judgments are recorded by the server when a search is completed. A
relevance judgments for that topic are nished. A searc h-ID is assigned to each search so that multiple
ments [4]. The system uses a Web-based server-side architecture. Searchers interact with the system using
experiment was based on an existing system that we had developed for our TREC-9 CLIR track
experidierences b y using the same user interface with both types of translation. The user interface for our
Because we wished to compare translation strategies, we sought to minimize the eect of presen tation
searches can be tracked simultaneously, but participants in the study completed the task individually so
search starts when a searcher selects a topic and a translation option (MT or Gloss) and ends when the
so no speed dierence bet ween translation types is apparent to the searcher. Again, query terms
that appeared in a translated document were highlighted in red and rendered in italics.
is selected by a searcher. All translations are performed in advance and cached within the server,
Display the translation of the full text of a document in a separate window whenever that document
to the ranked document list so that the searcher can read and understand it before making any
Display topic descriptions based on the searcher’s selection. The topic is displayed separately prior
relevance judgments, and it remains displayed at the top of the page once the ranked list is displayed
as a ready reference.</p>
    </sec>
    <sec id="sec-12">
      <title>Science and some familiarity with cross-language retrieval and is working as a user interface programmer.</title>
      <p>summer session limited the pool of potential participants, however, and the 3-hour search session made
ment, since we expect that librarians could make extensive use of CLIR systems when conducting searches
We had originally intended to recruit graduate students in library science to participate in our
experiThe fourth participants (umd04) has a Bachelors degree in religion, is currently working as a nancial
were doctoral students in the College of Information Studies, and both have interests in information
on behalf of people with dieren t language skills. The fact that the experiments were performed during
deadline approach, we therefore became somewhat less selective. Of the four participants in our
experparticipation less appealing even though we oered a cash pa yment ($20) to each participant. As the
retrieval and human-computer interaction. A third subject (umd02) has a Masters degree in Computer
iments, two (umd01 and umd03) held a Masters degree in Library Science. Both of those participants
controller, and professed no interest in the technical details of what we were doing.
umd01. Participant umd01 reported 14 years of searching experience, much more than any other
participant.
beginning of the session, we were therefore somewhat surprised. Unfortunately, there was not
umd03. Participant umd03 was the only one to report good reading skills in French (the others reporting
French in high school. Clearly we need to give more thought to how we conduct language skills
sucien t time remaining before the deadline to recruit an additional participant. Interestingly,
screening.
mentioned this when recruiting subjects. When we saw this answer on the questionnaire at the
poor skills or none). Knowledge of French was disallowed by the track guidelines, and we had
after the experiment, participant umd03 mentioned in a casual conversation that they had studied
had at least v e years of online searching experience. All four participants reported a great deal of
The ages of the four participants range were between 28 and 35 at the time of the experiment. None
experience searching the World Wide Web and a great deal of experience of using a point-click interface.
Our observations during the experiment agreed with their assessments on this point.
In addition to the backgrounds described above, the following self-reported characteristics of
distinof the participants had been involved in previous interactive retrieval experiments of this sort, but all
guished an individual participants from the group:
systems, the only one to report typically searching less than once a day (for umd04, the response
umd04. Participant umd04 was the only one of the four with no experience searching online commercial
was twice a week) and the only one to give a neutral response to the question of how they feel about
searching (the others reporting that they either enjoy or strongly enjoy searching).</p>
    </sec>
    <sec id="sec-13">
      <title>After a few further changes, we froze the conguration of the in terface for the experiments reported in</title>
      <p>The four search sessions were conducted individually by the rst author of this paper. Upon arriv al,
hour peer review session with several graduate students who were working on computational linguistics.
point-click interface, and reading the document language. Following that was a 30-minute tutorial in
For each search, the experimenter would tell the participant which topic and system to select, and then
along with the searcher, pointing out specic details that migh t have been incompletely understood when
break was necessary, and none took it. The rst searc h then started.
this step, the searcher was asked to take a 10-minute break. Interestingly, no participant thought this
the two systems and provided an unstructured space for additional comments.
this paper.
a searcher was rst giv en a 10-minute brief introduction to the goal of the study, the procedure of the
occasionally ask questions of the experimenter, but we tried to minimize this tendency. Each search was
the experimenter would quietly observe the search process and take observation notes. Participants did
experiment, the tasks he or she was expected to complete, and the time allocation for each step. Then a
5which the two systems were introduced. The tutorial was conducted in a hands-on fashion|the searcher
basic demographic information and information about the searcher’s experience with searching, using
that they had made. When two searches with the same system were completed, a questionnaire regarding
the searcher’s experience with that system was conducted. That was followed by a 10-minute break and
The iCLEF experiment in Maryland started on June 27, 2001, and ended on July 9, 2001. We began with a
necessary. We found that all the searchers learned how to use the systems in less than 30 minutes. After
started with making relevance judgments for that topic, and their degree of condence in the judgmen ts
practiced using the systems while reading printed instructions line-by-line. The experimenter followed
followed by a 5-minute questionnaire regarding the searcher’s familiarity with the topic, the ease of getting
exit questionnaire was completed. That questionnaire sought the participant’s subjective comparison of
minute pre-search questionnaire was completed. The major purpose of that questionnaire was to collect
small (two-user) pilot study, after which we made some changes to our system. We then conducted a
halfthen the process was repeated with the second system. After all the four searches were completed, an
Table 2 shows the ocial results on a per-searc h basis, and Table 3 shows the result of averaging the
p &lt; 0:05). This is probably due to the an insucien t numbers of degrees of freedom in our test (i.e., too
better with MT than gloss translation on broad topics, and all four searchers did better with MT on
implementation of gloss translation when scored using the ocial measure.</p>
      <p>A couple of observations are easily made from Table 3. The values of for narrow topics are F0:8
hypothesis, the preponderance of the evidence suggests that MT is better for this task than our present
make relevance judgments more accurately for narrow topics than for broad ones. Another interesting
few participants), since the trend seems quite clear. So although we cannot reject the second of our null
measures of the two participants that experienced each condition. Three of the four searchers did F0:8
narrow topics. A two-tail paired t-test (p&lt;0.05), found no signican t dierence in either case, ho wever, at
consistently higher than the values for broad topics. This suggests that searchers are typically able to
0.13
0.20
0.10
0.27
0.28
GLOSS
0.78
0.41
0.83
0
GLOSS
0
each category). That may, however, be an artifact of the presence of a greater density of truly relevant
narrow topics that helps users to make more total judgments and to get the balance between relevant and
judgments as a third \system" for which only two types of judgment were provided. Clearly, many more
make is that for broad topics, our participants seemed to exhibit a greater proclivity to assess documents
judgments on nonrelevant documents is particularly striking, suggesting that there is something about
not relevant judgments about right, regardless of the system type. One other observation that we could
As Figure 2 (b) shows, \unsure" and \somewhat relevant" judgments took longer on average than
\relExamining the time required to make relevance judgments provides another perspective on our results.
analysis to further explore our results. Figure 2 (a) shows the average number of documents to receive
No single measure can reect ev ery interesting aspect of the data, so we performed some descriptive data
documents near the top of any well constructed ranked list.
documents were left unjudged for broad topics than for narrow ones. The highly skewed distribution of
as relevant than as not relevant (based on the fraction of the ocial judgmen ts that they achieved in
each type of relevance judgment by topic and system type. In that gure, w e treat the ocial CLEF
ability to use gloss translations for this purpose. The rst part of this conclusion is ten tative because we
least a factor of two when using the MT system, and two of the four participants beat it by that much
of documents for broad and narrow topics) that might produce higher values for F0:8.
zero and one.
be useful, but that there is substantial variation across the population of searchers with regard to their
as relevant. That guarantees a recall of 1.0 (since we compute recall over the relevant documents in
In order to test our rst n ull hypothesis, we must construct some simple strategy that does not require
fairly well around the mean, for narrow topics the values have a bimodal distribution with peaks near
looking at the documents. One way to do this is to simply selects all 50 documents in the ranked list
the top-50, not over all relevant documents known to CLEF). The precision is then the fraction of the
when using gloss translation. From this we tentatively conclude that both MT and gloss translation can
average over all topics for when computed in this way is 0.26. All participants beat that value by at F0:8
observation is that the values of for broad topics exhibit a strong central tendency by clustering F0:8
have not yet tried some other rules (e.g., always select the top 10 documents, or select dieren t numbers
entire list that happens to be relevant, which is much larger for broad topics than narrow ones. The
relevant:" 398, \somewhat relevant:" 57, \relevant:" 89, and \unsure:" 20. Comparing these numbers
this speculation. We observed that some searchers often modied their relev ance judgment, either right
judged document. We observed that other searchers rarely changed their relevance judgments, however,
learn to recognize documents in a category based on their recollection of documents that have been
previously assigned to that category. Our observation of search behavior oers some evidence to support
evant" judgments, and \not relevant" judgments could be performed the most quickly. This was true
had fewer \not judged" cases. The seemingly excessive time required to reach a judgment of \somewhat
It is interesting to note that the track guidelines did not provide any formal denition for the t ypes
judgment types to our participants, and no searcher expressed any confusion regarding this terminology.
them based on the common meanings of the terms. In our study, we provided no further explanation of the
of relevance judgments, presumably assuming that both experimenters and searchers would understand
so it is not clear how pervasive this eect is.
for any sort of inference.
category. One possible explanation for this would be a within-topic learning eect, in whic h searchers
afterwards or later when they worked on a dieren t document. In that second case, presumably their
is the focus of the next subsection.
relevant" when using gloss translation results from a single data point, and therefore provides little basis
judgment of the relevance of the later document seemed to be related to the relevance of a previously
for both topic types, and it helps to explain why narrow topics (which have few relevant documents)
For this reason, we decided to explore whether the participants interpreted these terms consistently. That
inverse relationship between the number of documents and time required to assign a document to that
The total number of documents of each relevance judgment type (across both topic types) is: \not
with the average amount of time per document of each relevance type in Figure 2 (b), we see a clear</p>
    </sec>
    <sec id="sec-14">
      <title>4.3 Comparing Strict and Loose Relevance Judgments</title>
      <p>(a) (b)
9
that loose judgments produce higher values and values below the axis indicating that strict judgments
For the ocial results \somewhat relev ant" was treated as \not relevant." For the sake of brevity, we
measure increased, it would indicates that on average the participants were being stricter than necessary in
would have been better. Two trends are evident in this data. First, broad topics benet more from
loose judgments than narrow topics. Second, the improvement for gloss translation was more consistent
Figure 3 depicts this dierence for eac h of the 16 searches, with bars above the X axis indicating
relevant" as \relevant," a scenario that we call \loose" relevance judgments. Our key idea was simple: we
will refer to that as \strict" relevance judgment. We could equally well choose to treat \somewhat
relevant" judgments that people made with MT and and gloss translation were actually dieren t in some
than the improvement for MT. There were 40 judgments of \somewhat relevant" for MT, but only 17
recomputed the measure with all \somewhat relevant" judgments treated as \relevant," and if the F0:8
for gloss translation, so more does not seem to be better in this case. It seems that the \somewhat
making their relevance judgments. Table 4 shows the value by search with loose relevance judgments, F0:8
and Table 5 compares the average value by systems and judgment type. Higher values are obtained F0:8
from loose judgments in both cases, but the improvement is far larger for gloss translation than for MT.
that recall may not be a discriminating factor.
surprising, however, since there are so few relevant documents to be found in the case of narrow topics
0.01
0.68
0.05
0.31
1
0.93
1
0.70
0.03
0.43
0.08
0.09
Figure 4 shows, that participant actually achieved the lowest average values for three of the four measures
was that some of the subjects might actually know quite a bit about one of the topics. This actually did
The rst of these w as that participant umd03 reported that they had good reading skills in French. As
(although two or three other participants were close in every case), and Table 3 shows that this poor
performance was consistent for both MT and gloss translation. The other factor we had concern about
Two other factors that had been of potential concern to us turned out not to make much of a dierence.
search was zero for both values of . Go gure.
happen in one case, again with searcher umd03, for topic 29. As it turned out, the value of F for that
0.55
0
0.93
0
helping with the peer review of our system. This work has been supported in part by DARPA cooperative
study, Gina Levow for help with gloss translation, Bob Allen for advice on statistical signicance testing,
The authors would like to thank Clara Cabezas for assistance in setting up the systems used in the
agreement N660010028910.
our participants for their willingness to invest their time in this study, and everyone from the CLIP lab
http://www.glue.umd.edu/ oard/research.html.</p>
      <p>In Carol Peters, editor, Proceedings of the First Cross-Language Evaluation Forum. 2001. To appear.
[2] Douglas W. Oard. Evaluating interactive cross-language information retrieval: Document selection.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>