<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cross-language relevance assessment and task context</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Appendix A</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jussi Karlgren and Preben Hansen Swedish Institute of Computer Science</institution>
          ,
          <addr-line>SICS Box 1263, SE-164 29 Kista</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>An experiment on how users assess relevance in a foreign language they know well is reported. Results show that relevance assessment in a foreign language takes more time and is prone to errors compared to assessment in the reader's first language. The results are related to task and context and an enhanced methodology for performing context-sensitive studies is reported.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Cross-linguality and reading</title>
      <sec id="sec-1-1">
        <title>1.1 People are naturally multi-lingual</title>
        <p>For people in cultures all around the world competence in more than one language is quite common
and the European cultural area is typical in that respect. Many people, especially those engaged in
intellectual activities are familiar with more than one language and have some acquaintance with
several.</p>
      </sec>
      <sec id="sec-1-2">
        <title>1.2 People are good at making relevance assessments</title>
        <p>Information access systems deliver results which on a good day hold up to forty per cent relevant
items. It is up to the reader to winnow out the good stuff from the bad.</p>
        <p>We know that readers are excellent at making relevance assessments for texts. Both assessment
efficiency and precision are very impressive. But how we go about it we know very little about.
Practice seems to improve both assessment speed, assessment precision, and assessor confidence, but
what features a reader focuses on and how they are combined has not been studied in any great detail.</p>
      </sec>
      <sec id="sec-1-3">
        <title>1.3 Linguistic competence is a continuum</title>
        <p>Languages are tools tied to tasks. For any one task, typically people have one language they prefer to
perform it in. In general, while people may have working knowledge of more than one language, it is
not common for people to have equal competence in many; the first language, or the school language,
or the workplace language will tend to be stronger for whatever task they are engaged in. Linguistic
competence is not a binary matter: people know a language to some extent, greater or lesser. What bits
of competence are important in any given situation is an ongoing discussion in the field of language
teaching – we will here concentrate on some aspects of reading, related to situation, task, and domain.</p>
      </sec>
      <sec id="sec-1-4">
        <title>1.4 Assessing relevance in a strange language is hard – and important</title>
        <p>We do know that reading about strange things in strange genres takes more time than familiar genres,
and that reading a language we do not know well is hard work, and something we only attempt if we
believe it is worth the effort.</p>
        <p>Judging trustworthiness and usefulness of documents in a foreign language is difficult and a noticeably
less reliable process than doing it in a language and cultural context we are familiar with.
These starting points have immediate ramifications for the design of cross-lingual and multi-lingual
information access systems. Presenting large numbers of documents to users if it is likely they will not
be able to determine their usefulness is a waste at best and a trustworthiness and reliability risk at
worst.</p>
      </sec>
      <sec id="sec-1-5">
        <title>1.5 Finding out more – does language make a difference?</title>
        <p>We need more data about reading and related processes. To find out more we set up an experiment
where Swedish-speaking subjects, fluent in English as determined by self-report, were presented with
retrieval results both languages, and given the task of rating the results by relevance. Our hypotheses
were that results for a foreign language would be more time-consuming and less competent than those
for the first language.</p>
      </sec>
      <sec id="sec-1-6">
        <title>1.6 Task-based approach to query construction and relevance assessment</title>
        <p>Generally, topicality has been the main criteria for relevance in information retrieval experiments. Our
approach suggests that other criteria may come into play, especially criteria related to the task and
domain at hand. For interactive information retrieval experiments, we propose to expand the original
query with information about context. In this study, we want to relate the relevance assessment to a
specific task situation, i.e. the subject will be given a semi-realistic situation including a domain
description, and then we will investigate if the relevance assessment situation involves criteria beyond
topicality.
2. Experiment
2.1 Set-up
- Participants: The study involved 12 participants divided into 3 groups. Groups A and B were
given a workplace scenario involving a domain with relevant work-tasks. Group C was given the
i-CLEF queries without context information.
- Scenario: Each scenario had 4 participants.
- Language. 2 languages were used: English and Swedish.
- Queries. The four CLEF queries used in this year’s interactive track were used in both languages:
queries 53, 56, 65, and 80. Query 86 was used for a practice run.
- Result list. Sets of ranked result lists of length between one and two hundred were produced in
Swedish using Siteseeker, a commercial web-based search system by Euroseek AB, on the TT
CLEF corpus and English using Inquery on the LA Times CLEF corpus.
- Presentation. The ranked lists were presented to the participants, varied by order and language
(cf. Table 1) in a simulated search interface.
- System. The experiment infrastructure was built using HTTP and was deployed over the WWW.</p>
        <p>The canned ranked results were put up as html pages and linked to the actual documents, which
were displayed with four buttons to be used for the relevance ranking. A simple cgi-bin based
logging tool noted the relevance assessment made and the time taken to make the assessment after
display of the document.
- Questionnaires. The participants filled out questionnaires at various points in the study. The data
was collected either by semi-structured questions or measured by a Likert scale of 1 to 5 or 1 to 3.
- Relevance categories. The participants could in the interface indicate for each document one of
four assessments: “not relevant” “somewhat relevant”, “relevant”, and “don’t know”.</p>
      </sec>
      <sec id="sec-1-7">
        <title>2.2 Simulated Domain and Work-Task Scenarios</title>
        <p>In this study we use the Simulated Domain and Work-Task Scenario (SDWS) methodology, an
evaluation methodology with simulated contexts that include description of domains and work-tasks.
The method is an extension of the notion of simulated work-tasks (Brajnic et al., 1995; Borlund, 2000;
Ruthven et. al., 2002) among others. Borlund and Ruthven enhanced the context of standard queries
using two fields with descriptive information. We extend this design to include a domain description
and a general work-task description. The goal of the method is to give the experimental query a context
closer to a real-life information-seeking situation. In this way, the SDWS would allow the user a) a
broader understanding of the situation, and b) a subjective interpretation of the relevance.
Constructing a SDWS query within a context was done by creating two levels of description (cf. Figure
2): a general description including a short description of the domain and a short description of general
work-tasks or routines that are performed. The next level contains a situational description including
the topic of the query (in this case the I-clef query) and a search task description, which also include
parts of the description field of the actual I-clef query (cf. appendix A for a SDWS for query CO53).
Results</p>
        <p>General descriptions:</p>
        <p>Domain:
Results Work-task description:
JUSSitSuaI:tiAonbaoludtestchreipCtiLonE:F-runs</p>
        <p>Topic:</p>
        <p>Search task description:</p>
      </sec>
      <sec id="sec-1-8">
        <title>2.3 Procedure</title>
        <p>The participants were asked to answer some initial questions. After that, participants in groups A or B
were asked to read through a workplace scenario carefully and try to act within the assigned scenario
as well as possible. Then participants were asked to read through the first work-task related query and
to assess the ranked list for it pursuant time constraints as per the scenario, or in the case of group C, to
keep the time about constant around fifteen to twenty minutes per query. After the assessment
participants were asked to answer a fixed set of questions related to the query and the work task. This
fixed set of questions was repeated after each of the four queries. Finally, after the last query,
participants were asked to answer a last set of questions.</p>
      </sec>
      <sec id="sec-1-9">
        <title>2.4 Participant</title>
        <p>The 12 participants in this study had a variety of academic and professional backgrounds. 5
participants were male and 7 female, with an average age of 36,5. The participants had an overall high
experience searching web-based search engines such as Google (4,33) and an overall low experience in
searching commercial databases (2,16) and using machine translation tools such as Babel-fish (2.00).
2/3 of the participants used some kind of search engine 1-2 times every day. Average on overall
knowledge in English was 4,25 (see app. B for a full version and table of the pre-questionnaire). Note
that this information is based on the participants’ own subjective judgments.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Results</title>
      <sec id="sec-2-1">
        <title>3.1 Foreign-language texts took longer to assess and were assessed less well</title>
        <p>Assessing texts in English (30 s average assessment time) took longer than for Swedish (19 s). Given
the extra effort invested into reading the English texts it is somewhat surprising to find that the results
of the assessments were significantly less reliable for English than for Swedish as well (cf. Figure 2; all
differences between English and Swedish significant by Mann Whitney U; p &gt; 0,95). Assessments
were judged by how well they correspond to the CLEF official assessments; precision and recall are
calculated with respect to the known relevant documents found in the retrieved and presented set of
documents. In general, the precision is reasonably high for both languages, which can be taken to
indicate that participants went through the list and found most relevant documents in the presented list.
All documents are very short. The Swedish documents are from a wire service and the English
documents from a newspaper. The average length of an English article is over seven hundred words,
whereas the Swedish articles are of an average length of just over four hundred. The difference in
averages is partially due to the English average being highly skewed from a few very long feature
articles, a genre almost entirely missing from the Swedish corpus. The length difference could account
for part of the assessment time difference, but since the length of the article correlates very weakly
with assessment time (Spearman’s Rho = 0,3) that explanation can be discounted</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2 Task focus may have an effect on assessment performance</title>
        <p>No significant differences between scenarios (cf. Figure 3) could be found, other than a tendency for
group B to perform better (p &gt; 0,75; Mann Whitney U) than group A or the control group. As found by
questionnaire, group B invested less effort in topic and more in task related aspects of relevance than
did group A, which may be a tentative explanation for the tendency; this relation needs to be
investigated further before any conclusions can be drawn, however.</p>
      </sec>
      <sec id="sec-2-3">
        <title>3.3 Relevance judgment aspects</title>
        <p>We assumed that aspects of the relevance judgment taken into account would extend beyond
traditional topicality. In order to see if aspects other than topicality were taken into account, we added
two more levels related to our domain and task-based scenario approach. After each query, the
participants were asked what aspects of relevance judgments were of any importance for their
assessment. We present the results for groups A and B in Table 2. Merged, the two groups used the
domain related aspect in 12% of the cases, the task related aspect in 46% of the cases, and the
topicrelated aspect in 42% of the cases. All observations were done over all four i-CLEF queries given to
the participants. Notable is that 36% in the A-group and 61 % in the B-group marked that their
assessments were related to task. Another interesting observation is that nobody in the group B
reported using the domain-related aspect in assessments. Group A had a level of 44% on topic-related
aspect and 36% on task-related aspects.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Discussion</title>
      <p>The results are quite convincing. Time matters. Relevance assessment in a foreign language, even a
familiar one, is more time-consuming and more difficult than in one’s first language. Tasks seem to
matter. Generally, traditional information retrieval experiments are based on algorithmic and topical
relevance. In this study we have seen that other aspects do count in the relevance assessment.
Furthermore, we have a weak but interesting indication that the Simulated Domain and Work-Task
Scenario applied may have an effect on the assessment performance. This is but a first step in this
direction; we intend to pursue this avenue of inquiry further, and investigate its effects on design.
Specifically, during the coming year we will investigate if adding more information to the interface
will improve results for the foreign language assessment task.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>We thank Tidningarnas Telegrambyrå AB, Stockholm, for providing us with the Swedish text
collection, and Heikki Keskustalo, University of Tampere, Johan Carlberger and Hercules Dalianis,
Euroseek AB, for kind support in producing the ranked lists of documents. Furthermore, we gratefully
acknowledge funding provided to us the European Commission under contracts IST-2000-29452
(DUMAS) and IST-2000-25310 (CLARITY). And finally we thank our patient subjects for the time
they spent reading really old news.
Hansen, P., &amp; Järvelin, K. (2000). The Information Seeking and Retrieval process at the Swedish
Patent- and Registration Office. Moving from Lab-based to real life work-task environment.
Proceedings of the ACM-SIGIR 2000 Workshop on Patent Retrieval, Athens, Greece, July 28, 2000,
pp. 43-53.</p>
      <p>Ruthven, I., Lalmas, M. and van Rijsbergen, K. (2002). Ranking Expansion Terms with Partial and
Ostensive Evidence. Proceedings of the Fourth International Conference on Conceptions of Library
and Information Science – CoLIS4, Seattle, USA, July, 2002, pp. 199-220.</p>
      <sec id="sec-4-1">
        <title>The SDWS framework description</title>
        <p>The following is a full version of a simulated Domain and Work-task Scenario (SDWS) (translated
from the Swedish original) for I-clef query C053
General descriptions:</p>
        <p>Domain: Monitoring news and translation services
Work task: Among your daily work-tasks you monitor and translate news information within a
specific areas based on profiles set up by external customers. Your customers are
usually companies and public institutions.</p>
        <p>Situational description</p>
        <p>Topic: Genes and Diseases
Search task: You have been assigned to monitor incoming news items that describe genes, which
cause disease on humans. The customer especially wants documents that identify or
report the discovery of a gene that is the source of any type of disease, syndrome,
behavioural or developmental disorder in humans. Any information or document
that reports the discovery of a defective gene that causes problems in humans is
relevant. Documents that describe diseases and disorders caused by the absence of a
gene are not relevant</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Appendix B</title>
      <sec id="sec-5-1">
        <title>The pre-questionnaire</title>
        <p>Searching online library catalogues?
Searching commercial databases
(such as Dialog)
Searching Internet-based search
engines such as Google
Using tools for machine translation
(such as Babelfish)
How often do you use any kind of
search engine?
I like searching for information
My reading skills in English ...
Neutral
3
Agree</p>
        <p>4</p>
        <p>Strongly agree</p>
        <p>5
Very good
5
7
SUM</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>