<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLIR at NTCIR Workshop 3</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Noriko Kando National Institute of Informatics (NII)</institution>
          ,
          <addr-line>Tokyo</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper introduces NTCIR Workshop, a series of evaluation workshops, which are designed to enhance research in information retrieval and related text-processing techniques, such as summarization, question answering and extraction, by providing large-scale test collections and a forum for researchers. A brief history, tasks, participants, test collections, CLIR evaluation at the workshops, and brief overviews at the third NTCIR workshop are described in this paper. To conclude, some thoughts on future directions are suggested.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The NTCIR Workshop [1] is a series of evaluation workshops designed to enhance research in information access (IA)
technologies including information retrieval (IR), cross-lingual information retrieval (CLIR), information extraction
(IE), automatic text summarization, question answering, etc.</p>
      <p>The aims of the NTCIR project are;</p>
      <p>to encourage research in information access technologies by providing large-scale test collections reusable
for experiments and common evaluation infrastructures
to provide a forum of research groups interested in cross-system comparison and exchanging research ideas
in an informal atmosphere, and
to investigate methodologies and metrics for evaluation of information access technologies and methods for
constructing large-scale reusable test collections.</p>
    </sec>
    <sec id="sec-2">
      <title>1.1. Information Access</title>
      <p>A term "information access (IA)" includes a whole process to make information in the documents usable to the users. A
traditional IR system returns a ranked list of retrieved documents which are likely containing relevant information to
the user's information needs. This retrieves relevant documents from a vast document collection and makes these
documents usable for the users, and is one of the most fundamental and core process of IA. It is however not the end of
the story for the users. After obtaining a ranked list of retrieved documents, the user skims the documents, does
relevance judgments, locates the relevant information, reads, analyses, summarizes, compares the contents with other
documents, integrates, summarizes and does an information work such as decision making, problem solving, writing,
etc., based on the information obtained from the retrieved documents. We have looked such IA technologies to support
the users to utilize the information in the large-scale documents in document collections. The scope of the IA is more
closely related to the theory and practice of the digital libraries than IR.</p>
      <p>In the following, the next section provides definition of several terms used in this paper. Section 3 describes the NTCIR
Workshop and test collections. Section 4 discusses the continuum of the system- to user-oriented evaluation and the
context of evaluation design. Section 5 summarizes the discussion.
3. NTCIR</p>
    </sec>
    <sec id="sec-3">
      <title>3.1 Brief History of the NTCIR</title>
      <p>The NTCIR Workshop is periodical events which have been taken place once per about one and half years. It was
co-sponsored by the Japan Society for Promotion of Science (JSPS) as part of the JSPS "Research for Future" Program
(JSPS-RFTF 96P00602) and the National Center for Science Information Systems (NACSIS) since 1997. In April
2000, NACSIS was reorganized and changed its name to the National Institute of Informatics (NII). NTCIR was
co-sponsored by the JSPS and the Research Center for Information Resources at NII (RCIR/NII,) in FY 2000, and by
the RCIR/NII and Japanese MEXT1 Grant-in-Aid for Scientific Research on Informatics (#13224087) in and after
FY2001.The tasks, test collection constructed, participants of the previous workshops are summarized in Table 1.
For the First NTCIR Workshop, the process started with the distribution of the training data set on 1st November 1998,
and ended with the workshop meeting, which was held from 30 August to 1st September 1999 in Tokyo, Japan [5]. The
IREX [6], another evaluation workshop of IR and IE (named entities) using Japanese newspaper articles, and NTCIR
joined forces in 2000 and have worked together to organize the NTCIR Workshop since then. The challenging tasks of
Text Summarization and Question Answering became feasible with this collaboration.</p>
      <p>An international collaboration to organize Asian languages IR evaluation was proposed at the 4th International
Workshop on Information Retrieval with Asian Languages (IRAL'99). In accordance with the proposal, the Chinese
Text Retrieval Tasks are organized by Hsin-Hsi Chen and Kuang-hua Chen, National Taiwan University, at the second
workshop, and Cross Language Retrieval of Asian languages at the third workshop.</p>
      <p>For the Second Workshop, the process was started from June 2000 and the meeting was held on 7-9 March 2001, NII,
Tokyo [7]. The process of the Third NTCIR Workshop starts from August, 2001 and the meeting will be held on 8-10
October 2002, NII, Tokyo[8].</p>
    </sec>
    <sec id="sec-4">
      <title>3.2 Focus of the NTCIR</title>
      <p>Through the series of the NTCIR Workshops, we have looked at both traditional laboratory-typed IR system testing and
evaluation of challenging technologies. For the laboratory-typed testing, we have placed emphasis on 1) information
retrieval (IR) with Japanese or other Asian languages and 2) cross-lingual information retrieval. For the challenging
issues, 3) shift from document retrieval to technologies to utilize "information" in documents, and 4) investigation for
evaluation methodologies, including evaluation of automatic text summarization; multi-grade relevance judgments for
IR; evaluation methods appropriate to the retrieval and processing of a particular document-genre and its usage of the
1 MEXT: Ministry of Education, Culture, Sports, Science and Technology
user group and so on.</p>
      <p>From the beginning, CLIR is one of the central interests of the NTCIR. It was because CLIR between English and own
languages are critical for international information transfer in Asian countries, and it was challenging that CLIR
between languages with completely different structures and origins such as English and Chinese, or English and
Japanese. It was also partly because CLIR techniques are needed even for monolingual text retrieval [9]. For example,
a part of a document is sometimes written in English (ex. A Japanese document often contains an English abstract or</p>
    </sec>
    <sec id="sec-5">
      <title>3.3 Test Collections</title>
      <p>A test collection is a data set used in system testing or experiments. In the NTCIR project the term "test collections" are
used for any kind of data sets usable for system testing and experiments however it often means IR test collections used
in search experiments.</p>
      <p>
        The test collections constructed through NTCIR Workshops are listed in Table 2.
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) a set of topics: written statements of user's search request.
lang
      </p>
      <p>J
CE
JE
J</p>
      <p>J</p>
      <sec id="sec-5-1">
        <title>CKJE</title>
      </sec>
      <sec id="sec-5-2">
        <title>CKJE J J J</title>
        <p>
          topic
#
83
50
49
30
50
31
110
200+800
(
          <xref ref-type="bibr" rid="ref1 ref3">3</xref>
          ) relevance judgments: a list of relevant documents for each topic (right answers)
In the retrieval experiments, relevance judgments are most expensive procedure. However, once test collections are
created, they can be independent from the settings of the original experiment and can be repeatedly used in the different
experiments. Some of the NTCIR test collections contain additional data such as, tagged corpus in NTCIR-1(sentences
in the selected documents are segmented manually) and segmented data in NTCIR-2 (every sentence in a whole
documents set is automatically segmented into words and phrases beforehand).
        </p>
        <sec id="sec-5-2-1">
          <title>3.3.1 Documents</title>
          <p>Documents were collected from various domain or genres. Each task carefully selected the appropriate domain of
document collection And the task (experiment) design and relevance judgment criteria are set according to each
document collection and the supposed user community who use the type of documents in the everyday tasks.
Fig 1 shows an sample record in NTCIR-1 JE. More than half of the documents in the NTCIR-1 JE Collection are
English-Japanese paired. Documents are plain text with SGML-like tags in the NTCIR collections. A record may
contain document ID, title, a list of author(s), name and date of the conference, abstract, keyword(s) that were assigned
by the author(s) of the document, and the name of the host society.
&lt;DOC&gt;
&lt;DOCNO&gt;ctg_xxx_19990110_0001&lt;/DOCNO&gt;
&lt;LANG&gt;EN&lt;/LANG&gt;
&lt;HEADLINE&gt; Asia Urged to Move Faster inShoring Up Shaky Banks &lt;/HEADLINE&gt;
&lt;DATE&gt;1999-01-10&lt;/DATE&gt;
&lt;TEXT&gt;
&lt;P&gt;HONG KONG, Jan 10 (AFP) - Bank for International Settlements (BIS) general manager Andrew Crockett has urged Asian
economies tomove faster in reforming their shaky banking sectors, reports said Sunday. Speaking ahead of Monday's meeting at the
BIS office here of international central bankers including US Federal Reserve chairman Alan Greenspan, Crockett said he was
encouraged by regional banking reforms but "there is still some way to go." Asian banks shake off their burden of bad debt if they
were to be able to finance recovery in the crisis-hit region, he said according to the Sunday Morning Post. Crockett added that more
stable currency exchange rates and lower interest rates had paved the way for recovery. "Therefore I believe in the financial area, the
crisis has in a sense been contained and that now it is possible to look forward to real economic recovery," he was quoted as saying by
the Sunday Hong Kong Standard.&lt;/P&gt;
&lt;P&gt;"It would not surprise me, given the interest I know certain governors have, if the subject of hedge funds was discussed during the
meeting," Crockett said. &lt;/P&gt;
&lt;P&gt;He reiterated comments by BIS officials here that the central bankers would stay tight-lipped about their meeting, the first to be
held at the Hong Kong office of the Swiss-based institution since it opened last July. &lt;/P&gt;
&lt;/TEXT&gt;
&lt;/DOC&gt;
3.3.2 Topics
A sample topic record used in the CLIR at the NTCIR Workshop 3 is shown in Fig. 2. Topics are defined as statements
of "user's requests" rather than "queries", which are the strings actually submitted to the system, since we wish to allow
both manual and automatic query construction from the topics.</p>
          <p>The topics contain SGML-like tags. A topic consists of the title of the topic, a description (question), a detailed
narrative, and a list of concepts and field(s). The title is a very short description of the topic and can be used as a very
short query that resembles those often submitted by users of Internet search engines. Each narrative may contain a
detailed explanation of the topic, term definitions, background knowledge, the purpose of the search, criteria for
judgment of relevance, etc.
&lt;TOPIC&gt;
&lt;NUM&gt;013&lt;/NUM&gt;
&lt;SLANG&gt;CH&lt;/SLANG&gt;
&lt;TLANG&gt;EN&lt;/TLANG&gt;
&lt;TITLE&gt;NBA labor dispute&lt;/TITLE&gt;
&lt;DESC&gt;
&lt;/DESC&gt;
&lt;NARR&gt;
&lt;/NARR&gt;
&lt;CONC&gt;
&lt;/CONC&gt;
&lt;/TOPIC&gt;
To retrieve the labor dispute between the two parties of the US National Basketball Association at the end of 1998 and the agreement that they reached.
The content of the related documents should include the causes of NBA labor dispute, the relations between the players and the management, main
controversial issues of both sides, compromises after negotiation and content of the new agreement, etc. The document will be regarded as irrelevant if
it only touched upon the influences of closing the court on each game of the season.</p>
          <p>NBA (National Basketball Association), union, team, league, labor dispute, league and union, negotiation, to sign an agreement, salary, lockout, Stern,
Bird Regulation.</p>
          <p>Fig. 2. A Sample Topic (CLIR at NTCIR WS 3)</p>
        </sec>
        <sec id="sec-5-2-2">
          <title>3.3.3 Relevance Judgments (Right Answers)</title>
          <p>The relevance judgments were conducted using multi-grades. In relevance judgment files contained not only the
relevance of each document in the pool, but also contained extracted phrases or passages showing the reason the
analyst assessed the document as "relevant". These statements were used to confirm the judgments, and also in the hope
of future use in experiments related to extracting answer passages.</p>
          <p>In addition, we proposed new measures, weighted R precision and weighted average precision, for IR system testing
with ranked output based on multi-grade relevance judgments [10]. Intuitively, the highly relevant documents are more
important for users than the partially relevant, and the documents retrieved in the higher ranks in the ranked list are
more important. Therefore, the systems producing search results in which higher relevance documents are in higher
ranks in the ranked list, should be rated as better. Based on the review of existing IR system evaluation measures, it was
decided that both of the proposed measures be single number, and can be averaged over a number of topics.
Most IR systems and experiments have assumed that the highly relevant items are useful to all users. However, some
user-oriented studies have suggested that partially relevant items may be important for specific users and they should
not be collapsed into relevant or irrelevant items, but should be analyzed separately [11]. More investigation is required.</p>
        </sec>
        <sec id="sec-5-2-3">
          <title>3.3.4 Linguistic analysis (additional data)</title>
          <p>NTCIR-1 contains a "Tagged Corpus". This contains detailed hand-tagged part-of-speech (POS) tags for 2,000
Japanese documents selected from NTCIR-1. Spelling errors are manually collected. Because of the absence of explicit
boundaries between words in Japanese sentences, we set three levels of lexical boundaries (i.e., word boundaries, and
strong and weak morpheme boundaries).</p>
          <p>In NTCIR-2, the segmented data of the whole J (Japanese document) collection are provided. They are segmented into
three levels of lexical boundaries using a commercially available morphological analyzer called HAPPINESS. An
analysis of the effect of segmentation is reported in Yoshioka et al. [12].</p>
        </sec>
        <sec id="sec-5-2-4">
          <title>3.3. 5 Robustness of the System Evaluation using the Test Collections</title>
          <p>The test collections NTCIR-1 and -2 have been tested for the following aspects, to enable their use as a reliable tool for
IR system testing:
exhaustiveness of the document pool
inter-analyst consistency and its effect on system evaluation
topic-by-topic evaluation.</p>
          <p>The results have been reported and published on various occasions [13–16]. In terms of exhaustiveness, pooling the top
100 documents from each run worked well for topics with fewer than 100 relevant documents. For topics with more
than 100 relevant documents, although the top 100 pooling covered only 51.9% of the total relevant documents,
coverage was higher than 90% if combined with additional interactive searches. Therefore, we conducted additional
interactive searches for the topics with more than 50 relevant documents in the first workshop, and those with more
than 100 relevant documents in the second workshop.</p>
          <p>When the pool size was larger than 2500 for a specific topic, the number of documents collected from each submitted
run was reduced to 90 or 80. This was done to keep the pool size practical and manageable for assessors to keep
consistency in the pool. Even though the numbers of documents collected in the pool were different for each topic, the
number of documents collected from each run is exactly the same for a specific topic.</p>
          <p>
            A strong correlation was found to exist between the system rankings produced using different relevance judgments and
different pooling methods, regardless of the inconsistency of the relevance assessments among analysts and regardless
of the different pooling methods used [13–15,17]. It served as an additional support to the analysis reported by
Voorhees [18].
The first NTCIR Workshop [5] hosted three tasks below;
1. Ad Hoc Information Retrieval Task: to investigate the retrieval performance of systems that search a static
set of documents using new search topics (J&gt;JE).
2. Cross-Lingual Information Retrieval Task: an ad hoc task in which the documents are in English and the
topics are in Japanese (J&gt;E).
3. Automatic Term Recognition and Role Analysis Task: (1) to extract terms from titles and abstracts of
documents, and (
            <xref ref-type="bibr" rid="ref2">2</xref>
            ) to identify the terms representing the "object", "method", and "main operation" of the
main topic of each document.
          </p>
          <p>In the Ad Hoc Information Retrieval Task, the document collection containing Japanese, English and Japanese-English
paired documents is retrieved by Japanese search topics. In Japan, document collections often naturally consist of such
a mixture of Japanese and English. Therefore, the Ad Hoc IR Task at the NTCIR Workshop 1 is substantially CLIR,
although some of the participating groups discarded the English section and performed the task as a Japanese
monolingual IR.</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>Communications Research Laboratory (Japan)</title>
        <p>
          Fuji Xerox (Japan)
Fujitsu Laboratories (Japan)
Central Research Laboratory, Hitachi Co. (Japan)
JUSTSYSTEM Corp. (Japan)
Kanagawa Univ. (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) (Japan)
KAIST/KORTERM (Korea)
Manchester Metropolitan Univ. (UK)
Matsushita Electric Industrial (Japan)
NACSIS (Japan)
National Taiwan Univ. (Taiwan ROC)
NEC (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) (Japan)
NTT (Japan)
        </p>
      </sec>
      <sec id="sec-5-4">
        <title>RMIT &amp; CSIRO (Australia)</title>
        <p>Tokyo Univ. of Technology (Japan)
Toshiba (Japan)
Toyohashi Univ. of Technology (Japan)
Univ. of California Berkeley (US)
Univ. of Lib. and Inf. Science (Tsukuba, Japan),
Univ. of Maryland (US)
Univ. of Tokushima (Japan)
Univ. of Tokyo (Japan)
Univ. of Tsukuba (Japan)
Yokohama National Univ. (Japan)</p>
        <p>Waseda Univ. (Japan).
3.5 NTCIR Workshop 2 (June 2000 -- March 2001)
The second workshop [7] also hosted three tasks, and each task was proposed and organized different research group
on the topic.</p>
      </sec>
      <sec id="sec-5-5">
        <title>ATT Labs &amp; Duke Univ. (US)</title>
        <p>Communications Research Laboratory (Japan),
Fuji Xerox (Japan)
Fujitsu Laboratories (Japan)
Fujitsu R&amp;D Center (China PRC)
Central Research Laboratory, Hitachi Co. (Japan)
Hong Kong Polytechnic (Hong Kong, China PRC)
Institute of Software, Chinese Academy of Sciences
(China PRC)
Johns Hopkins Univ. (US)
JUSTSYSTEM Corp. (Japan)
Kanagawa Univ. (Japan)
Korea Advanced Institute of Science and Technology
(KAIST/KORTERM) (Korea)
Matsushita Electric Industrial (Japan)
National TsinHua Univ. (Taiwan, ROC)
NEC Media Research Laboratories (Japan)</p>
      </sec>
      <sec id="sec-5-6">
        <title>National Institute of Informatics (Japan)</title>
        <p>
          NTT-CS &amp; NAIST (Japan)
OASIS, Aizu Univ. (Japan)
Osaka Kyoiku Univ. (Japan)
Queen College-City Univ. of New York (US)
Ricoh Co. (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) (Japan)
Surugadai Univ. (Japan)
Trans EZ Co. (Taiwan ROC)
Toyohashi Univ. of Technology (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) (Japan)
Univ. of California Berkeley (US)
Univ. of Cambridge/Toshiba/Microsoft (UK)
Univ. of Electro-Communications (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) (Japan)
Univ. of Library and Information Science (Japan)
Univ. of Maryland (US)
Univ. of Tokyo (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) (Japan),
Yokohama National Univ. (Japan)
        </p>
        <p>
          Waseda Univ. (Japan).
3.6 NTCIR Workshop 3 (Sept. 2001 -- Oct. 2002)
The third NTCIR Workshop started with the document data distribution in September 2001 and the workshop meeting
will be held in October 2002. We selected five areas of research as tasks; (1) Cross-language information retrieval of
Asian languages (CLIR), (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) Patent retrieval (PATENT), (
          <xref ref-type="bibr" rid="ref1 ref3">3</xref>
          ) Question answering (QAC), (4) Automatic text
summarization (TSC2), and (
          <xref ref-type="bibr" rid="ref4">5</xref>
          ) Web retrieval (WEB). The updated information is available at
http://research.nii.ac.jp/ntcir/workshop/.
3.6.1 Cross-Language Retrieval Task (CLIR)
After the second NTCIR workshop [1, 2], researchers from Japan, Korea, and Taiwan have discussed a much more
complicated cross-language information retrieval (CLIR) evaluation task, which is closer to the realistic application of
IR environment and is a real challenge to IR researchers. We propose a CLIR task in the third NTCIR workshop and
organize an executive committee to fulfill this task. The CLIR Task Executive Committee consists of 9 researchers
from Japan, Korea, and Taiwan. These members meet 3 times in Japan to discuss the details of CLIR Task, to make the
schedule, and to arrange the agenda. Topic creation and relevance judgments on each language documents are done by
each country group. Evaluation and report writing were done by Kuang-hua Chen and pooling was done by Kazuko
Kuriyama.
        </p>
        <p>Documents and topics are in four languages (Chinese, Korean, Japanese and English). Fifty topics for the collections of
1998–1999 (Topic98) and 30 topics for the collection of 1994 (Topic94). Both topic sets contain four languages
(Chinese, Korean, English and Japanese). Context of the experimental design is "report writing".</p>
        <p>Multilingual CLIR (MLIR): Search the document collection of more than one language by one of four languages
of topics, except the Korean documents because of the time range difference (Xtopic98&gt;CEJ).</p>
        <p>Bilingual CLIR (BLIR): Search of any two different languages as language and documents, except the searching
of English documents (Xtopic98&gt;C, Xtopic94&gt;K, Xtopic98&gt;J).</p>
        <p>Single Language IR (SLIR): Monolingual Search of Chinese, Korean, or Japanese. (Ctopic98&gt;C, Ktopic94&gt;K,</p>
        <p>Jtopic98&gt;J).</p>
        <p>When we think of the "layers of CLIR technologies"[19], the CLIR of newspaper articles closely related to the
"pragmatic layer (social, cultural convention, etc) " and cultural/social differences among the countries is the issues we
should attack in both topic creation and retrieval. For the scientific information transfer, CLIR between English and
own language is the one of the biggest interests in East Asian countries. In these years, interests towards social/cultural
aspects in East Asia is increasing especially in younger generation. Also technological information transfer among Asia
is one of the critical issues in business and industrial sector. According to these changes in the social needs, the CLIR
task has changed from English-Japanese Scientific documents to multilingual newspapers and patent documents.
For the next NTCIR Workshop, Korean newspaper articles published in 1998-99 in both English and Korean language
will be added, then multilingual CLIR of four languages of Chinese, Korean, English which is published in Asia and
Japanese will be feasible.</p>
        <p>After executing relevance judgment, we find the number of relevant documents for some topics is small or is zero. For
example, the topics created by Japanese member have few relevant Chinese documents. As a result, the members of
Executive Committee of CLIR Task have discussed how to screen out the unsuitable topics for each combination of
target languages based on a basic idea of keeping as many topics as possible. Here, the source language is query
language and the target language(s) is/are document language(s). We adopt the so-called “3-in-S+A” criterion. The
“3-in-S+A” means that a qualified topic must have at least 3 relevant documents with ‘S’ or ‘A’ score. Based on this
criterion, we identify various topic sets for each combination of target language document set no matter what source
(query) languages are used.</p>
        <p>
          We construct a NTCIR-3 Formal Test Collection for the Evaluation of NTCIR-3 CLIR Task which we describe in
the following.For Chinese, Japanese, and English Document Set (note that these are 1998-1999 news articles) and the
accompanying 1998-1999 Topic Set with 50 topics as we send to you in the beginning of CLIR task, we create the
following sub-test collection based on “3-in-S+A” criterion. In the FORMAL Test Collection, each target document set
(C, J, E, CJ, CE, JE, CJE, and K) has different set of topics.
(1) NTCIR-3 Formal Chinese Test Collection
It contains 381,681 Chinese documents and 42 topics in source language of Chinese, Japanese, Korean, and English.
The IDs of topics in 1998-1999 Topic Set used in this collection are;
1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 18, 19, 20, 21, 22, 23, 24, 25, 27, 32, 33, 34, 35, 36, 37, 38,
39, 40, 42, 43, 45, 46, 47, 48, 49, and 50.
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) NTCIR-3 Formal Japanese Test Collection
It contains 220,078 Japanese documents and 42 topics in source language of Chinese, Japanese, Korean, and English.
The IDs of topics in 1998-1999 Topic Set used in this collection are;
2, 4, 5, 7, 8, 10, 12, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37,
38, 39, 40, 41, 42, 43, 44, 45, 46, 47, and 50.
(
          <xref ref-type="bibr" rid="ref1 ref3">3</xref>
          ) NTCIR-3 Formal English Test Collection
It contains 22,927 English documents and 32 topics in source language of Chinese, Japanese, Korean, and English. The
IDs of topics in 1998-1999 Topic Set used in this collection are;
2, 4, 5, 7, 9, 12, 13, 14, 18, 19, 20, 21, 23, 24, 26, 27, 28, 29, 31, 32, 33, 34, 35, 36, 37, 38, 39, 42, 43, 45, 46,
and 50.
(4) NTCIR-3 Formal CJ Test Collection
It contains 601,759 Chinese and Japanese documents and 50 topics in source language of Chinese, Japanese, Korean,
and English. All of the 50 topics in 1998-1999 Topic Set are used in this collection.
(
          <xref ref-type="bibr" rid="ref4">5</xref>
          ) NTCIR-3 Formal CE Test Collection
It contains 404,608 Chinese and English documents and 46 topics in source language of Chinese, Japanese, Korean, and
English. The IDs of topics in 1998-1999 Topic Set used in this collection are;
1 ,2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 31, 32 ,33, 34,
35, 36, 37, 38, 39, 40, 42, 43, 45, 46, 47, 48, 49, and 50.
(
          <xref ref-type="bibr" rid="ref5">6</xref>
          ) NTCIR-3 Formal JE Test Collection
It contains 243,005 Japanese and English documents and 45 topics in source languages of Chinese, Japanese, Korean,
and English. The topics in 1998-1999 Topic Set used in this collection are;
2, 4, 5, 7, 8, 9, 10, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35,
36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, and 50.
(
          <xref ref-type="bibr" rid="ref6">7</xref>
          ) NTCIR 3 Formal CJE Test Collection
It contains 624,686 Chinese, Japanese, and English documents and 50 topics in source language of Chinese, Japanese,
Korean, and English. All of the 50 topics in 1998-1999 Topic Set are used in this collection 2. For Korean Document
Set (note that these documents are 1994 news articles) and the accompanying 1994 Topic Set with 30 topics as we send
to you in the beginning of CLIR task, we create one sub-test collection also based on “3-in-S+A” criterion.
(
          <xref ref-type="bibr" rid="ref7">8</xref>
          ) NTCIR-3 Formal Korean Test Collection
It contains 66,146 Korean documents and 30 topics in Chinese, Japanese, Korean, and English. All of the 30 topics in
1994 Topic Set are used in this collection. To sum up, we will have a NTCIR-3 Formal Test Collection with 8 sub-test
collections for all of the possible combination of target language documents except that Korean documents could not be
combined with Chinese, Japanese, and English documents.
3.6.2 Patent Retrieval Task (PATENT)
Context of the experimental design is "search for technological trend survey". Regarding "Cross Genre Retrieval", we
assumes that someone send a newspaper article clip to a patent intermediary and ask to retrieve the related patents.
Search using ordinary topic fields such as &lt;DESC&gt;, &lt;NARR&gt;, etc. are accepted as well as non-mandatory runs. Topic
creation and relevance judgments were conducted by professional patent intermediaries who are members of the
Information Retrieval Committee at the Japan Intellectual Property Association. The task design was also done with
close collaboration with these professionals.
        </p>
        <p>Main Task</p>
        <p>Cross-language Cross-Genre Retrieval: retrieve patents in response to newspaper articles associated
with technology and commercial products. Thirty query articles with a short description of the search
request. Topics are available in Japanese, English, Chinese (simplified, traditional), and Korean.</p>
        <p>Monolingual Associative Retrieval: retrieve patents associated with a Japanese patent as input. Thirty
query patents with a short description of search requests.</p>
        <p>Optional task: Any research reports are invited on patent processing using the above data, including, but not
limited to: generating patent maps, paraphrasing claims, aligning claims and examples, summarization for
patents, clustering patents.
document:
-</p>
        <p>Japanese patents: 1998–1999 (ca. 17GB, ca 700K docs)
JAPIO patent abstracts: 1995–1999 (ca. 1750K docs)
Patent Abstracts of Japan (English translations for JAPIO patent abstracts): 1995–1999 (ca. 1750K)
Newspaper articles (included in topics)
3.6.3 Question Answering Challenge (QAC)</p>
        <p>Task 1: System extracts five answers from the documents in some order. One hundred questions. The system is
required to return support information for each answer to the questions. We assume the support information
is a paragraph, hundred-character passage or document that includes the answer.</p>
        <p>Task 2: System extracts only one answer from the documents. One hundred questions. Support information is
required.</p>
        <p>Task 3: evaluation of a series of questions. The related questions are given for 30 of the questions of Task 2.
3.6.4 Text Summarization Challenge (tsc2)</p>
        <p>Task A (single-document summarization): Given the texts to be summarized and summarization lengths, the
participants submit summaries for each text in plain text format.</p>
        <p>Task B (multi-document summarization): Given a set of texts, the participants produce summaries of it in plain
text format. The information, which was used to produce the document set, such as queries, as well as
summarization lengths, is given to the participants.
3.6.5 Web Retrieval Task (web)
Survey retrieval is a search for survey and aims to retrieve many relevant documents as possible. Target retrieval is a
search aiming a few highly relevant documents to get a quick answer for the search request represented as a topic.
"Topic retrieval" is a search in response to a search request and "similarity retrieval" is a search by given relevant
document(s). In the relevance judgments, one-hop linked documents were also included in the consideration. A topic
contain several extra fields specialized to Web retrieval such as a) known relevant documents, b) information on topic
author, who is basically relevance assessors.</p>
      </sec>
      <sec id="sec-5-7">
        <title>A. Survey Retrieval (both recall and precision are evaluated) A1. Topic Retrieval A2. Similarity Retrieval B. Target Retrieval (precision-oriented)</title>
        <p>C. Optional Task</p>
        <p>C1.Search Results Classification
C2. Speech-Driven Retrieval</p>
        <p>
          C3. Other
3.6.6 Features of the NTCIR Workshop 3 Tasks
For the next workshop, we planed some new ventures, including:
(1)
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          )
(
          <xref ref-type="bibr" rid="ref1 ref3">3</xref>
          )
(4)
(
          <xref ref-type="bibr" rid="ref4">5</xref>
          )
(
          <xref ref-type="bibr" rid="ref5">6</xref>
          )
(
          <xref ref-type="bibr" rid="ref6">7</xref>
          )
(
          <xref ref-type="bibr" rid="ref7">8</xref>
          )
        </p>
        <p>Multilingual CLIR (CLIR)
Search by Document (Patent, Web)
Passage Retrieval or submit "evidential passages", passages to show the reason the documents are supposed
to be relevant (Patent, QA, Web)
Optional Task (Patent, Web)
Multi-grade Relevance Judgments (CLIR, Patent, Web)
Various Relevance Judgments (Web)
Precision-Oriented Evaluation (QA, Web).</p>
        <p>
          Various types of relevance judgments
For (1), it is our first trial of the CLEF [20] model in Asia. We would like to invite any other language groups who wish
to join us by providing document data and relevance judgments or by providing query translation. For (
          <xref ref-type="bibr" rid="ref1 ref3">3</xref>
          ), we suppose
that identifying the most relevant passage in the retrieved documents is required when retrieving longer documents
such as Web documents or patents. The primary evaluation will be done from the document base, but we will use the
submitted passages as secondary information for further analysis.
(4). For patent and Web tasks, we invite any research groups who are interested in the research using the document
collection provided in the tasks for any research projects. Those document collections are new to our research
community and many interesting characteristics are included. We also expect that this venture will explore the new
tasks possible for future workshops.
        </p>
        <p>
          For (
          <xref ref-type="bibr" rid="ref4">5</xref>
          ), we have used multi-grade relevance judgment so far since it is more natural to the users than binary judgments
although most of the standard metrics used for search effectiveness are calculated based on the binary relevance
judgments. We uses "cumulated gain" [21] and proposes new metrics, "weighted average precision" [10] for that
purpose. We will continue this line of investigation and will add "top relevant" for the Web task, as well as standard
metrics can be calculated by trec_eval.
        </p>
      </sec>
      <sec id="sec-5-8">
        <title>The some results will be reported at the Workshop.</title>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>4. Discussion</title>
      <p>The multilingual CLIR of Asian languages has just started from this year. So far, the needs for CLIR in East Asia was
focused on the CLIR between own language and English, or other "international language". Mutual interests towards
each other culture are increasing quite recently, especially among younger generation. Information transfer in the
technological domain is also very acute. For example, Japanese technological information is a great interest for Korean
industry and in other side, Korea is the second largest patent import, just next to the Unite State, to Japan. Regardless of
the close interaction through the long history, each language is quite different. Even Taiwan ROC, Japan, People's
Republic of Chine use "Chinese Characters" in the own languages, the pronounce and sentence structures are
completely different between Chinese and Japanese and the Characters are simplified and modified in own way in PRC
and Japan, so that we can not understand each other using own languages. Also there are no umbrella organization like
the Europe Union in Europe. However the needs for CLIR among these languages are increasing in rather informal, or
a grass-root movement like way.</p>
      <p>In responding to such movement, some search engine provide simple CLIR functionalities and an operational
information providers specialized to patent plan to start CLIR of East Asian languages from the next year. Therefore, it
was several years behind from the situation in Europe, now we are in the stage to initiate the multilingual CLIR of east
Asian languages.</p>
      <p>The multilingual CLIR of East Asian languages have to tackle to new problems, like character codes (there are 4
standards for Japanese character codes, 2 in Korea and the character codes used in the simplified Chinese and the
traditional Chinese are different and can not be 100% certain convert), less available resources and research staff who
can understand the contents of other languages, transliteration of proper names in English documents, and so on. The
structure of languages are quite different each other. In summary, the barriers for cross language information access in
East Asia are various in every layer of the CLIR technologies shown below, and this bring new challenges to CLIR
research and CLIR evaluation design and organization as well.</p>
      <p>pragmatic layer: cultural &amp; social aspects, convention
semantic layer: concept mapping
lexical layer: language identify, indexing
symbol layer: character codes
physical layer: network
For the future direction, Korean newspaper articles and English newspaper articles publish in Korea in 1998-1999, the
same year of Chinese, Japanese and Taiwan-English and Japan-English. So the real multilingual CLIR of Chinese,
Korea, Japanese, and English's in each country will be feasible for the next NTCIR.</p>
      <p>Moreover, we may think of the possibility of the following CLIR, which are rather aiming to the application for the
operational CLIR systems. In the following, focuses placed on the directions towards implications in operational or real
life CLIR systems.</p>
      <p>Cumulating the experiences
Switching language CLIR
Task/genre oriented CLIR
Pragmatic layer of CLIR technologies and identifying the differences</p>
      <p>Towards CL information access
Cumulating the experiences: In the Internet environment, we have to cope with wide varieties of languages and
combination of languages in operational setting. IR contains lots of language dependent process and the effective IR
technologies can vary according to the languages or the combination of query and target languages. Therefore it is
critical to cumulate the experiences on each language or combination of languages by reviewing and summarizing the
findings of the CLIR research on them like building blocks of understanding on the effectiveness of CLIR technologies
on each languages and language combination. Evaluation campaigns or workshops which wide variety of CLIR
systems tested various approached on the same collection on the common infrastructures for cross-system comparison
are expected to contribute to such activities to cumulate and summarize the experiences.</p>
      <p>Switching (Pivot) language CLIR: In the Internet environment which wide varieties of languages are included, the
shortness of the resource for translation knowledge is one of the critical problems for CLIR. In the real world
environment, we can often find parallel or quasi-parallel document collections of own languages and English in
non-English speaking countries. It has been proposed by many CLIR researchers to utilize such parallel document
collections and connecting them by using English as switching or key language to obtain translation knowledge, but the
actual testing of such direction of researches are seldom evaluated. It was partly because relevance judgments on such
multilingual document collections are difficult to be done by an individual research group. International collaboration
of evaluation activities can contribute to this direction.</p>
      <p>Pragmatic layer of CLIR technologies which cope with social and cultural aspects of languages are ones of the most
challenging issues of real world CLIR. CLIR research has so far placed emphasis on technologies to provide access to
the relevant information across the different languages. Identifying the differences of viewpoints expressed in different
languages or in documents produced in different cultural or social backgrounds is also critical to improve the realistic
global information transfer across the languages.</p>
      <p>Towards CL information access: In order to widen the scope of the CLIR in the whole process of the information
access, we can think of the technologies to make information in the document more usable for users, for example, cross
language summarization, cross language question answering, cross language text mining, and so on. The technologies
to enhance the interaction between systems and users or to support the query construction on CLIR systems are also
included in this direction.</p>
    </sec>
    <sec id="sec-7">
      <title>References:</title>
      <p>[1] NTCIR Project: http://research.nii.ac.jp/ntcir/
[2] TREC. http://trec.nist.gov/
[3] Smeaton, A.F. and Harman, D. "The TREC (IR) experiments and their impact on Europe", Journal of Information Science, No.</p>
      <p>23, pp 169-174, 1997.
[4] Sparck Jones, K., Rijsbergen, C.J. Report on the need for and provision of an 'ideal' information retrieval test collection,</p>
      <p>Computer laboratory, Univ. Cambridge., 1975 (BLRDD Report)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          3.4 NTCIR Workshop 1 (Nov.
          <fpage>1998</fpage>
          --
          <lpage>Sept</lpage>
          .
          <year>1999</year>
          )
          <article-title>1</article-title>
          .
          <string-name>
            <given-names>Chinese</given-names>
            <surname>Text Retrieval</surname>
          </string-name>
          <article-title>Task (CHTR): including English-Chinese CLIR (ECIR; E&gt;C) and Chinese monolingual IR (CHIR tasks, C&gt;C) using the test collection CIRB010, consisting of newspaper articles from five newspapers in Taiwan R</article-title>
          .O.C.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Japanese-English IR</surname>
          </string-name>
          <article-title>Task (JEIR): using the test collection of NTCIR-1 and -2, including monolingual retrieval of Japanese</article-title>
          and
          <string-name>
            <surname>English (J&gt;J</surname>
            ,
            <given-names>E</given-names>
          </string-name>
          &gt;E),
          <article-title>and CLIR of Japanese</article-title>
          and
          <string-name>
            <surname>English (J&gt;E</surname>
            , E&gt;
            <given-names>J</given-names>
            , J
          </string-name>
          &gt;JE,
          <string-name>
            <surname>E</surname>
          </string-name>
          &gt;JE).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Text</given-names>
            <surname>Summarization</surname>
          </string-name>
          <article-title>Task (TSC: Text Summarization Arrange): text summarization of Japanese newspaper articles of various kinds</article-title>
          .
          <source>The NTCIR-2 Summ Collection was used.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[5] NTCIR Workshop 1: Proceedings of the First NTCIR Workshop on Research in Japanese Text Retrieval and Term Recognition</source>
          , Tokyo, Japan,
          <volume>30</volume>
          Aug.-1
          <string-name>
            <surname>Sept</surname>
          </string-name>
          .,
          <year>1999</year>
          .
          <fpage>ISBN4</fpage>
          -924600-77-6. (http://research.nii.ac.jp/ntcir/workshop/ OnlineProceedings/).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <surname>IREX</surname>
            <given-names>URL</given-names>
          </string-name>
          :http://cs.nyu.edu/cs/projects/proteus/irex/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>[7] NTCIR Workshop 2: Proceedings of the Second NTCIR Workshop on Research in Chinese &amp; Japanese Text Retrieval and Text Summarization</source>
          , Tokyo, Japan, June 2000-March
          <year>2001</year>
          .
          <fpage>ISBN4</fpage>
          -924600-96-2. (http://research.nii.ac.jp/ntcir/workshop/OnlineProceedings/)
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[8] NTCIR Workshop 3 Meeting: Working Note of the Third NTCIR Workshop Meeting</source>
          , Tokyo, Japan, Oct.
          <volume>8</volume>
          -
          <fpage>10</fpage>
          ,
          <year>2002</year>
          . 6 vols.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>"Cross-linguistic scholarly information transfer and database services in Japan"</article-title>
          .
          <article-title>Presented in the panel on Multilingual Database in the Annual Meeting of the American Society for Information Science</article-title>
          , Washington DC. , USA,
          <year>November 1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoshioka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>"Evaluation based on multi-grade relevance judgments"</article-title>
          .
          <source>IPSJ SIG Notes</source>
          , Vol.
          <volume>2001</volume>
          <source>-FI-63</source>
          , pp.
          <fpage>105</fpage>
          -
          <lpage>112</lpage>
          ,
          <year>July 2001</year>
          . (in Japanese w/English abstract)
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Spink</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greisdorf</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <article-title>"Regions and levels: Measuring and mapping users' relevance judgments"</article-title>
          .
          <source>Journal of the American Society for Information Sciences</source>
          , Vol.
          <volume>52</volume>
          , No.
          <issue>2</issue>
          , pp.
          <fpage>161</fpage>
          -
          <lpage>173</lpage>
          ,
          <year>2001</year>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Yoshioka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
          </string-name>
          , N.:
          <article-title>"Analysis on the usage of Japanese segmented</article-title>
          texts in
          <source>the NTCIR Workshop 2." In NTCIR Workshop 2: Proceedings of the Second NTCIR Workshop on Research in Chinese &amp; Japanese Text Retrieval and Text Summarization</source>
          , Tokyo, June 2000-March
          <source>2001 ISBN 4-924600-96-2).</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N</given-names>
          </string-name>
          , Nozue,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Kuriyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Oyama</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          ,
          <article-title>"NTCIR-1: Its policy and practice"</article-title>
          ,
          <source>IPSJ SIG Notes</source>
          , Vol.
          <volume>99</volume>
          , No.
          <volume>20</volume>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          ,
          <year>1999</year>
          . (in Japanese w/English abstract)
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nozue</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>"Pooling for a large scale test collection: Analysis of the search results for the pre-test of the NTCIR-1 Workshop"</article-title>
          ,
          <source>IPSJ SIG Notes</source>
          , Vol.
          <volume>99</volume>
          -FI-54, pp.
          <fpage>25</fpage>
          -
          <issue>32</issue>
          <year>May</year>
          ,
          <year>1999</year>
          [in Japanese].
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <article-title>"Construction of a large scale test collection: Analysis of the training topics of the NTCIR-1"</article-title>
          ,
          <source>IPSJ SIG Notes</source>
          , Vol.
          <volume>99</volume>
          -FI-55, pp.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          ,
          <year>July 1999</year>
          . (in Japanese w/English abstract)
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eguchi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <article-title>"Construction of a large scale test collection: Analysis of the test topics of the NTCIR-1"</article-title>
          ,
          <source>In Proceedings of IPSJ Annual Meeting [in Japanese]</source>
          . pp.
          <fpage>3</fpage>
          -
          <lpage>107</lpage>
          -- 3-
          <issue>108</issue>
          , 30 Sept.-3
          <string-name>
            <surname>Oct</surname>
          </string-name>
          .
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Kuriyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoshioka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <article-title>"Effect of cross-lingual pooling"</article-title>
          .
          <source>In NTCIR Workshop 2: Proceedings of the Second NTCIR Workshop on Research in Chinese &amp; Japanese Text Retrieval and Text Summarization</source>
          , Tokyo, June 2000-March
          <source>2001 ISBN 4-924600-96-2)</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Voorhees</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          ,
          <article-title>"Variations in relevance judgments and the measurement of retrieval effectiveness"</article-title>
          ,
          <source>In Proceedings of 21st Annual International ACM-SIGIR Conference on Research and Development in Information Retrieval</source>
          .,
          <string-name>
            <surname>Melbourne</surname>
          </string-name>
          , Australia,
          <year>August 1998</year>
          , pp.
          <fpage>315</fpage>
          -
          <lpage>323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Kando</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <article-title>"Towards real multilingual information discovery and access "</article-title>
          .
          <source>Presented at ACM Digital Libraries and ACM-SIGIR Joint Workshop on Multilingual Information Discovery and Access</source>
          .
          <article-title>Panel on the Evaluation of the Cross-Language Information Retrieval</article-title>
          . Berkeley, CA, USA,
          <year>August 15</year>
          ,
          <year>1999</year>
          . (http://www.clis2.umd.edu/ conferences/midas/papers/kando2.ppt)
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20] CLEF:
          <string-name>
            <surname>Cross-Language Evaluation</surname>
          </string-name>
          Forum, http://www.iei.pi.cnr.it/DELOS/CLEF
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Jarvelin</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Kekalainen</surname>
            <given-names>ACM-SIGIR</given-names>
          </string-name>
          <year>2000</year>
          p.
          <fpage>41</fpage>
          -
          <lpage>48</lpage>
          J.:
          <article-title>IR evaluation methods for retrieving highly relevant documents</article-title>
          .
          <source>Proceedings of 2000</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Nozue</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kando</surname>
          </string-name>
          , N. “
          <article-title>Primary considerations in the concept of relevance: Relevance judgement of NTCIR”</article-title>
          .
          <source>IPSJ SIG Notes, 99-FI-53</source>
          , Vol.
          <volume>99</volume>
          , No.
          <volume>20</volume>
          ,
          <year>March 1999</year>
          , p.
          <fpage>49</fpage>
          -
          <lpage>56</lpage>
          . (in Japanese w/English abstract)
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>