<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluating Cross-Language Explicit Semantic Analysis and Cross Querying at TEL@CLEF 2009</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maik Anderka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nedim Lipka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benno Stein</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Cross-Language Information Retrieval, Cross-Language Explicit Semantic Analysis, Wikipedia,
Cross Querying</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Media, Media Systems Bauhaus University Weimar 99421 Weimar</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our participation in the TEL@CLEF task of the CLEF 2009 adhoc track. The task is to retrieve items from various multilingual collections of library catalog records, which are relevant to a user's query. Two different strategies are employed: (i) the Cross-Language Explicit Semantic Analysis, CL-ESA, where the library catalog records and the queries are represented in a multilingual concept space that is spanned by aligned Wikipedia articles, and, (ii) a Cross Querying approach, where a query is translated into all target languages using Google Translate and where the obtained rankings are combined. The evaluation shows that both strategies outperform the monolingual baseline and achieve comparable results. Furthermore, inspired by the Generalized Vector Space Model we present a formal definition and an alternative interpretation of the CL-ESA model. This interpretation is interesting for real-world retrieval applications since it reveals how the computational effort for CL-ESA can be shifted from the query phase to a preprocessing phase.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Cross-language information retrieval, CLIR, is the task of retrieving documents from a target
collection written in a language different from the language of a user’s query. CLIR systems give
multilingual users the possibility to express queries in any language, e.g., their native language,
and to obtain result documents in all languages they are familiar with. Since CLIR is not restricted
to collections in the query language more sources can be included in the retrieval process, and
the chance to fulfill a particular information need of a multilingual user is higher. Another use
case for CLIR techniques is cross-language plagiarism detection, where the query corresponds to
a suspicious document and the target collection is a reference corpus with original documents [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The Cross-Language Evaluation Forum, CLEF, provides an infrastructure for the evaluation
of information retrieval systems, both monolingual and cross-lingual. We participated in the
TEL@CLEF task of the CLEF 2009 ad-hoc track, which aims at the evaluation of systems to
retrieve relevant items from multilingual collections of library catalog records. The main challenges
of this task are the multilinguality and the sparsity of the dataset. We used two different CLIR
approaches to tackle this task; the paper in hand outlines and discusses these approaches and the
achieved results.</p>
      <p>
        The first approach is Cross-Language Explicit Semantic Analysis, CL-ESA, which is a
multilingual retrieval model to access cross-language similarity between text documents [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The CL-ESA
model exploits a document-aligned comparable corpus such as Wikipedia in order to map the
query and the documents into a common multilingual concept space [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. We also present a
formal definition and an alternative interpretation for the CL-ESA model, which is inspired by the
Generalized Vector Space Model, GVSM. Our view is mathematically equivalent to the original
idea of the CL-ESA model; it reveals how the computational effort for CL-ESA can be shifted
from the query phase to a preprocessing phase.
      </p>
      <p>In the second approach, called Cross Querying, each query is translated into all target
languages. The particular rankings are used in a combined fashion considering the most likely
language of the documents. The evaluation on the TEL@CLEF collections shows that both CLIR
approaches are able to outperform the monolingual baseline. In the bilingual subtask, querying
with a foreign language, Cross Querying achieves nearly the same or even higher results compared
to the monolingual subtask; the performance of the CL-ESA is lower compared to the monolingual
results.</p>
      <p>The paper is organized as follows. Section 2 describes the target collection used in the
TEL@CLEF task along with the evaluation procedure. Section 3 defines the general CL-ESA
model, our formalization, and details of the CL-ESA implementation employed in the
experiments. Section 4 presents the Cross Querying approach, Section 5 discusses the evaluation, and
Section 6 concludes with an outlook.
2</p>
    </sec>
    <sec id="sec-2">
      <title>TEL@CLEF dataset and Evaluation Procedure</title>
      <p>In this year’s TEL@CLEF task three target collections, provided by The European Library1,
TEL, are used. The collections are labeled BL, ONB, and BNF, and mainly contain information
in English, German, and French respectively (see Table 1). The collections are comprised of library
catalog records, referring to different types of items such as articles, books, or videos. The data is
provided in structured form and represented in XML. Each library catalog record has several fields
containing meta information and content information that describe the particular item. Typical
meta information fields are author, rights, or publisher, and typical content information fields
are title, description, subject, or alternative. In our experiments we focus on the content
information fields. A major difficulty is the sparsity of the available information: for many records
only few fields are given.</p>
      <p>The user’s information need is specified by 50 topics that are provided by CLEF in the three
main languages of the target collections, namely English, German, and French. A topic consists of
two fields: a title, containing 2-4 keywords, and a description, containing 1-2 sentences that
specify the item of interest in greater detail. The topics are used to construct the queries.</p>
      <p>
        The TEL@CLEF task is divided into a monolingual and a bilingual subtask. The aim in both
subtasks is to retrieve documents (library catalog records) from the target collections, which are
most relevant to a query; for each query the results are submitted as a ranked list of documents.
In the monolingual subtask the language of the query and the main language of the collection
are the same, while in the bilingual subtask the language of the query is different from the main
language of the collection. We submitted runs for both subtasks and for all three languages.
1The European Library: http://www.theeuropeanlibrary.org/.
Cross-Language Explicit Semantic Analysis, CL-ESA, is a generalization of the Explicit Semantic
Analysis, ESA [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and was proposed by Potthast et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This section presents a formal definition
of the CL-ESA model that reveals its close connection to the Generalized Vector Space Model,
GVSM [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: the ESA model and the GVSM can be transformed into each other [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. It follows
immediately that this is also true for the CL-ESA model and the cross-lingual extension of the
Generalized Vector Space Model, CL-GVSM [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
3.1
      </p>
      <sec id="sec-2-1">
        <title>Formal Definition</title>
        <p>
          Let di be a real-world document written in language Li, and let di be a bag-of-word-based
representation of di, encoded as a vector of normalized term frequency weights over a universal term
vocabulary Vi. Vi contains all used terms for language Li. A set Di of document representations
defines a term-document matrix ADi , where each column in ADi corresponds to a vector di ∈ Di.
Definition 1 (ESA Representation [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]) Let Di∗ be a collection of index documents written in
language Li. The ESA representation diESA of a document di with representation di is defined
as follows:
diESA = ATDi∗ · di,
(1)
where AT designates the matrix transpose of A.
        </p>
        <p>The rationale of this definition becomes clear if one considers that the weight vectors di∗ ∈ Di∗
and di are normalized: ||di∗|| = ||di|| = 1, for each di∗ ∈ Di∗. Hence, each entry in the ESA
representation diESA of a document di is the cosine similarity between di and some vector di∗ ∈ Di∗.
Put another way, di is compared to each index document in Di∗, and diESA is comprised of the
respective cosine similarities.</p>
        <p>Definition 2 (CL-ESA Similarity) Let L = {L1, . . . , Lk} denote a set of natural languages,
and let D∗ = {D1∗, . . . , Dk∗} be a set of index collections where each Di∗ ∈ D∗ is a list of index
documents written in language Li ∈ L. D∗ is a document-aligned comparable corpus, i.e., for each
language Li ∈ L the n-th index document in Di∗ ∈ D∗ describes the same concept. The CL-ESA
similarity, ϕCL−ESA(qj , di), between a query qj in language Lj and a document di in language Li
is computed as cosine similarity ϕ of the ESA representations of qj and di:
ϕCL−ESA(qj , di) = ϕ(qj ESA, diESA) = ϕ(ATDj∗ · qj , ATDi∗ · di)
(2)</p>
        <p>
          Due to the alignment of the index collections Dj∗ and Di∗ the ESA representations of qj
and di are comparable. Definition 2 is equivalent to the definition of the CL-GSVM
similarity ϕCL−GVSM (qj , di) given in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], which means that, in analogy to [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], the CL-ESA model and
the CL-GVSM can be directly transformed into each other:
        </p>
        <p>ϕCL−ESA(qj, di) = ϕ(ATDj∗ · qj , ATDi∗ · di) = ϕCL−GVSM (qj, di)
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Alternative Interpretation</title>
        <p>The original idea of the CL-ESA model is to map both query and documents into a multilingual
concept space, as it is expressed in Equation 2. Note that Equation 2 can be rearranged as follows:
ϕCL−ESA(qj, di) = ϕ(ATDj∗ · qj, ATDi∗ · di) = qjT · ADj∗ · ATDi∗ · di
In particular, the matrix ADj∗ · ATD∗ = Gj,i can be computed in advance since it is independent
i
from a particular qj or di. Hence:</p>
        <p>ϕCL−ESA(qj, di) = qjT · Gj,i · di</p>
        <p>The rationale of Equation 5 becomes apparent if one recognizes Gj,i = ADj∗ · ATDi∗ as |Vj | × |Vi|
term co-occurrence matrix. The n-th row in ADj∗ corresponds to the distribution of the n-th
term tn ∈ Vj over the index documents in Dj∗; likewise, the m-th row in ADi∗ corresponds to the
∗
distribution of the m-th term tm ∈ Vi over the index documents in Di . Recall that the index
documents in Dj∗ and Di∗ are aligned. I.e., the value in the n-th row and the m-th column of Gj,i
quantifies the similarity between the distributions of tj and ti given the concepts described by the
index documents in Dj∗ and D∗.</p>
        <p>i</p>
        <p>The CL-ESA similarity computation of Equation 5 can be viewed in two ways:
(i) As a translation of the query representation qj into the space of the document
representation di: ϕCL−ESA(qj , di) = (qjT · Gj,i) · di, or,
(ii) as a translation of the document representation di into the space of the query
representation qj : ϕCL−ESA(qj , di) = qjT · (Gj,i · di).</p>
        <p>These views are different from the original idea of the CL-ESA model where both the query
representation and the document representation are mapped into a common multilingual concept
space (see Equation 2). From a mathematical standpoint Equation 2 and Equation 5 are
equivalent; however, implementing CL-ESA based on the alternative interpretation yields a considerable
runtime improvement in practical retrieval applications. Table 2 contrasts the interpretations and
the related runtime complexities. Here, we assume a closed retrieval situation where from a given
target collection Di in language Li the most similar documents to a query qj in language Lj are
desired. CLIR with CL-ESA is straightforward: computation of ϕCL−ESA(qj, di) for each di ∈ Di
and ranking by decreasing CL-ESA similarity.</p>
        <p>Under the original interpretation the ESA representations diESA of the documents di ∈ Di
can be computed in advance. At retrieval time the query is mapped into the concept space in
O(l · |D∗|), where l denotes the number of query terms. The computation of the cosine similarity
(3)
(4)
(5)
between the ESA representations qj ESA and diESA requires O(|D∗|). Under the alternative
interpretation the matrix Gj,i can be computed in advance. Note that in practical applications
l ≪ |D∗|, since a reasonable index collection size |D∗| is 10 000, which shows the substantial
performance improvement under the alternative interpretation and View (ii) .
3.3
In this subsection we describe implementation details of the CL-ESA model we used in our
submission. The best parameter setting was determined by analyzing unofficial experiments of the
TEL@CLEF 2008 dataset.</p>
        <p>Query and Document Construction. We use the original words of both topic fields, title and
description, as queries. The documents are constructed by merging the text of the three record
fields title, subject, and alternative. We assume that the language of these fields is the same
within one record; however, this assumption may be violated in some cases since the collections
contain multilingual records. Records without these fields are omitted in the experiments (see
Table 1).</p>
        <p>Index Collection. As index collection Wikipedia is employed. We restrict the multilinguality of our
model to the three main languages of the target collections: English, German, and French. Based
on a Wikipedia snapshot from March 2009 about 169 000 articles per language can be aligned and
fulfill several filter criteria, e.g., to contain more than 100 words or not to be a disambiguation or
redirection page. All articles are used as index documents.</p>
        <p>As term weighting schema tf · idf is used. Query and document words are stemmed using the
Snowball stemmers. To speed-up the CL-ESA similarity computation all values below a threshold
of ǫ = 0.025 are discarded.</p>
        <p>Language Detection. While the language of the queries is determined by the corresponding topics
the language of the documents is unknown since the collections are multilingual and no language
meta information is provided. In the experiments we resort to a simple “detection by stop words”
approach for the three main languages; if the detection fails the main language of the collection is
assumed.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Cross Querying</title>
      <p>Cross querying is a straightforward approach for CLIR systems. We subsume the fields of a topic
in one query which is translated in the other languages. With each of the translations we compute
a set of rankings by retrieving against each document field. The rankings are merged with respect
to their cosine similarities. Additionally, the scores are multiplied by a boosting constant.
Definition 3 (Cross Querying) Let L = {L1, . . . , Lk} denote a set of natural languages and
let F = {F1, . . . , Fk} denote a set of document fields. lang : D → L, lang(d) 7→ Li estimates the
language of a document d. d, q, and qLi are the representations of a document d, a query q and
the translation of q in language Li. Then the cross querying similarity, ϕCQ (q, d), of a query q
and a document d is defined as follows:
ϕCQ (q, d) = X</p>
      <p>Fi∈F</p>
      <p>X</p>
      <p>Li∈L,</p>
      <p>Li6=lang(d)
b · ϕ(qlang(d), dFi) +
ϕ(qLi , dFi ) ,
(6)
where ϕ is the cosine similarity and b the boosting constant.</p>
      <p>The name “Cross Querying” reflects the fact that |L| × |F | rankings are merged by querying
in each language in each field. The applied parameters are as follows:
Query and Document Construction. The words of both topic fields, title and description, are
used as queries and translated to each Li ∈ L, with L = {German, F rench, English}. The
selection of the document fields corresponds to title and subject.</p>
      <p>As term weighting schema tf · idf is used. Query and document words are stemmed using
the Snowball stemmers while stop words are removed. The queries are translated with Google
Translate; the boosting constant b is based on the unofficial evaluation on the TEL@CLEF 2008
dataset.</p>
      <p>Language Detection. In order to estimate the language of d with lang(d) we take the corpus
language of the associated evaluation run.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Results</title>
      <p>The results of the monolingual subtask and the bilingual subtask are shown in Figure 1 and
Figure 2 respectively.</p>
      <p>70
60
50
%
in40
n
o
i
ics 30
e
r
P20
10
0
70
60
50
%
in40
n
o
i
ics 30
e
r
P20
10
0</p>
      <sec id="sec-4-1">
        <title>Monolingual English</title>
        <p>Baseline
Cross Querying</p>
        <p>CL-ESA
CL-ESA-LD
0 10 20 30 40 50 60 70 80 90 100</p>
        <p>Recall in %</p>
      </sec>
      <sec id="sec-4-2">
        <title>Monolingual French</title>
        <p>Baseline
Cross Querying</p>
        <p>CL-ESA
CL-ESA-LD</p>
      </sec>
      <sec id="sec-4-3">
        <title>Monolingual German</title>
        <p>Baseline
Cross Querying</p>
        <p>CL-ESA</p>
        <p>CL-ESA-LD
70
60
50
%
in40
n
o
i
ics 30
e
r
P20
10
0
0 10 20 30 40 50 60 70 80 90 100</p>
        <p>Recall in %
English</p>
        <p>German</p>
        <p>French
Baseline
Cross Querying
CL-ESA
CL-ESA-LD
0 10 20 30 40 50 60 70 80 90 100</p>
        <p>Recall in %</p>
        <p>We submitted an additional baseline to the monolingual subtask using state-of-the-art retrieval
technology: since in this subtask the language of the topics is equal to the main language of the
target collection, the ranking is based on the cosine similarities of the tf ·idf -weighted bag-of-words
representations of the topics and the documents.</p>
        <p>Each plot in Figure 1 corresponds to one target collection and shows the baseline along with the
results achieved under Cross Querying, CL-ESA, and CL-ESA with automatic language detection,
CL-ESA-LD. Both Cross Querying and CL-ESA clearly outperform the baseline. The variation
between the two approaches is small, except for the German collection where Cross Querying
outperforms CL-ESA at low recall levels. At higher recall levels CL-ESA is better, which explains
a slightly higher mean average precision on the English and the French collections. Using
CLESA along with the automatic language detection improves the performance only for the French
collection, which indicates that this collection contains a larger fraction of non-French documents.</p>
        <p>In the bilingual subtask the language of the queries is different from the main language of
the target collection. Each plot in Figure 2 corresponds to one target collection that is queried
in the two other languages, using both Cross Querying and CL-ESA. For example, in the plot
50
%
in40
n
o
i
ics30
e
r
P20
10
0
70
60
50
%
in40
n
o
i
ics30
e
r
P20
10
0</p>
      </sec>
      <sec id="sec-4-4">
        <title>Bilingual English</title>
        <p>Cross Querying-de
Cross Querying-fr</p>
        <p>CL-ESA-de
CL-ESA-fr
0 10 20 30 40 50 60 70 80 90 100</p>
        <p>Recall in %</p>
      </sec>
      <sec id="sec-4-5">
        <title>Bilingual French</title>
        <p>Cross Querying-de
Cross Querying-en
CL-ESA-de
CL-ESA-en
60
50
%
in40
n
o
i
ics30
e
r
P20
10
0</p>
      </sec>
      <sec id="sec-4-6">
        <title>Bilingual German</title>
        <p>Cross Querying-en
Cross Querying-fr</p>
        <p>CL-ESA-en</p>
        <p>CL-ESA-fr
0 10 20 30 40 50 60 70 80 90 100</p>
        <p>Recall in %</p>
        <p>English German French
Cross Querying-en
Cross Querying-de
Cross Querying-fr
CL-ESA-en
CL-ESA-de
CL-ESA-fr
0 10 20 30 40 50 60 70 80 90 100</p>
        <p>Recall in %
“Bilingual English” the graph for “CL-ESA-de” shows the results of querying the English collection
with German topics using the CL-ESA. Cross Querying achieves nearly the same or even higher
results compared to the monolingual situation, whereas the performance of the CL-ESA is lower
in contrast to the monolingual results.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>The evaluation results for the TEL@CLEF task show that both CLIR approaches CL-ESA and
Cross Querying are able to outperform the monolingual baseline—though the absolute results are
still improvable. Furthermore, we have presented a formal definition and an alternative
interpretation for the CL-ESA model, which is interesting for real-world retrieval applications since it reveals
how the computational effort for CL-ESA can be shifted from the query phase to a preprocessing
phase.</p>
      <p>As for future work, CL-ESA and Cross Querying will benefit if more languages are taken into
account. Currently, German, English, and French are used, but the target collections comprise
more languages. For documents from other languages an inconsistent CL-ESA representation is
computed. CL-ESA also needs a reliable language detection mechanism in order to compute a
consistent representation; note that we used a rather simple approach in our experiments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Maik</given-names>
            <surname>Anderka</surname>
          </string-name>
          and
          <string-name>
            <given-names>Benno</given-names>
            <surname>Stein</surname>
          </string-name>
          .
          <article-title>The ESA Retrieval Model Revisited</article-title>
          . In Mark Sanderson, James Allan ChengXiang Zhai, Justin Zobel, and Javed A. Aslam, editors,
          <source>32th Annual International ACM SIGIR Conference</source>
          , pages
          <fpage>670</fpage>
          -
          <lpage>671</lpage>
          . ACM,
          <year>July 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Evgeniy</given-names>
            <surname>Gabrilovich</surname>
          </string-name>
          and
          <string-name>
            <given-names>Shaul</given-names>
            <surname>Markovitch</surname>
          </string-name>
          .
          <article-title>Computing Semantic Relatedness using Wikipedia-based Explicit Semantic Analysis</article-title>
          .
          <source>In Proceedings of The 20th International Joint Conference for Artificial Intelligence</source>
          , Hyderabad, India,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Martin</given-names>
            <surname>Potthast</surname>
          </string-name>
          , Benno Stein, and
          <string-name>
            <given-names>Maik</given-names>
            <surname>Anderka</surname>
          </string-name>
          .
          <article-title>A Wikipedia-Based Multilingual Retrieval Model</article-title>
          . In Craig Macdonald, Iadh Ounis, Vassilis Plachouras, Ian Ruthven, and Ryen W. White, editors,
          <source>30th European Conference on IR Research</source>
          , ECIR
          <year>2008</year>
          ,
          <article-title>Glasgow</article-title>
          , volume
          <volume>4956</volume>
          <source>LNCS of Lecture Notes in Computer Science</source>
          , pages
          <fpage>522</fpage>
          -
          <lpage>530</lpage>
          , Berlin Heidelberg New York,
          <year>2008</year>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Sorg</surname>
          </string-name>
          and
          <string-name>
            <given-names>Philipp</given-names>
            <surname>Cimiano</surname>
          </string-name>
          .
          <article-title>Cross-lingual information retrieval with explicit semantic analysis</article-title>
          .
          <source>In Working Notes for the CLEF 2008 Workshop</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>K. M. Wong</surname>
          </string-name>
          , Wojciech Ziarko, and
          <string-name>
            <surname>Patrick</surname>
            <given-names>C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Wong</surname>
          </string-name>
          .
          <article-title>Generalized vector spaces model in information retrieval</article-title>
          .
          <source>In SIGIR '85: Proceedings of the 8th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          , pages
          <fpage>18</fpage>
          -
          <lpage>25</lpage>
          , New York, NY, USA,
          <year>1985</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Yiming</given-names>
            <surname>Yang</surname>
          </string-name>
          , Jaime G. Carbonell, Ralf D.
          <string-name>
            <surname>Brown</surname>
            , and
            <given-names>Robert E.</given-names>
          </string-name>
          <string-name>
            <surname>Frederking</surname>
          </string-name>
          .
          <article-title>Translingual information retrieval: learning from bilingual corpora</article-title>
          .
          <source>Artif</source>
          . Intell.,
          <volume>103</volume>
          (
          <issue>1-2</issue>
          ):
          <fpage>323</fpage>
          -
          <lpage>345</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>