<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Query and Document Translation by Automatic Text Categorization: A Simple Approach to Establish a Strong Textual Baseline for ImageCLEFmed 2006</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julien Gobeill</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Henning Muller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Patrick Ruch</string-name>
          <email>patrick.ruchg@sim.hcuge.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University and Hospitals of Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we report on the fusion of simple retrieval strategies with thesaural resources in order to perform document and query translation for cross{language retrieval in a collection of medical cases. The collection contains textual and visual contents. In this paper, we focus on the textual contents of the collection, which contains documents in three languages: French, English and German. The fusion of visual and textual content will also be treated. Unlike most automatic categorization systems, which rely on training data in order to infer text{to{concept relationships, our approach can be applied with any controlled vocabulary and does not use any training data. For the 2006 ImageCLEFmed experiments we use the Medical Subject Headings (MeSH), a terminology maintained by the National Library of Medicine and which exists in a dozen languages. The basic idea consists of annotating every textual content of the collection (documents and queries) with a set of MeSH concepts using an automatic text categoriser. Thus, allowing an interlingual mapping between queries and documents. For tuning purposes, the system uses a sample of MEDLINE from the OHSUMED collection. Our results, con rmed that such a simple approach is competitive with best performing cross-language retrieval methods for such a collection. Several simple linear approaches were used to combine textual and visual features</p>
      </abstract>
      <kwd-group>
        <kwd>Image retrieval</kwd>
        <kwd>Text categorization</kwd>
        <kwd>multimodal retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Cross{Language Information Retrieval (CLIR) is increasingly relevant as network{based resources
become commonplace. In the medical domain it is of strategic importance in order to ll the gap
between clinical records, written in national languages and research reports massively written in
English. Images are also getting increasingly important and varied in the medical domain, and they
become available in digital form. Despite the fact that images are language{independent, they are
most often accompanied by textual notes in various languages and these textual notes can strongly
improve retrieval quality [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. There are several ways for handling CLIR. Historically, the most
traditional approach to IR in general and to multilingual retrieval in particular, uses a controlled
vocabulary for indexing and retrieval. In this approach, a librarian selects for each document a few
descriptors taken from a closed list of authorised terms. A good example of such a human indexing
is found in the MEDLINE database, where records are manually annotated with Medical Subject
Headings (MeSH). Ontological relations (synonyms, related terms, narrower terms, broader terms)
can be used to help choose the right descriptors, and solve the sense problems of synonyms and
homographs. The list of authorised terms and semantic relations between them are contained in
a thesaurus. A problem remains, however, since concepts expressed by one single term in one
language sometime are expressed by distinct terms in another. We can observe that terminology{
based CLIR is a common approach in well{delimited elds for which multilingual thesauri already
exist (not only in medicine but also in the legal domain, energy, etc.) as well as in multinational
organisations or countries with several o cial languages. This controlled vocabulary approach is
often associated with Boolean{like engines, and it gives acceptable results but prohibits precise
queries that cannot be expressed with these authorised keywords. The two main problems are:
it can be di cult for users to think in terms of a controlled vocabulary. Therefore, the use of
these systems { like most Boolean{supported engines | is often performed by professionals
rather than general users;
this retrieval method ignores the free{text portions of documents during indexing.
      </p>
      <p>
        A detailed task description on the medical image retrieval task can be found in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The
non{medical tasks of ImageCLEF are described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
1.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Translation{based Approaches</title>
      <p>
        A second approach to multilingual interrogation is to use existing machine translation (MT)
systems to automatically translate the queries [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] or even the entire textual database [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] from
one language to another, thereby transforming the CLIR problem into a mono{lingual information
retrieval (MLIR) problem.
      </p>
      <p>This kind of method would be satisfactory if current MT systems did not make errors. A certain
amount of syntactic error can be accepted without disturbing results of information retrieval
systems but MT errors in translating concepts can prevent relevant documents, indexed on the
missing concepts, from being found. For example, if the word traitement in French is translated by
processing instead of prescription, the retrieval process would yield wrong results. This drawback
is limited in MT systems that use huge transfer lexicons of noun phrases by taking advantage of
frequent co{locations to help disambiguation but in any collection of text ambiguous nouns will
still appear as isolated noun phrases untouched by this approach.
1.2</p>
    </sec>
    <sec id="sec-3">
      <title>Using Parallel Resources</title>
      <p>
        A third approach receiving increasing attention is to automatically establish associations between
queries and documents independent of language di erences. Seminal researches were using latent
semantic indexing [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The general strategy when working with parallel or comparable texts is
the following: if some documents are translated into a second language these documents can be
observed both in the subspace related to the rst language and the subspace related to the second
one; using a query expressed in the second language, the most relevant documents in the translated
subset are extracted (usually using a cosine measure of proximity). These relevant documents are
in turn used to extract close untranslated documents in the subspace of the rst language. This
approach use implicit dependency links and co{occurrences that better approximate the notion of
concept. Such a strategy has been tested with success on the English-French language pair using a
sample of the Canadian Parliament bilingual corpus. It is reported that for 92% of the English text
documents the closest document returned by the method was its correct French translation. Such
an approach presupposes that the sample used for training is representative of the full database,
and that su cient parallel/comparable corpora are available or acquired.
      </p>
      <p>
        Other approaches are usually based on bilingual dictionaries and terminologies, sometimes
combined with parallel corpora. These approaches attempt to infer a word by word transfer
function: they typically begin by deriving a translation dictionary, which is then applied to query
translation. Of interest for comparison with our experiments, the study reported by Eichmann
and al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] uses the OHSUMED collection and relies on a transfer lexicon built from the Uni ed
Medical Language System (UMLS1). Finally, [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] described a system which combines thesaurus{
based approaches and machine translation for translating queries in a monolingual collection.
      </p>
      <p>
        To synthesise, we can consider that the performance of CLIR systems typically ranges between
60% and 90% of the corresponding monolingual run [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. CLIR ratios above 100% have been
reported [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], however, such results were obtained by computing a weak monolingual baseline.
2
      </p>
      <sec id="sec-3-1">
        <title>Data and Strategies</title>
        <p>The 2006 imageCLEF collection contains a set of 50'000 medical images accompanied by textual
reports (mainly pathology and radiology) that are all well-structure in XML. Some reports describe
an entire case containing several images and other reports describe a single image. In this article
we focus on experiments made on the textual part. A short description will also describe the visual
and multi{modal retrieval parts. The text of the collection contains over 40'000 textual documents
in three languages: French, English and German. There are 30 English queries, translated into
French and German. Because the document collection is multilingual, we decided to map each
document and each query to a set of MeSH concepts, thus transforming a fairly usual text retrieval
task into a category retrieval task.</p>
        <p>
          Soergel describes a general framework for the use of multilingual thesauri in CLIR [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ], noting
that a number of operational European systems employ multilingual thesauri for indexing and
searching. However, except for very early work [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], there has been little empirical evaluation of
multilingual thesauri in the context of free{text CLIR, particularly when compared to dictionary
and corpus{based methods. This may be due to the cost of constructing multilingual thesauri, but
this cost is unlikely to be more than that of creating bilingual dictionaries or even realistic parallel
collections. It seems that multilingual thesauri can be built quite e ectively by merging existing
monolingual thesauri, as shown by the current development of the Uni ed Medical Language
System (UMLS) or the SemanticMining Multilingual Lexicon[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Our approach to CLIR in MEDLINE exploits the UMLS resources and its multilingual
components. The core technical component of our cross{language engine is an automatic text categoriser,
which associates a set of MeSH terms to any input text. The experimental design is the following:
1. each document and all queries of the imageCLEF collection are annotated by our automatic
text categoriser, which contains MeSH in French, English, and German;
2. each query is annotated by three MeSH categories and we evaluate the impact of varying the
number of categories for the document collection: three, ve and eight categories are tested;
3. the MeSH annotated imageCLEF document collection is indexed using a standard engine.
2.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>MeSH{driven Text Categorization</title>
      <p>Automatic text categorization has been studied largely and has led to an impressive amount of
papers. A partial list2 of machine learning approaches applied to text categorization includes naive
1See http://umlsks.nlm.nih.gov.</p>
      <p>2See http://faure.iei.pi.cnr.it/~fabrizio/ for an updated bibliography.</p>
      <p>
        Bayes [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], k{nearest neighbours [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], boosting [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], and rule{learning algorithms [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. However,
most of these studies apply text classi cation to a small set of classes; usually a few hundred, as
in the Reuters collection [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In comparison to this our system is designed to handle large class
sets [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]: retrieval tools used are only limited by the size of the inverted le, but 105 6 documents
is still a modest range 3.
      </p>
      <p>Our approach is data{poor because it only demands a small collection of annotated texts
for ne tuning: instead of inducing a complex model using large training data, our categoriser
indexes the collection of MeSH terms as if they were documents and then it treats the input
as if it was a query to be ranked regarding each MeSH term. The classi er is tuned by using
English abstracts and English MeSH terms. Then, we apply the system on the ImageCLEFmed
collection. For tuning the categoriser, the top 15 returned terms are selected because it is the
average number of MeSH terms per abstract in the OHSUMED collection. When applied on the
ImageCLEFmed collection, the number of categories to be attached to every document will be an
important parameter.
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Collection and Metrics</title>
      <p>
        The mean average precision (MAP): is the main measure for evaluating ad hoc retrieval tasks (for
both monolingual and bilingual runs). Following [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], we also use this measure to tune the automatic
text categorization system. We tune the categorization system on a small set of OHSUMED
abstracts: 1200 randomly selected abstracts were used to select the weighting parameters of the
vector space classi er and the best combination of these parameters with the regular expression{
based classi er.
2.3
      </p>
    </sec>
    <sec id="sec-6">
      <title>Visual retrieval techniques</title>
      <p>
        The technology used for the visual retrieval of images is mainly taken from the Viper 4 project.
Much information about this system is available [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Outcome of the Viper project is the GNU
Image Finding Tool, GIFT 5. This tool is open source and can be used by other participants of
ImageCLEF. A ranked list of visually similar images for every query topic was made available for
participants and will serve as a baseline to measure the quality of submissions. Feature sets used
by GIFT are:
      </p>
      <p>Local color features at di erent scales by partitioning the images successively into four
equally sized regions (four times) and taking the mode color of each region as a descriptor;
global color features in the form of a color histogram, compared by a simple histogram
intersection;
local texture features by partitioning the image and applying Gabor lters in various scales
and directions, quantised into 10 strengths;
global texture features represented as a simple histogram of responses of the local Gabor
lters in various directions and scales.</p>
      <p>
        A particularity of GIFT is that it uses many techniques well{known from text retrieval. Visual
features are quantised and the feature space is similar to the distribution of words in texts. A
simple tf/idf weighting is used and the query weights are normalised by the results of the query
itself. The histogram features are compared based on a histogram intersection [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        3In text categorization based on learning methods, the scalability issue is twofold: it concerns both the ability
of these data{driven systems to work with large concept sets, and their ability to learn and generalise regularities
for rare events: [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] shows how the frequency of concepts in the collection is a major parameter for learning systems.
4http://viper.unige.ch/
5http://www.gnu.org/software/gift/
      </p>
      <p>System or
parameters</p>
      <p>
        RegEx
lnc.atn
anc.atn
ltc.atn
ltc.lnn
.1601
.1421
.1418
.1341
.111
In this section, we present the basic classi ers and their combination for the categorization task.
Three main modules constitute the skeleton of our system: the regular expression (RegEx)
component, the vector space (VS) component, and the visual retrieval part. Each of the basic classi ers
implements known approaches to document retrieval. The rst tool is based on a regular
expression pattern matcher [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], it is expected to perform well when applied on very short documents
such as keywords: MeSH terms do not contains more than 5 tokens. The second classi er is based
on a vector space engine6. This second tool is expected to provide high recall in contrast to the
regular expression{based tool, which should privilege precision. The former component uses
tokens as indexing units and can be merged with a thesaurus, while the latter uses stems (Porter).
Table 1 shows the results of each classi er.
      </p>
      <p>
        Regular expressions and MeSH thesaurus. The regular expression search tool is applied
on the canonic MeSH collection augmented with the MeSH thesaurus (120'020 synonyms). In this
system, string normalisation is mainly performed by the MeSH terminological resources when the
thesaurus is used. Indeed, the MeSH provides a large set of related terms, which are mapped to a
unique MeSH representative in the canonic collection. The related terms gather morpho-syntactic
variants, strict synonyms, and a last class of related terms, which mixes up generic and speci c
terms: for example, Inhibition is mapped to Inhibition (Psychology). The system cuts the
abstract into 5{token{long phrases and moves the window through the abstract: the edit{distance
is computed between each of these 5 token sequences and each MeSH term. Basically, the
manually crafted nite{state automata allow two insertions or one deletion within a MeSH term, and
ranks the proposed candidate terms based on these basic edit operations: insertion costs 1, while
deletion costs 2. The resulting pattern matcher behaves like a term proximity scoring system [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ],
but is restricted to a 5{token matching window.
      </p>
      <p>Vector space classi er. The vector space module is based on a general IR engine with the
tf.idf 7 weighting schema. The engine uses a list of 544 stop words.</p>
      <p>
        As for setting the weighting factors, we observed that cosine normalisation was especially
effective for our task. This is not surprising, considering the fact that cosine normalisation performs
well when documents have a similar length [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. As for the respective performance of each basic
classi er, table 1 shows that the RegEx system performs better than any tf.idf schema used by
the VS engine, so the pattern matcher provides better results than the vector space engine for
automatic text categorization. However, we also observe in table 1 that the VS system gives
better precision at high ranks (P recisionat Recall=0 or mean reciprocal rank ) than the RegEx system:
this di erence suggests that merging the classi ers could be e ective. The idf factor also seems
to be an important parameter. As shown in table 1 the four best weighting schemas use the idf
76TWhee uesaesytIhRe eSnMgiAneR,Tanrdeptrheeseanuttaotmionatifcorcaetxepgroersisziantgionstatotiosltkicitalarweeaigvhatiilnabglefaocntotrhs:e aautfohromr'aslpdaegsecsr.iption can be
found in [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
factor. This observation suggests that even in a controlled vocabulary, the idf factor is able to
discriminate between content{ and non{content{bearing features (such as syndrome and disease).
      </p>
      <p>
        Classi er fusion. The hybrid system combines the regular expression classi er with the
vector{space classi er. Unlike [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] we do not merge our classi ers by linear combination, because
the RegEx module does not return a scoring consistent with the vector space system. Therefore
the combination does not use the RegEx's edit distance, and instead it uses the list returned by
the vector space module as a reference list (RL), while the list returned by the regular expression
module is used as boosting list (BL), which serves to improve the ranking of terms listed in RL.
A third factor takes into account the length of terms: both the number of characters (L1) and
the number of tokens (L2, with L2 &gt; 3) are computed, so that long and compound terms, which
appear in both lists, are favoured over single and short terms. We assume that the reference list
has good recall, and we do not set any threshold on it. For each concept t listed in the RL, the
combined Retrieval Status Value (cRSV , equation 1) is:
cRSVt =
      </p>
      <p>RSVV S(t) Ln(L1(t) L2(t) k) if t 2 BL,
RSVV S(t) otherwise.
(1)</p>
      <p>The value of the k parameter is set empirically. Table 2 shows that the optimal tf.idf
parameters (lnc.atn) for the basic VS classi er does not provide the optimal combination with RegEx.
Measured by MAP. The optimal combination is obtained with ltc.lnn settings (.1818) 8, whereas
atn.ntn maximises the P recisionat Recall=0 (.9143).
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Cross{Language Categorization and Indexing</title>
      <p>
        To translate the ImageCLEFmed textual content (queries and documents), we transform the
English MeSH mapping tool described above, attributes MeSH terms to English abstracts. Thus,
the English, French, and German version of the MeSH are simply merged in the categoriser. We
use the weighting schema and system combination (ltc.lnn + RegEx) as described above. Then,
the annotated collection is indexed using the vector{space engine used by the categoriser. For the
document indexing, we rely on weighting schemas based on pivoted normalisation: because the
documents have a very variable length in the collection such a factor can be important. A slightly
modi ed version of dtu.dtn [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which has shown some e ectiveness for the TREC Genomics, is
used for full{text indexing and retrieval. The English stop word list is merged with a French and
a German stop word list. Porter stemming is used for all documents.
3.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>Visual and multimodal retrieval</title>
      <p>
        For the visual retrieval, one quantisation of four grey levels was used that has shown to be e cient
[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. To combine visual and textual runs we choose English as the main language and a number of
8For the augmented term frequency factor (noted a, which is de ned by the function
value of the parameters is = = 0:5.
+
(tf=max(tf)), the
ve terms. Visual inspection of some results indicated that this might lead to good results. The
combination is simply done linearly by normalising the output of visual and textual results and
then adding them up in various ratios. A second approach for a multimodal combination was to
take the results from the visual retrieval side and increase the value of those results in the rst
1000 images that also appear in the visual results by simple re{ranking.
4
      </p>
      <sec id="sec-8-1">
        <title>Results and Discussion</title>
        <p>We submitted several runs including runs combining textual and visual features. The visual runs
have a low overall performance but do perform well on the visual topics. The mixed runs had a
problem in the combination part and are in large part broken. The text retrieval was based on
the cases and thus needed to be expanded towards
4.1</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Textual retrieval results</title>
      <p>
        The runs were generated with the very same strategy but using respectively, eight, ve and three
MeSH categories to annotate the document collection. The runs were generated automatically
and do not use visual features. For each of them, queries were expanded using the top three
MeSH categories provided by the categoriser. Following previous experiments dedicated to query
translation [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], the optimal number of categories for query translation is around two or three.
      </p>
      <p>Run
GE-8EN
GE-5EN
GE-3EN</p>
      <p>MAP
0.2255
0.1967
0.1913</p>
      <p>Run
GE-8DE
GE-5DE
GE-3DE</p>
      <p>MAP
0.0574
0.0416
0.0433</p>
      <p>Run
GE-8FR
GE-5FR
GE-6FR</p>
      <p>MAP
0.0417
0.0346
0.0323</p>
      <p>Evaluations are computed by retrieving the rst 1'000 documents for each query. In Table 3,
we observe that the maximum MAP is reached when eight MeSH terms are selected per document.
This result suggests that a large number of categories can be selected to annotate a document,
although it must be observed that the precision of the system is low beyond the top one or two
categories. It means that annotating a document with several potentially irrelevant concepts does
not hurt the matching power of our interlingual concepts! This result is somehow consistent with
known observations made on query expansion: a certain number of inappropriate expansion is
acceptable and still can improve retrieval e ectiveness of modern search engines. This statement
should be emphasised in the case of the ImageCLEFmed collection, which can be regarded as a
small collection (about 40'000 cases for 50'000 images), where recall plays a more important role
than in larger collections, which can rely on information redundancy with a less important role
for recall.</p>
      <p>The English retrieval results were the second best group results of all participants. For
languages other than English it seems to be much harder to obtain very good results as the majority
of documents in the collection is in English.
4.2</p>
    </sec>
    <sec id="sec-10">
      <title>Visual and multi-modal runs</title>
      <p>Table 4 shows the results for our visual run and the best mixed runs. The visual run is performance{
wise in the middle of the submissions and the best purely visual runs are approximately 30% better.
GIFT performance better in early precision than other systems with a higher MAP. For the visual
topics the results are very satisfactory whereas semantic topics do not perform well.</p>
      <p>One big problem shows up in all the mixed runs submitted. They are not as good as the
underlying textual runs even when only a very small percentage of visual information is used. One
problem that might have caused this is the use of a wrong le for the textual runs. The English</p>
      <p>Run
GE-GIFT
GE-vt10
GE-vt20
runs perform much better than French and German runs, and so mixing up these two can cause
such trouble. We need to further investigate into this to nd the reasons and allow for better
multimodal image retrieval.
5</p>
      <sec id="sec-10-1">
        <title>Conclusion and Future Work</title>
        <p>We have presented a cross language information retrieval engine for the ImageCLEFmed image
retrieval task, which capitalises from the availability of a multilingual controlled vocabulary to
translate user requests and documents. The system relies on a text categoriser, which maps
queries into a set of prede ned concepts. For the ImageCLEFmed 2006 collection, optimal
precision is obtained when selecting three MeSH categories per query and eight MeSH categories per
documents, whereas a larger number leads to best MAP. Visual retrieval shows to work well on the
visual topics but fails on the semantic topics. Further experiments are needed to determine the
best number of concepts. The use of a dynamic threshold will be evaluated in the future. Another
problem is the combination of visual and textual features that really needs further analysis beyond
simple linear combinations.</p>
      </sec>
      <sec id="sec-10-2">
        <title>Acknowledgements</title>
        <p>The study has also been partially supported by the Swiss National Foundation (Grants 3200{
065228 and 205321{109304/1), the European Union (SemanticMining Network of Excellence,
INFS{CT{2004{507505) via an OFES Grant (No 03.0399, cf. http://www.genisis.ch/~natlang/semm/).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C</given-names>
            <surname>Apte</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F</given-names>
            <surname>Damerau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S</given-names>
            <surname>Weiss</surname>
          </string-name>
          .
          <article-title>Automated learning of decision rules for text categorization</article-title>
          .
          <source>ACM Transactions on Information Systems (TOIS)</source>
          ,
          <volume>12</volume>
          (
          <issue>3</issue>
          ):
          <volume>233</volume>
          {
          <fpage>251</fpage>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A</given-names>
            <surname>Aronson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Humphrey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Lin</surname>
          </string-name>
          , H Liu,
          <string-name>
            <given-names>P</given-names>
            <surname>Ruch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L</given-names>
            <surname>Tanabe</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J</given-names>
            <surname>Wilbur</surname>
          </string-name>
          .
          <article-title>Fusion of Knowledge-intensive and Statistical Approaches for Retrieving and Annotating Textual Genomics Documents</article-title>
          .
          <source>In TREC</source>
          <year>2005</year>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R</given-names>
            <surname>Baud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Nystrom</surname>
          </string-name>
          ,
          <string-name>
            <surname>L Borin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S</given-names>
            <surname>Schulz</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P</given-names>
            <surname>Zweigenbaum</surname>
          </string-name>
          .
          <article-title>Interchanging lexical information for a multilingual dictionary</article-title>
          .
          <source>AMIA Symposium Proceedings</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P</given-names>
            <surname>Clough</surname>
            , M Grubinger
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Deselaers</surname>
          </string-name>
          ,
          <article-title>A Hanbury, and</article-title>
          <string-name>
            <given-names>H.</given-names>
            <surname>Mu</surname>
          </string-name>
          <article-title>ller. Overview of the ImageCLEF 2006 photo retrieval and object annotation tasks</article-title>
          .
          <source>In CLEF working notes</source>
          , Alicante, Spain, Sep.
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M</given-names>
            <surname>Davis.</surname>
          </string-name>
          <article-title>Free resources and advanced alignment for cross-language text retrieval</article-title>
          .
          <source>In In proceedings of The Sixth Text Retrieval Conference (TREC6)</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S</given-names>
            <surname>Dumais</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Letsche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Littman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T</given-names>
            <surname>Landauer</surname>
          </string-name>
          .
          <article-title>Automatic cross-language retrieval using latent semantic indexing</article-title>
          . In D Hull and
          <string-name>
            <given-names>D</given-names>
            <surname>Oard</surname>
          </string-name>
          , editors,
          <source>AAAI Symposium on CrossLanguage Text and Speech Retrieval</source>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D</given-names>
            <surname>Eichmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and P</given-names>
            <surname>Srinivasan</surname>
          </string-name>
          .
          <article-title>Cross{language information retrieval with the UMLS metathesaurus</article-title>
          .
          <source>In SIGIR Conference</source>
          , pages
          <volume>72</volume>
          {
          <fpage>80</fpage>
          ,
          <string-name>
            <surname>Melbourne</surname>
          </string-name>
          , Australia,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P</given-names>
            <surname>Hayes and S Weinstein</surname>
          </string-name>
          .
          <article-title>A system for content-based indexing of a database of news stories</article-title>
          .
          <source>Proceedings of the Second Annual Conference on Innovative Applications of Intelligence</source>
          ,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>L</given-names>
            <surname>Larkey</surname>
          </string-name>
          and
          <string-name>
            <given-names>W</given-names>
            <surname>Croft</surname>
          </string-name>
          .
          <article-title>Combining classi ers in text categorization</article-title>
          .
          <source>In SIGIR</source>
          , pages
          <volume>289</volume>
          {
          <fpage>297</fpage>
          . ACM Press, New York, US,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>U</given-names>
            <surname>Manber and S Wu. GLIMPSE</surname>
          </string-name>
          :
          <article-title>A tool to search through entire le systems</article-title>
          .
          <source>In Proceedings of the USENIX Winter 1994 Technical Conference</source>
          , pages
          <volume>23</volume>
          {
          <fpage>32</fpage>
          , San Fransisco CA USA,
          <volume>17</volume>
          -
          <fpage>21</fpage>
          1994.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A</given-names>
            <surname>McCallum</surname>
          </string-name>
          and
          <string-name>
            <given-names>K</given-names>
            <surname>Nigam</surname>
          </string-name>
          .
          <article-title>A comparison of event models for naive bayes text classi cation</article-title>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J</given-names>
            <surname>McCarley</surname>
          </string-name>
          .
          <article-title>Should we translate the documents or the queries in cross-language information retrieval</article-title>
          .
          <source>ACL</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Deselaers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Lehmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P</given-names>
            <surname>Clough</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W</given-names>
            <surname>Hersh</surname>
          </string-name>
          .
          <article-title>Overview of the ImageCLEFmed 2006 medical retrieval and annotation tasks</article-title>
          .
          <source>In CLEF working notes</source>
          , Alicante, Spain, Sep.
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Henning</given-names>
            <surname>Mu</surname>
          </string-name>
          <article-title>ller, Antoine Geissbuhler, and Patrick Ruch</article-title>
          .
          <source>ImageCLEF</source>
          <year>2004</year>
          :
          <article-title>Combining image and multi{lingual search for medical image retrieval</article-title>
          .
          <source>In Cross Language Evaluation Forum (CLEF</source>
          <year>2004</year>
          ), Springer Lecture Notes in Computer Science (LNCS), Bath, England,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15] Henning Muller, Nicolas Michoux, David Bandon,
          <string-name>
            <given-names>and Antoine</given-names>
            <surname>Geissbuhler</surname>
          </string-name>
          .
          <article-title>A review of content{based image retrieval systems in medicine { clinical bene ts and future directions</article-title>
          .
          <source>International Journal of Medical Informatics</source>
          ,
          <volume>73</volume>
          :1{
          <fpage>23</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D</given-names>
            <surname>Oard</surname>
          </string-name>
          and
          <string-name>
            <given-names>P</given-names>
            <surname>Hackett</surname>
          </string-name>
          .
          <article-title>Document translation for cross-language text retrieval at the university of Maryland</article-title>
          . In
          <source>In Proceedings of The Sixth Text Retrieval Conference (TREC6)</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Y</given-names>
            <surname>Rasolofo and J Savoy</surname>
          </string-name>
          .
          <article-title>Term proximity scoring for keyword-based retrieval systems</article-title>
          .
          <source>In ECIR</source>
          , pages
          <volume>101</volume>
          {
          <fpage>116</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>P</given-names>
            <surname>Ruch</surname>
          </string-name>
          .
          <article-title>Using contextual spelling correction to improve retrieval e ectiveness in degraded text collections</article-title>
          .
          <source>COLING</source>
          <year>2002</year>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P</given-names>
            <surname>Ruch</surname>
          </string-name>
          .
          <article-title>Query translation by text categorization</article-title>
          .
          <source>COLING</source>
          <year>2004</year>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>P</given-names>
            <surname>Ruch</surname>
            , R Baud
          </string-name>
          ,
          <article-title>and A Geissbuhler. Learning-Free Text Categorization</article-title>
          .
          <source>LNAI 2780</source>
          , pages
          <fpage>199</fpage>
          {
          <fpage>208</fpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G</given-names>
            <surname>Salton</surname>
          </string-name>
          .
          <article-title>Automatic processing of foreign language documents</article-title>
          .
          <source>JASIS</source>
          ,
          <volume>21</volume>
          (
          <issue>3</issue>
          ):
          <volume>187</volume>
          {
          <fpage>194</fpage>
          ,
          <year>1970</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>R</given-names>
            <surname>Schapire</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y</given-names>
            <surname>Singer</surname>
          </string-name>
          .
          <article-title>BoosTexter: A boosting-based system for text categorization</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>39</volume>
          (
          <issue>2</issue>
          /3):
          <volume>135</volume>
          {
          <fpage>168</fpage>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>P</given-names>
            <surname>Scha</surname>
          </string-name>
          <article-title>uble and P Sheridan. Cross-language information retrieval (CLIR) track overview</article-title>
          . In
          <source>In Proceedings of The Sixth Text Retrieval Conference (TREC6)</source>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>A</given-names>
            <surname>Singhal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C</given-names>
            <surname>Buckley</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M</given-names>
            <surname>Mitra</surname>
          </string-name>
          .
          <article-title>Pivoted document length normalization</article-title>
          .
          <source>ACM-SIGIR</source>
          , pages
          <volume>21</volume>
          {
          <fpage>29</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>D</given-names>
            <surname>Soergel</surname>
          </string-name>
          .
          <article-title>Multilingual thesauri in crosslanguage text and speech retrieval</article-title>
          . In D Hull and
          <string-name>
            <given-names>D</given-names>
            <surname>Oard</surname>
          </string-name>
          , editors,
          <source>AAAI Symposium on Cross-Language Text and Speech Retrieval</source>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>David</given-names>
            <surname>McG. Squire</surname>
          </string-name>
          , Wolfgang Muller, Henning Muller, and Thierry Pun.
          <article-title>Content{based query of image databases: inspirations from text retrieval</article-title>
          .
          <source>Pattern Recognition Letters (Selected Papers from The 11th Scandinavian Conference on Image Analysis SCIA '99)</source>
          ,
          <volume>21</volume>
          (
          <fpage>13</fpage>
          - 14):
          <volume>1193</volume>
          {
          <fpage>1198</fpage>
          ,
          <year>2000</year>
          .
          <string-name>
            <given-names>B.K.</given-names>
            <surname>Ersboll</surname>
          </string-name>
          , P. Johansen, Eds.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Michael</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Swain</surname>
          </string-name>
          and
          <string-name>
            <surname>Dana H. Ballard</surname>
          </string-name>
          . Color indexing.
          <source>International Journal of Computer Vision</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <volume>11</volume>
          {
          <fpage>32</fpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>J</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A</given-names>
            <surname>Fraser</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R</given-names>
            <surname>Weischedel</surname>
          </string-name>
          .
          <article-title>Cross-lingual retrieval at bbn</article-title>
          .
          <source>In TREC</source>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Y</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>An evaluation of statistical approaches to text categorization</article-title>
          .
          <source>Journal of Information Retrieval</source>
          ,
          <volume>1</volume>
          :
          <fpage>67</fpage>
          {
          <fpage>88</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>