<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CoLesIR at CLEF 2007: from English to French via Character N -Grams</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jesus Vilares</string-name>
          <email>vilares@uvigo.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Campus de Elvin~a</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael P. Oakes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>St. Peter's Campus</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuel Vilares</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Campus As Lagoas s/n</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer Science</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computing</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of A Corun~a</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Sunderland</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Vigo</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1956</year>
      </pub-date>
      <abstract>
        <p>This work is an extension of our proposal originally presented in CLEF 2006, which, unfortunately, could not be ready on time for the workshop. We describe here a knowledge-light approach for query translation in Cross-Language Information Retrieval systems. This proposal itself can be considered as an extension of the previous work of the Johns Hopkins University Applied Physics Lab, preserving its advantages but avoiding its main drawbacks. As in their original proposal, our work is based on the direct translation of character n-grams, avoiding in this way the need for word normalization during indexing or translation, and also dealing with out-of-vocabulary words. Moreover, since such a solution does not rely on language-speci c processing, it can be used with languages of very di erent natures even when linguistic information and resources are scarce or unavailable. Nevertheless, in contrast with the original approach, our proposal is much faster and transparent, making extensive use of freely available resources. The system has been tested in the robust ad-hoc English-to-French bilingual task, obtaining encouraging results.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing|Indexing methods</kwd>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval| Query formulation</kwd>
        <kwd>I</kwd>
        <kwd>2 [Arti cial Intelligence]</kwd>
        <kwd>I</kwd>
        <kwd>2</kwd>
        <kwd>7 Natural Language Processing|Machine translation</kwd>
        <kwd>Text analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Sunderland { SR6 0DD (UK)
This work is an extension of the proposal originally presented by our group in the previous CLEF
edition, a new knowledge-light approach for query translation in Cross-Language Information
Retrieval (CLIR) systems based on the direct translation of character n-grams. Such a proposal
itself can be considered as an extension of the previous work of the Johns Hopkins University
Applied Physics Lab (JHU/APL) on the employment of overlapping character n-grams for indexing
documents [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
      </p>
      <p>The interest in using overlapping character n-grams comes from the fact that it provides a
surrogate means to normalize word forms and to allow to manage languages of very di erent
natures without further processing. Such a knowledge-light approach does not rely on
languagespeci c processing, and it can be used even when linguistic information and resources are scarce
or unavailable.</p>
      <p>In the case of monolingual retrieval, the employment of n-grams is quite straightforward, since
both queries and documents are just tokenized into overlapping n-grams instead of words: the
word tomato, for example, is split into -tom-, -oma-, -mat- and -ato-. The resulting n-grams
are then processed by the retrieval engine either for indexing or querying. Nevertheless, when
extending its use to the case of CLIR, an extra translation phase is needed during querying.</p>
      <p>
        Aiming to avoid some of the limitations of classic dictionary-based translation methods, such
as the need for word normalization or the inability to handle out-of-vocabulary words, JHU/APL
researchers developed a direct n-gram translation algorithm which allows translation not at the
word level but at the n-gram level [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. This n-gram translation algorithm takes as input a parallel
corpus, aligned at the paragraph (or document) level and extracts candidate translations as follows.
Firstly, for each candidate n-gram term to be translated, paragraphs containing this term in the
source language are identi ed. Next, their corresponding paragraphs in the target language are
also identi ed and, using an ad-hoc statistical measure, a translation score is calculated for each
of the terms occurring in the target language texts. Finally, the target n-gram with the highest
translation score is selected as the potential translation of the source n-gram. Nevertheless, the
whole process was found to be very slow, making the testing of new developments di cult: it
could take several days in the case of working with 5-grams, for example.
      </p>
      <p>This paper describes a new direct n-gram alignment proposal we have developed both to speed
up the process and to make the system more transparent. The article is structured as follows.
Firstly, Sect. 2 describes our approach. Next, in Sect. 3, our proposal is evaluated. Finally, in
Sect. 4, we present our conclusions and future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Description of the system</title>
      <p>
        Taking as our model the system designed by JHU/APL, we developed our own n-gram based
retrieval system, trying to preserve the advantages of the original proposal but avoiding its main
drawbacks. Moreover, instead of the ad-hoc resources developed for the original system [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ],
our system has been built using freely available resources when possible in order to make it
more transparent and to minimize e ort. This way, we use the open-source retrieval platform
Terrier [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] instead of the ad-hoc retrieval system employed by the original design, and the
well-known Europarl parallel corpus1 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is used as training data instead of the ad-hoc corpus
employed by JHU/APL.
      </p>
      <p>
        Nevertheless, the main di erence is the n-gram alignment algorithm itself, which now consists
of two phases. In the rst phase, the slowest one, the input parallel corpus is aligned at the
word-level using the well-known statistical tool GIZA++ [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], obtaining as output the translation
probabilities between the di erent source and target language words. In our case, after some initial
experiments [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], we have opted for a bidirectional alignment [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] which considers a (wEN , wSP )
English-to-Spanish word alignment only if there also exists a corresponding (wSP , wEN )
Spanishto-English alignment. This way the subsequent processing will be focused only on those words
whose translation seems less ambiguous, reducing both the number of input word pairs to be
1This corpus was extracted from the proceedings of the European Parliament, containing up to 28 million words
per language. It includes versions in 11 European languages: Romance (French, Italian, Spanish, Portuguese),
Germanic (English, Dutch, German, Danish, Swedish), Greek and Finnish.
processed and output n-gram pairs to be obtained by more than 60%. This reduction allows us
to reduce greatly both computing and storage resources |including processing time.
      </p>
      <p>
        Next, prior to the second phase, heuristics can be applied |if desired| for re ning or
modifying the word-to-word translation scores calculated by GIZA++. In our case, we have removed
those least-probable word alignments from the input (those with a word translation probability
less than a threshold W , with W =0.15) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Such pruning leads to a considerable extra reduction
of processing time and storage space: a reduction of over 90% in the number of both input word
pairs processed and output n-gram pairs aligned.
      </p>
      <p>
        Finally, in the second phase, n-gram translation scores are computed employing statistical
association measures [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], taking as input the translation probabilities previously calculated by
GIZA++.
      </p>
      <p>As we can see, this rst step acts as an initial lter, since only those n-gram pairs corresponding
to aligned words will be considered, whereas in the original JHU/APL approach all n-gram pairs
corresponding to aligned paragraphs were considered. This approach increases the speed of the
process by concentrating most of the complexity in the word-level alignment phase, allowing
ngram alignment techniques to be easily tested.
2.1</p>
      <p>
        Word-level alignment using association measures
Our n-gram alignment algorithm is an extension of the way association measures can be used for
creating bilingual word dictionaries taking as input parallel collections aligned at the paragraph
level [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. In this context, given a word pair (ws; wt) |ws standing for the source language
word, and wt for its candidate target language translation|, their cooccurrence frequency can be
organized in a contingency table resulting from a cross-classi cation of their cooccurrences in the
input aligned corpus:
      </p>
      <p>S = ws
S 6= ws</p>
      <p>T = wt</p>
      <p>T 6= wt
O11
O21
= C1</p>
      <p>O12
O22
= C2
= R1
= R2
= N
As shown, the rst row accounts for those instances where the source language paragraph contains
ws, while the rst column accounts for those instances where the target language paragraph
contains wt. The cell counts are called the observed frequencies: O11, for example, stands for the
number of aligned paragraphs where the source language paragraph contains ws and the target
language paragraph contains wt; O12 stands for the number of aligned paragraphs where the source
language paragraph contains ws but the target language paragraph does not contain wt; and so
on. The total number of word pairs considered |or sample size N | is the sum of the observed
frequencies. The row totals, R1 and R2, and the column totals, C1 and C2, are also called marginal
frequencies and O11 is called the joint frequency.</p>
      <p>Once the contingency table has been built, di erent association measures can be easily
calculated for each word pair. The most promising pairs, those with the highest association measures,
are stored in the bilingual dictionary.
2.2</p>
      <p>Adaptations for n-gram-level alignment
We have described how to compute and use association measures for generating bilingual word
dictionaries from parallel corpora. However, in our case we do not start with aligned paragraphs
composed of words, but aligned words |previously aligned through GIZA++| composed of
character n-grams. A rst choice for adapting the previous word-level alignment algorithm to the
case of n-grams could be just to adapt the contingency table to the new context, by considering
that we are managing n-gram pairs (gs; gt) cooccurring in aligned words instead of word pairs
(ws; wt) cooccurring in aligned paragraphs. So, contingency tables should be adapted accordingly:
O11, for example, should be re-formulated as the number of aligned word pairs where the source
language word contains n-gram gs and the target language word contains n-gram gt.</p>
      <p>
        However, we do not have real instances of n-gram cooccurrences at aligned words, but just
probable ones, since GIZA++ uses a statistical alignment model which computes a translation
probability for each cooccurring word pair [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. So, the same word may be aligned with several
translation candidates, each one with a given probability. Taking as example the case of the
English words milk and milky, and the Spanish words leche (milk), lechoso (milky) and
tomate (tomato), a possible output word-level alignment |with its corresponding probabilities|
would be:
source word
candidate translation
      </p>
      <p>prob.
milk
milky
milk
leche
lechoso
tomate
Our proposal consists of weighting the likelihood of a cooccurrence according to the probability of
its containing word alignments. So, the resulting contingency tables corresponding to the n-gram
pairs (-milk-, -lech-) and (-milk-, -toma-) are as follows:</p>
      <p>T =
-lech</p>
      <p>T 6=
-lechS =
-milk</p>
      <p>O11 = 0:98 + 0:92 =1.90</p>
      <p>O12 = 0:98 + 3 0:92 + 3 0:15 =4.19
S 6=
-milk</p>
      <p>S =
-milkS 6=
-milk</p>
      <p>O11 =0.15</p>
      <p>O12 = 2 0:98 + 4 0:92 + 2 0:15 =5.94
O21 =0.92
Notice that, for example, the O11 frequency corresponding to (-milk-, -lech-) is not 2 as might
be expected, but 1.90. This is because the pair appears in two word alignments |milk{leche
and milky{lechoso|, but each cooccurrence in an alignment has been weighted according to its
translation probability:</p>
      <p>O11 = 0.98 (for milk{leche) + 0.92 (for milky{lechoso) = 1.90 .</p>
      <p>
        Once the contingency tables have been generated, the association measures corresponding to each
n-gram pair can be computed. In contrast with the original JHU/APL approach [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ], which
used an ad-hoc measure, ours uses three of the most extensively used standard measures: the Dice
coe cient (Dice), mutual information (MI ), and log-likelihood (logl ), which are de ned by the
following equations [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]:
      </p>
      <p>Dice(gs; gt) =</p>
      <p>2O11
R1 + C1
: (1)</p>
      <p>M I(gs; gt) = log N O11 : (2)</p>
      <p>R1C1
logl(gs; gt) = 2 X Oij log N Oij : (3)
i;j RiCj
If using the Dice coe cient, for example, we nd that the association measure of the pair (-milk-,
-lech-) |the correct one| is much higher than that of the pair (-milk-, -toma-) |the wrong
one:</p>
      <p>Dice(-milk-, -lech-) = 6:09+2:82 = 0.43 .</p>
      <p>2 1:90
Dice(-milk-, -toma-) = 6:09+0:15 = 0.05 .</p>
      <p>2 0:15
Notice that if we consider that a real existing cooccurrence instance corresponds to a 100%
probability, we can think about the original word-based algorithm described in Sect. 2.1 as a particular
case of the generalized n-gram-based algorithm we have proposed here with n=1.
0.2
0
0.8
0.2</p>
      <p>
        0
Since the lack of time did not allow us to have our n-gram direct translation tool ready on time for
the past CLEF 2006 workshop [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], this year we have taken part again in the robust ad-hoc task,
speci cally in the English-to-French bilingual task, in order to present the current development of
our work.
      </p>
      <p>
        The robust task is essentially an ad-hoc task which re-uses the topics and collections from past
CLEF editions [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In this case, the French document collection is formed by 87,191 news reports
(243 MB) provided by Le Monde and SDA and corresponding to the year 1994. The English topics
set consists of 200 topics divided into two subsets: a training topics subset to be used for tuning
purposes, formed by 100 topics (C041{C140); and a test topics subset for testing purposes, formed
by the remaining 100 topics (C251{C350). Moreover, only title and description topic elds were
used in the submitted queries.
      </p>
      <p>
        With respect to the indexing process, documents were simply split into n-grams and indexed,
as were the queries. We have used 4-grams as a compromise n-gram size after studying the results
previously obtained by the JHU/APL group [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] using di erent lengths. Before that, the text
had been lowercased and punctuation marks were removed [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], but not diacritics. The open-source
Terrier platform [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] was used as retrieval engine with a InL22 ranking model [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. No stopword
removal or query expansion were applied at this point.
      </p>
      <p>
        For querying, the source language topic is rstly split into n-grams. Next, these n-grams are
replaced by their N most probable alignments.3 The resulting translated topics are then submitted
to the retrieval system.4 Because of the lack of time, we could not tune the N value for this new
set of English-to-French experiments, so we decided to take those values used in our previous
English-to-Spanish experiments [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]:
      </p>
      <p>Dice coe cient
Mutual Information
Log-likelihood</p>
      <p>N =1
N =10
N =1
Finally, Fig. 1 and Fig. 2 show the results obtained for each association measure: the Dice
coefcient (EN2FR Dice), Mutual Information (EN2FR MI), and log-likelihood (EN2FR logl). We also
show the results for two baselines: by querying the French index with the initial English topics
split into 4-grams (EN) |allowing us to measure the impact of casual matches|, and by querying
the index using the French topics split into 4-grams (FR) |i.e. a French monolingual run and our
ideal performance goal. Notice that mean average precision (MAP) values are also given.
2Inverse Document Frequency model with Laplace after-e ect and normalization 2.
3With N 2 f1; 2; 3; 5; 10; 20; 30; 40; 50; 75; 100g.</p>
      <p>
        4A second selection algorithm, consisting of taking those alignments with a probability greater or equal than
a threshold T , was also used in previous experiments [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Nevertheless, this threshold-based approach has been
dismissed because of the di culty for xing T and because if performed not as well as this top-rank-based approach.
0.8
10 15 20 30 100 200
      </p>
      <p>Documents retrieved (D)</p>
      <p>These results show that the log-likelihood measure obtains the best results for both topic sets,
although no signi cant di erence is found with respect to Dice.5 On the other hand, both approaches
perform signi cantly better than mutual information.</p>
      <p>Although we still need to improve our results in order to reach our ideal performance goal, our
current results are encouraging, since it must be taken into account that these are still our rst
experiments, so the margin for improvement is still great.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusions and Future Work</title>
      <p>
        This paper extends the proposal originally presented in CLEF 2006 for the development of a CLIR
system which uses character n-grams not only as indexing units, but also as translation units. This
system was inspired by the work of the Johns Hopkins University Applied Physics Lab [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ], but
tries to preserve its advantages while avoiding its main drawbacks. As in their original proposal,
our work is based on the direct translation of character n-grams, avoiding in this way the need for
word normalization during indexing or translation, and also dealing with out-of-vocabulary words.
      </p>
      <p>Moreover, since such a solution does not rely on language-speci c processing, it can be used
with languages of very di erent natures even when linguistic information and resources are scarce
or unavailable. Nevertheless, in contrast with the original approach, our proposal is much faster
and transparent, making extensive use of freely available resources.</p>
      <p>So, the n-gram alignment algorithm described consists of two phases. In the rst phase, the
slowest one, word-level alignment of the text is made through a statistical alignment tool. In the
second phase, n-gram translation scores are computed employing statistical association measures,
taking as input the translation probabilities calculated in the previous phase. This new approach
speeds up the training process, concentrating most of the complexity in the word-level alignment
phase, making the testing of new association measures for n-gram alignment easier.</p>
      <p>
        With respect to our future work, new tests with other languages of di erent characteristics
are being prepared in order to complete the tuning of the system, including the possibility of
removing high or low-frequency n-grams, the employment of relevance feedback, or the use of pre
or post-translation expansion techniques in the case of translingual runs [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This research has been partially funded by the European Union (FP6-045389), Ministerio de
Educacion y Ciencia and FEDER (TIN2004-07246-C03 and HUM2007-66607-C04), and Xunta de Galicia
5Two-tailed T-tests over MAPs with =0.05 have been used along this work.
(PGIDIT05PXIC30501PN, PGIDIT05SIN044E, and Rede Galega de Procesamento da Linguaxe e
Recuperacion de Informacion). The authors would also like to thank John I. Tait for his support.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] http://ir.dcs.gla.ac.uk/terrier/ (visited on
          <year>August 2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] http://www.clef-campaign.
          <source>org (visited on August</source>
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Amati and C. J. van Rijsbergen</surname>
          </string-name>
          .
          <article-title>Probabilistic models of information retrieval based on measuring divergence from randomness</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ):
          <volume>357</volume>
          {
          <fpage>389</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Koehn</surname>
          </string-name>
          .
          <article-title>Europarl: A parallel corpus for statistical machine translation</article-title>
          .
          <source>In Proc. of the 10th Machine Translation Summit (MT Summit X)</source>
          , pp.
          <volume>79</volume>
          {
          <issue>86</issue>
          ,
          <year>2005</year>
          . Corpus available in http://www.iccs.inf.ed.ac.uk/~pkoehn/publications/europarl/ (visited on
          <year>August 2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Koehn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Och</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Marcu</surname>
          </string-name>
          .
          <article-title>Statistical phrase-based translation</article-title>
          .
          <source>In NAACL '03: Proc. of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology</source>
          , pp.
          <volume>48</volume>
          {
          <issue>54</issue>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Schu</surname>
          </string-name>
          <article-title>tze. Foundations of statistical natural language processing</article-title>
          . The MIT Press,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>McNamee</surname>
          </string-name>
          and
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>May eld. Character n-gram tokenization for European language text retrieval</article-title>
          .
          <source>Information Retrieval</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          -2):
          <volume>73</volume>
          {
          <fpage>97</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>McNamee</surname>
          </string-name>
          and
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>May eld. JHU/APL experiments in tokenization and non-word translation</article-title>
          . Vol.
          <volume>3237</volume>
          of Lecture Notes in Computer Science, pp.
          <volume>85</volume>
          {
          <fpage>97</fpage>
          . Springer-Verlag,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Och</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Ney</surname>
          </string-name>
          .
          <article-title>A systematic comparison of various statistical alignment models</article-title>
          .
          <source>Computational Linguistics</source>
          ,
          <volume>29</volume>
          (
          <issue>1</issue>
          ):
          <volume>19</volume>
          {
          <fpage>51</fpage>
          ,
          <year>2003</year>
          . Source code available at http://www.fjoch.com/GIZA++.
          <source>html (visited on August</source>
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Vilares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Oakes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. I.</given-names>
            <surname>Tait</surname>
          </string-name>
          . CoLesIR at CLEF 2006:
          <article-title>rapid prototyping of a n-gram-based CLIR system</article-title>
          .
          <source>In Working Notes of the CLEF 2006 Workshop</source>
          ,
          <year>2006</year>
          . Available at [2].
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Vilares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Oakes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Vilares</surname>
          </string-name>
          .
          <article-title>Character n-grams translation in cross-language information tetrieval</article-title>
          . Vol.
          <volume>4592</volume>
          of Lecture Notes in Computer Science, pp.
          <volume>217</volume>
          {
          <fpage>228</fpage>
          .
          <string-name>
            <surname>SpringerVerlag</surname>
          </string-name>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>