<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>11%) Mid level of English Phrases .</institution>
          <addr-line>46(-25%) .25(-19%) .30(-26%) .24(-22%)</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Main (Low level of English) Phrases .</institution>
          <addr-line>47(-2%) .34(</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2001</year>
      </pub-date>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>is able to process gigabytes of text. Therefore, a fast approximation to tagging
(exible) pattern:
most frequent POS is assigned to all occurrences of a word.
2. Words are tagged for Part-Of-Speech (POS). No POS tagger, to our knowledge,
uses MACO+ [1], and the English processor uses TreeTager [4].
1. Words are lemmatized using morphological analyzers. The Spanish processor
3. A shallow parsing process identies noun phrases that satisfy the following
ensure maximal recall in the phrase detection phase. For other languages, the
is performed: in the case of Spanish, a set of heuristics has been devised to
results are presented and discussed. Finally, in Section 5 we draw some conclusions.
that software.
In Section 2, we describe our phrase-based approach to document translation.
such software with a phrase-aligning algorithm that exploits comparable corpora;
up the phrase extraction software to handle CLEF-size collections; second, enriching
and third, obtaining an indirect measure (via document selection) of the quality of
In Section 3, the experimental setup for the evaluation is explained. In Section 4,
Besides testing our main hypothesis, we had three additional goals: rst, scaling
We have used the phrase extraction software from the UNED WTB Multilingual
search engine [3]. This software performs robust and eÆcient noun phrase extraction
in several languages, and provides two kinds of indexes:
{ maps every noun phrase into documents that contain that phrase.
logical variant of the word.
{ maps every (lemmatized) word into every noun phrase that contains a
morphocorresponding to about 1Gb of text. Before attempting this iCLEF experiment,
and 26,700,000 dieren t candidate phrases are detected. From this set, we have
analyzer, and correspond to proper nouns, typos, foreign words, or words uncovered
process the EFE collection with our (limited) hardware resources, it was necessary
retained the 3,600,000 phrases that appear more than once in the collection.
Overall, 280,000 dieren t lemmas (including unknown words) are considered,
the largest collection processed with our system had 60,000 documents. In order to
These are the approximate gures for the indexing process: 375,000 dieren t
The collection of 200 English documents is very small and poses no problem
words were detected, from which 250,000 were not recognized by the morphological
for indexing. The EFE collection, however, consists of about 250,000 documents
to re-program most of the system.
by the dictionary.
every term of the original phrase. This subset of the Spanish related phrases forms
tain that word. The set of all phrases forms the pool of related Spanish phrases.
the set of candidate translations. In the previous example, the system nds:
For each word in the translations set, we consider all Spanish phrases that
conThen we search all phrases that contain only (and exactly) one translation for</p>
    </sec>
    <sec id="sec-2">
      <title>If the set of candidate translations is empty, two steps are taken:</title>
      <p>abortion issue )
2.2 Phrase alignment
birth control
abortion issue
religious and cultural
English
last year
In the WTB search engine, such indexes are used to provide multilingual
phrasein iCLEF documents.
this data is used as statistical information to provide translations for English phrases
browsing capabilities in an interactive CLIR setting. In the present work, however,
2
12
2
16
2
frequency
5
(correctly) chosen as translation for \abortion issue". Note that all other candidate
phrases also disambiguated \issue" correctly as \tema, asunto’ ’.
Other alignment examples include:
If the subset is non-empty (as in the example above), the system selects the
most frequent phrase as the best phrasal translation. Therefore \tema del aborto" is
# candidates
52
6
3
10
between \birth" and \natalidad". The selected term \control de los nacimientos",
however, is unusual but understandable (in context) for a Spanish speaker.
la natalidad" (with a frequency of 107), but the dictionary does not provide a link
The most appropriate translation for \birth control" would rather be \control de
asunto aborto
asuntos como el aborto
temas como el aborto
tema del aborto
asuntos del aborto
asunto del aborto
phrase
emitir, expedir, dar, promulgar
phrase: "abortion issue"
translations: abortion -&gt; aborto
expedicion, descendencia, publicar,
lemmas: abortion, issue
issue -&gt; asunto, tema, edicion, numero, emision,
a bilingual dictionary. For instance:
For each English phrase, we start translating all content words in the phrase using
selected
control de los nacimientos
tema del aborto
culturales y religiosos
an~o pasado
el tema del aborto domino las nueve jornadas del Congreso Internacional sobre
Poblacion y Desarrollo.</p>
    </sec>
    <sec id="sec-3">
      <title>English sentence</title>
      <p>2.3 Phrase-based document translation
For instance:
As an example, let us consider this sentence from one of iCLEF documents:
Manual translation
while Systran produces:
A valid manual translation of the above sentence would be:
subphrase alignments:
development -&gt; desarrollo
population -&gt; poblacion
day international -&gt; jornadas internacionales
final translation:
word by word translations:
day international conference -&gt; jornada del congreso internacional
"jornada del congreso internacional poblacion desarrollo"
Note that, while the indexed phrase is not an optimal noun phrase (\day" should
ument selection).
be removed) and the translation is not fully grammatical, the lexical selection is
accurate, and the result is easily understandable for most purposes (including
docday -&gt; da, jornada, epoca, tiempo
lemmas: day, international, conference, population, development
development -&gt; desarrollo, avance, cambio, novedad, explotacion,
phrase: "day international conference on population and development"
international -&gt; internacional
conference -&gt; congreso, reunion
urbanizacion, revelado
possible translations:
population -&gt; poblacion, habitantes
lation and Development.
the abortion issue dominated the nine-day International Conference on
Popu{ Phrases that have an optimal alignment (boldface).
2. List the translations obtained for each original phrase according to the alignment
{ Phrases containing query terms (bright colour).
phase, highligting:
1. Find all maximal (i.e., not included in bigger units) phrases in the document,
and sort them by order of appearance in the document.
the alignment process. The basic process is:
The pseudo-translation of the document is made using the information obtained in
la edicion del aborto domino el de nueve das Conferencia in ternacional sobre
la poblacion y el desarrollo.
less noisy translations. If any of the phrases contain a (morphological variant of) a
query term for a particular search, the phrase is further highlighted.
where boldface is used for optimal phrase alignments, which are supposed to be</p>
    </sec>
    <sec id="sec-4">
      <title>Aside from grammatical correctness, Systran translation only makes one rele</title>
      <p>Our phrase indexing process, on the other hand, identies t wo maximal phrases:
aborto"(meaningless) instead of \tema del aborto".
vant mistake, interpreting \issue" as in \journal issue" and producing \edicion del</p>
    </sec>
    <sec id="sec-5">
      <title>Phrasal pseudo-translation</title>
      <p>Fig. 1. Search interface: MT system
Systran MT translation
tema del aborto
jornada del congreso internacional poblacion desarrollo
day International Conference on Population and Development
abortion issue
of our system is:
which receive the translations showed in the previous section. The nal displa y
level and high-level English skills.
we recruited 8 volunteers with low or no prociency at all in the English language.
We made three experiments with dieren t searcher proles: for the main experimen t,
For purposes of comparison, we formed two additional 8-people groups with
midtheir study center (with the presence of the same monitor).
but v e of them (UNED students) carried on the experiments via Internet from
Figure 1 shows an example of document displayed in the Systran MT system.
controlled by the system interface. Most of the searchers used the system locally,
The time for each search, and the combination of topics and systems, were fully
We followed closely the search protocol established in the iCLEF guidelines [2].
reliable phrasal translations (boldface).
Figure 2 shows the same document paragraph in our phrase-based system. The
latter shows less information (only noun phrases extracted and translated by the
system), highlights phrases containing query terms (bright green) and emphasizes
recall is lower, and the absolute gures are higher both for MT and phrasal
transjudge documents faster without loss of accuracy.
{ Mid-level English speakers have lower precision and recall for the phrasal
transure 3 for a comparison between low and high English skills.
that made the experiment remotely (see discussion below).
recognize them, these results are coherent with the main experiment. See
Fig{ In the main experiment with monolingual searchers (\Low level of English"),
precision is very similar, but phrasal translations get 52% more recall. Users
{ Users with good knowledge of English show a similar pattern, but the gain in
ysis of the data revealed that this experiment was spoiled by the three searchers
lations. As unknown words remain untranslated and English-speaking users may
lation system, contradicting the results for the other two groups. A careful
anal</p>
    </sec>
    <sec id="sec-6">
      <title>1. J. Carmona, S. Cervell, L. Marquez, M. A. Mart, L. P adro, R. Placer, H. Rodrguez,</title>
      <p>M. Taule, and J. Turmo. An environment for morphosyntactic processing of
unrelar phrasal equivalents in the searcher’s language, might be more appropriate for
Although the number of searchers does not allow for clear-cut conclusions, the
document selection than full-edged MT. Our purpose is to reproduce a similar
shallow NLP techniques is feasible for large-scale IR collections. The major
bottleresults of the evaluation indicate that summarized translations, and in
particuAs a side conclusion, we have proved that phrase detection and handling with
experiment with more users, and better-controlled experimental conditions, to have
neck, Part-Of-Speech tagging, can be overcome with heuristic simplications that
a better testing of our hypothesis in a near future.
do not compromise the usability of the results, at least in the present application.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>