<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Benefits of deep NLP-based lemmatization for information retrieval</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>P ́eter Hal ́acsy Budapest University of Technology and Economics Centre for Media Research</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper reports on our system used in the CLEF 2006 ad hoc mono-lingual Hungarian retrieval task. Our experiments focus on the benefits that deeper NLP-based lemmatization (as opposed to simpler stemmers) can contribute to mean average precision. Our results show that these benefits counterweight the disadvantage of using an off-the-shelf retrieval toolkit.</p>
      </abstract>
      <kwd-group>
        <kwd>Information retrieval</kwd>
        <kwd>Stemming</kwd>
        <kwd>Morphological Analysis</kwd>
        <kwd>Hungarian Language</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Background for the Hungarian experiments</title>
      <p>The first Hungarian test collection for information retrieval appeared last year at the CLEF 2005
conference. Previously there have been no empirical experiments measuring the effect of Hungarian
language features on retrieval performance.</p>
      <p>The Hungarian language1 is highly inflectional, rich in compound words, and has an extensive
inflectional and derivational morphology. Only nominals and verbs can be suffixed with inflectional
suffixes. Nominals and verbs can also be suffixed with productive derivational suffixes but there
are few derivational suffixes which can attach to adverbs, preverbs and postpositions.</p>
      <p>The set of inflectional suffixes after nominals are more or less the same, therefore nouns,
adjectives and numerals cannot be defined solely on the base of inflectional morphology. The stem
can be followed by 7 types of Possessive (POSS), 3 Plural, 3 Anaphoric Possessive (ANP) and 17
Case suffixes that results in as many as 1134 possible forms. In grammar textbooks we often find
examples such as: botjaitok´einak = ’for the sg. of your (pl) sticks’, which is analyzed as:
bot/NOUN&lt;PLUR&gt;&lt;POSS&lt;2&gt;&lt;PLUR&gt;&gt;&lt;ANP&lt;PLUR&gt;&gt;&lt;CAS&lt;DAT&gt;&gt;
1A more detailed descriptive grammar of Hungarian is available at http://mokk.bme.hu/resources/ir
The plural and case suffixes have several alternants. The alternation is governed by several
morpho-phonological processes such as vowel harmony and depends on the stem and on the types
of the case suffixes. For example the nominal plural is expressed with the suffix -k, but after
consonant-final stems it is preceded by a non-high vowel (a e o ¨o), the so-called linking vowel.
The quality of the linking vowel depends on the stem-vowel(s) and other lexical properties of the
stem. Table 1 shows the variants of the plural suffix for different stems.</p>
      <p>Singular
kar
v´ar
b´er
b¨or
kapu
lufi
haj´o
fa
kefe</p>
      <p>Plural
karok
v´arak
b´erek
b¨or¨ok
kapuk
lufik
haj´ok
f´ak
kef´ek</p>
      <p>GLOSS
’arm(s)’
’castle(s)’
’wage(s)’
’skin(s)’
’gate(s)’
’balloon(s)’
’boat(s)’
’tree(s)’
’brush(es)’</p>
      <p>Suffixing of foreign names presents a specific challenge for information retrieval because of the
effort needed to maintain proper name lexicons. If a foreign name is suffixed, the selection of
suffix allomorphs can be sensitive to the normal Hungarian pronunciation of the stem as Table
1 shows some examples for vowel harmony (a ∼ e), consonant assimilation, and for final vowel
lengthening (a ∼ ´a, e ∼ ´e, o ∼ ´o). Note that in the last type the stem-final a e o changes to the
long (accented) versions.</p>
      <p>suffixed form
Clintonnal
Reagannel
Austerrel
O’Connorral
Balzackal
Lucasszal
Bachhal
Hess´evel
Tzar´aval
Hug´oval</p>
      <p>GLOSS
’with Clinton’
’with Reagan’
’with Auster’
’with O’Connor’
’with Balzac’
’with Lucas’
’with Bach’
’with Hesse’
’with Tzara’
’with Hugo’</p>
      <p>Verbs have fewer forms than nouns: 104 forms (in our classification) that express various
person, number, tense, and transitivity distinctions. The phonological alternations are very similar
to those in the nominal system. For example: v´ar+ok, k´er+ek, u¨t+¨ok = ’I wait, I request, I hit’.</p>
      <p>Similarly to German and Finnish, compounding is very productive in Hungarian. Almost any
two (or three) nominals next to each other can form a single compound written without an
intervening whitespace. Examples are vizumk¨otelezetts´eg = ’obligation (to carry) visa’, ´opiumel˝o´all´ıt´as
= ’opium manufacture’, u¨vegh´azhat´as = ’glass house effect (greenhouse effect)’.</p>
      <p>The use of hyphens can also cause problems, as it is governed by complex orthographic rules.
Some examples are given in Table 1.</p>
    </sec>
    <sec id="sec-2">
      <title>Stemming algorithms</title>
      <p>It should be evident from the foregoing that extensive stemming is especially beneficial for
Hungarian information retrieval. All of last year’s top five systems (Table 2) had some method for
handling the rich morphology of Hungarian: either words were tokenized to n-grams or an
algorithmic stemmer was used.</p>
      <p>part
jhu/apl
unine
miracle
humminngbird
hildesheim
run
aplmohud
UniNEhu3
xNP01ST1
humHU05tde
UHIHU2
map
41.12%
38.89%
35.20%
33.09%
32.64%
stemming method
4gram
Savoy’s stemmer + decompounding
Savoy’s stemmer
Savoy’s stemmer + 4gram
5gram</p>
      <p>
        The best result was achieved by JHU/APL in the run called aplmohud [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. They used a
character 4-gram based tokenization in a language modeling information retrieval system. The
n-gram technique solves the problem of rich agglutinative morphology and compounding. For
example the word atomenergia = ’atomic energy’ in the query is tokenized to atom, tome, omen,
mene, ener, nerg, ergi, rgia strings. When the text only contains the form atomenergi´aval = ’with
atomic energy’, the system still finds the relevant document.
      </p>
      <p>Although the Snowball stemmer was also used together with the n-gram tokenization for the
English and French tasks, the Hungarian results were nearly as good: English 43.46%, French
41.22% and Hungarian 41.12%. From these results it seems that the difference between the
isolating and agglutinating languages can be eliminated by character n-gram methods. But note
that JHU/APL used a state-of-the-art retrieval model, achieving much better results than other
contestants using n-gram methods.</p>
      <p>
        Unine [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Miracle [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and Hummingbird [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] employ the same algorithmic Hungarian stemmer
that removes the nominal suffixes corresponding to the different cases, the possessive and the
number (plural). Above this, UniNEhu3 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] also utilizes a language independent decompounding
algorithm that tries to segment compounds according to corpus statistics calculated from the
document collection.[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] This is only triggered by long words, composed by more than 8 characters.
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] isolates the effect of the stemmer which increases the mean average precision from 18.24%
to 27.4%. However, they also mention that in some cases the aggressive overstemming causes a
drop in performance. The bank =’bank’ word is stemmed to ban=’in’ because the stemmer assumes
this this k is the plural suffix. The other example is n´emet =’German’ which is stemmed after
accent removal to nem which is a homonym word meaning ’no’ or ’gender’. In their best runs
they also use 4-grams instead of stemming and decompounding.
      </p>
      <p>
        This suggests that a lexical stemmer – that incorporates a lexicon and the morphological rules
– can be more beneficial for retrieval. The n-gram experiments show that the decompounding is
more than salutary. This hypothesis is confirmed by [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. They measured that for Finnish, which
is very similar to Hungarian, lexicon based lemmatization increase mean average precision from
32.8% to 56.1%. A pure algorithmic stemmer can only achieve as high as 42.4% score. They also
emphasize the positive effect of decompounding.
      </p>
      <p>
        [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] got a poor performance at the Multilingual Web Track when using a stemmer somewhat
similar to ours. This stemmer was based on the Hungarian ispell wordlist, created for the Hunspell
spellchecker. It reduced the accuracy of the retrieval. They do not mention specific faults in their
report, therefore we cannot reconstruct what might have caused the decline. However we suggest
that the ’take-the-first-stem’ heuristic did not work well with the spellchecker’s default settings.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Our system</title>
      <p>
        Our research group has been working on a Hungarian morphological analyzer for three years. First
we extended the codebase of MySpell, a reimplementation of the well-known Ispell spellchecker,
yielding a generic word analysis library[
        <xref ref-type="bibr" rid="ref15 ref9">9, 15</xref>
        ]. At this point the development of the library has
forked. Now the extended MySpell, called HunSpell, is part of the OpenOffice.org multilingual
office suite. Hunmorph is the program tuned to morphological analysis.
      </p>
      <p>Our technology, like the Ispell family of spellcheckers it descends from, enforces a strict
separation between the language-specific resources (known as dictionary and affix files), and the
runtime environment, which is independent of the target natural language.</p>
      <p>Compiling accurate wide coverage machine-readable dictionaries and coding the morphology
of a language can be an extremely labor-intensive task, so the benefit expected from reusing
the language-specific input database across tasks can hardly be overestimated. To facilitate this
resource sharing and to enable systematic task-dependent optimizations from a central lexical
knowledge base, we designed and implemented a powerful offline layer we call hunlex. Hunlex
offers an easy to use general framework for describing the lexicon and morphology of any language.
Using this description it can generate the language-specific aff/dic resources, optimized for the
task at hand.</p>
      <p>
        morphdb.hu[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] is the Hungarian lexical database and morphological grammar for the hunmorph
framework. It is the outcome of a several-year collaborative effort and represents a resource with
the widest coverage and broadest range of applicability presently available for Hungarian. The
grammar resource is the formalization of well-founded theoretical decisions handling inflection and
productive derivation.
      </p>
      <p>
        The coverage of morphdb.hu was measured running the analyzer on two Hungarian Corpora.
One of them is the Szeged Corpus [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] which contains 1 million words, and on which the recall of
our analyzer is 90%. The missing 10% are mostly proper names and acronyms not analyzed partly
due to the difficulty of multi-word named entity tokenization. The other corpus is the 700 million
word Hungarian Webcorpus [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ] on which the proportion of out-of-vocabulary items is 7%.
      </p>
      <p>
        For the CLEF 2006 experiments we have developed two types of lemmatizers. The first is
context independent: we choose the analysis with the shortest lemma, ie. we try to strip the
longest possible suffix. The second algorithm is more sophisticated: the choice between alternative
morphological analyses is resolved using the output of a POS tagger[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. When there are several
analyses that match the output of the tagger, we choose one with the biggest number of identified
morphemes.
      </p>
      <p>The morphological analyzer is able to guess possible analyses even if the input word is out
of vocabulary. This feature works similarly to resourceless algorithmic stemmers, but hunmorph
knows much more about suffix rules (also including derivations and compounding), and it only
gives results in accordance with Hungarian morpho(-phono)logical rules. For example it is known
that the word bank cannot be plural as suggested by Savoy’s stemmer (see Section 2), because
following final-consonant a linking-vowel is needed before the ’k’.</p>
      <p>Decompounding is similar to guessing, as hunmorph has to analyze words that are not
included in their lexicon. An unknown word is split if its components are known nominals. It is a
grammatical rule that non-nominals can’t be combined.</p>
      <p>Another feature of hunmorph is blocking. If this option is set, the algorithm gives back less
analyses. A lexical (non-affixed) partial analysis always blocks one that involves affixation. Out of
two partial analyses, only the ones that are not equivalent are kept. Blocking effectively implements
the idea that productive generation of an item by affixation or compounding is a fallback option
in case the item is not found lexicalized. The ‘blocking’ and ‘compounds’ options can be used
alongside in which case blocking also suppresses a compound analysis if the compound is entered
as a lexical item.</p>
      <p>morphdb.hu contains lexicalized words even with derivational suffixes and compounds. For
example the word u¨vegh´az =‘glass house (greenhouse)’ is represented as one entry in the lexicon,
therefore the blocking option suppresses the decompounded analysis. Although blocking is only
an option of hunmorph, during our experiments we always used it, as it can prevent the stemmer
from overstemming.</p>
      <p>For lemmatization we implemented a post-hoc filter that first chooses only one analysis which
the lemma is extracted from in the second step. If there is only one possible analysis (as in 50%
of the cases), use this analysis. Otherwise the ambiguity can be reduced by the following rules:
• If there is at least one analysis that is neither compound nor guessing then only this result
is considered.
• If there is no such ‘simple’ output, the compounds of known words dominate over the guessed
lemmas (analyses of unknown words).</p>
      <p>The used analysis is the one with the shortest lemma. The next option is to decide whether to
split the compound words or not. In decompounding mode, the components are added to the
index next to each other as different tokens. Finally, and independently from decompounding,
there is an option whether to strip derivations or not. Because of the blocking option mentioned
above, derivation stripping has no effect on lexicalized words.</p>
      <p>
        The more sophisticated lemmatization procedure involves a POS tagger[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] in the first step.
This is a time consuming task that attaches an inflectional tag to all tokens in the corpus. To
decrease ambiguity, only the matching analyses (with the same POS tags) are chosen from the
output of the morphological analyzer. After this, the same post-hoc filter is applied.
      </p>
      <p>
        We used Lucene (off-the-shelf) for indexing and retrieval with its standard vector space model.
Only the stemmed tokens were added to the index. All fields of documents were concatenated.
We used our own rule based tokenizer based on Lucene’s StandardTokenizer. We prohibited
the recognition of web hosts and acronyms, as these can be confused with periods between two
sentences (wrongly) written without an intervening whitespace. And we allowed words to contain
hyphens, since foreign names often are suffixed with linking hyphens. But after lemmatization the
words were split at remaining hyphens. Before indexing the text was converted to lowercase (the
morphological analyzer can handle uppercase words) but the accented characters were unmodified.
In addition, we used the same stopword list as [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]4. Both the index the queries were stopped.
      </p>
      <p>
        The disambiguation needs sentence boundary detection which is unusual in information
retrieval. For this we trained a maximum entropy model on the Hungarian Webcorpus [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Evalutation</title>
      <p>All measurements were performed with the topics of CLEF 2005 and 2006 to be able to compare
our results to others’. Table 5 shows the performance of our baseline system which doesn’t use
any stemming.</p>
      <p>stem
no
deriv</p>
      <p>comp
no
no
year
2005
2006</p>
      <p>MAP
21.31
18.30</p>
      <p>MRR
48.17
44.95</p>
      <p>In ad hoc retrieval experiments, the most used evaluation measure is ”average precision”.
For a topic, it is the average of the precision figures obtained after each new relevant document
is observed (using zeros as the precision for relevant documents which are not retrieved). By
convention, it is based on the first 1000 retrieved documents for the topic. The score ranges from
0.0 (no relevant found) to 1.0 (all relevant ones were at the top of the list). The Mean average
precision (MAP) is the mean of average precision scores over all of the test topics. This measure
is higher for systems which retrieve relevant documents early in the ranking list.</p>
      <p>Reciprocal rank (RR) and Precision at 10 documents (P10) focus only on the head of the ranked
list. For a topic, RR equals to 1r , where r is the rank of the first retrieved relevant documents.
If RR = 1 the first retrived document is relevant. If RR = 0.5 then the second. If no relevant
document was found, RR = 0. Mean reciprocal rank is the mean of the reciprocal ranks over
all topics. P10 is the precision after the first 10 document retrieved. This is important for web
applications, as it’s well known that the average user doesn’t typically look at the second page
of search results. The last measure is overall recall, ret/rel, i.e. the number of retrieved relevant
documents divided by the number of all relevant documents.</p>
      <p>Table 6 shows the performance gain caused by stemming. The second and third rows isolate the
impact of stripping derivations (deriv) and splitting compounds (compound). As the results show
decompounding has much more (positive) effect on accuracy. But an unexpected result was that
stripping derivations also increase precision. This might be caused by the fact that hunmorph
blocks stripping of non-productive derivations, as lexicalized words are contained in the high
coverage lexical database. So during stemming only productive and transparent derivations are
stripped, like the frequent verbal derivations, e.g. kl´on+oz = ‘to clone’.</p>
      <p>It is encouraging that in the 2005 (Hungarian monolingual ad hoc) task we achieved better
MAP than all other CLEF participants except JHU/APL. Since we used the off-the-shelf Lucene
toolkit for retrieval, we can deduce that the good results are due to lemmatization.</p>
      <p>Table 7 shows the retrieval performance when the POS tagger is utilized. We can see the
unexpected result that morphological disambiguation has no significant positive effect on performance.
What’s more, in some cases a decline is experienced.</p>
      <p>Figure 2 shows the overall recall vs. precision graph. At some recall levels the simple rule-based
lemmatizer performs better, at others the disambiguator-based does.</p>
      <p>Retrospectively, it is not too surprising that the morphological disambiguator did not bring
improvements. First, for most of the tokens, there is nothing to disambiguate, as the lemma is
unique. Second, when the disambiguator solves some nontrivial problem, it is often irrelevant from
4The stopword list is downloadable at http://ilps.science.uva.nl/Resources/HungarianStemmer/
deriv</p>
      <p>comp
deriv</p>
      <p>comp
ion 60%
s
i
c
e
rP 50%
e
g
a
r
va 40%
A</p>
      <p>MAP
0.3227
0.2797
0.3361
0.2933
0.3746
0.3317
0.3926
0.3482
year
2005
2006
2005
2006
2005
2006
2005
2006
a retrieval point of view. E.g. disambiguation of the frequent egy homonym (‘one/NUM’ or ‘a/DET’)
is a hard task, but either of the lemmata is discarded by stopword filtering.</p>
      <p>Another example is the classification of definite or indefinite verbal conjugation.5 The tagger
has to decide on the definiteness of csin´altam=‘I did’ based on the context, although the same
lemma will be used in the two cases. So the effort of the tagger is irrelevant in retrieval.</p>
      <p>Actual morphological homonym resolution also does not seem to be very useful for Hungarian
IR. For example, the word falunk can be segmented as falu+unk = ‘our village’ or as fal+unk
=‘our wall’. But such ambiguities are more likely to appear in theoretical textbooks than in
real-world language usage, very seldom affecting search results.</p>
      <p>Using the POS tagger can even lead to new kinds of errors. For example, in Topic 367, an
error of the morphological analyzer amplified an error of the POS tagger: The analyzer gave two
analyses for the word drogok. The drog/PLUR inflected and the (arguably overgenerated) drog+ok
(drug-cause) compounded one. The post-hoc filter presented in Section 3 would throw away the
compounded version, but the POS tagger chooses it, preferring the singular to the plural form.
This problem could be resolved by integrating the post-hoc filter into the POS tagger.</p>
      <p>Table 8 shows our official results submitted to CLEF 2006. Only after the CLEF
submission deadline did we realize that the simple rule-based lemmatizer can achieve the precision
of the disambiguator-based, more advanced lemmatizer. So all submitted results are for the
disambiguator-based lemmatizer. The values reported here are slightly different from the fourth
row of Table 7, since we’ve reimplemented the lemmatizer meanwhile, and a bug has been
eliminated.</p>
      <p>map
MRR
P10
ret/rel
The experiments on which we report in this paper confirm that a deep NLP-based
lemmatization in Hungarian greatly improves retrieval accuracy. Our system outperformed all CLEF 2005
systems that use algorithmic stemmers even though our retrieval toolkit is using a
not-state-ofthe-art off-the-shelf basic vector space model. The good results are due to the high coverage
lexical resources and the morphological analyzer which allow us the aggressive stemming and
decompounding without overstemming.</p>
      <p>We compared two different morphological analyzer-based lemmatization methods. We have
found that a more advanced method based on high-precision context-sensitive morphological
disambiguation does not bring improvements compared to a simpler context-insensitive greedy
lemmatization algorithm.</p>
      <p>Our Hungarian lemmatizer (together with its morphological analyzer and a Hungarian
descriptive grammar) is released under a permissive LGPL-style license, and can be freely downloaded
from http://mokk.bme.hu/resources/ir. We hope that members of the CLEF community will
incorporate these into their IR systems, closing the gap in effectivity between IR systems for
Hungarian and for major European languages.</p>
      <p>5The verb forms have to be agreed with the definiteness of the object in the sentence. If the verb is intransitive
or the object is an indefinite noun phrase, the indefinite value has to be used. The definite value on a verb refers a
definite object.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D</given-names>
            <surname>´ora Csendes</surname>
          </string-name>
          , Csaba Hatvani,
          <source>Zolt´an Alexin, J´anos Csirik</source>
          , Tibor Gyim´othy, G´abor
          <article-title>Pr´osz´eky, and Tam´as V´aradi. K´ezzel annot´alt magyar nyelvi korpusz: a Szeged Korpusz</article-title>
          .
          <source>In II. Magyar Sz´am´ıt´og´epes Nyelv´eszeti Konferencia</source>
          , pages
          <fpage>238</fpage>
          -
          <lpage>245</lpage>
          . Szegedi Tudom´anyegyetem,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] Jos´e Miguel Gon˜i Menoyo, Jos´e C.
          <article-title>Gonz´alez, and Julio Vilena-Rom´an. Miracle's 2005 approach to monolingual information retrieval'</article-title>
          .
          <source>Working Notes for the CLEF 2005 Workshop</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          September, Vienna, Austria,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>P</given-names>
            <surname>´eter Hal</surname>
          </string-name>
          <article-title>´acsy, Andr´as Kornai, L´aszl´o N´emeth, Andr´as Rung, Istv´an Szakad´at, and Viktor Tro´n. Creating open language resources for Hungarian</article-title>
          .
          <source>In Proceedings of Language Resources and Evaluation Conference (LREC04)</source>
          .
          <source>European Language Resources Association</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P</given-names>
            <surname>´eter Hal</surname>
          </string-name>
          <article-title>´acsy, Andr´as Kornai, Csaba Oravecz</article-title>
          , Viktor Tr´on, and D´aniel Varga.
          <article-title>Using a morphological analyzer in high precision POS tagging of Hungarian</article-title>
          .
          <source>In Proceedings of LREC 2006</source>
          , pages
          <fpage>2245</fpage>
          -
          <lpage>2248</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] Andr´as Kornai, P´eter Hal´acsy, Viktor Nagy, Csaba Oravecz, Viktor Tr´on, and D´aniel Varga.
          <article-title>Web-based frequency dictionaries for medium density languages</article-title>
          .
          <source>In Proceedings of the EACL 2006 Workshop on Web as a Corpus</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Craig</given-names>
            <surname>Macdonald</surname>
          </string-name>
          , Vassilis Plachouras, He Ben, Lioma Christina, and Ounis Iadh. University of Glasgow at WebCLEF 2005:
          <article-title>Experiments in per-field normalisation and language specific stemming</article-title>
          .
          <source>Working Notes for the CLEF 2005 Workshop</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          September, Vienna, Austria,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Craig</given-names>
            <surname>Macdonald</surname>
          </string-name>
          , Vassilis Plachouras, Ben He,
          <string-name>
            <surname>Christina Lioma</surname>
            , and
            <given-names>Iadh</given-names>
          </string-name>
          <string-name>
            <surname>Ounis</surname>
          </string-name>
          . Finnish,
          <article-title>Portuguese and Russian Retrieval with Hummingbird SearchServerTM at CLEF</article-title>
          <year>2004</year>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Paul</given-names>
            <surname>McNamee</surname>
          </string-name>
          .
          <source>Exploring New Languages with HAIRCUT at CLEF 2005. Working Notes for the CLEF 2005 Workshop</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          September, Vienna, Austria,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>L´</surname>
          </string-name>
          <article-title>aszl´o N´emeth, Viktor Tr´on, P´eter Hal´acsy, Andr´as Kornai, Andr´as Rung, and Istv´an Szakada´t. Leveraging the open-source ispell codebase for minority language analysis</article-title>
          .
          <source>In Proceedings of SALTMIL 2004. European Language Resources Association</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Jacques</given-names>
            <surname>Savoy</surname>
          </string-name>
          .
          <article-title>Report on CLEF-2003 monolingual tracks: Fusion of probabilistic models for effective monolingual retrieval. Results of the CLEF-2003, cross-language evaluation forum</article-title>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Jacques</given-names>
            <surname>Savoy</surname>
          </string-name>
          and
          <string-name>
            <surname>Pierre-Yves Berger</surname>
          </string-name>
          .
          <source>Report on CLEF-2005 Evaluation Campaign: Monolingual, Bilingual, and GIRT Information Retrieval. Working Notes for the CLEF 2005 Workshop</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          September, Vienna, Austria,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Tomlinson</surname>
          </string-name>
          .
          <article-title>Europan Ad hoc retireval experiments with HummingbirdTM at CLEF 2005'</article-title>
          .
          <source>Working Notes for the CLEF 2005 Workshop</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          September, Vienna, Austria,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Anna</given-names>
            <surname>Tordai</surname>
          </string-name>
          and Maarten de Rijke.
          <source>Hungarian Monolingual Retrieval at CLEF 2005. Working Notes for the CLEF 2005 Workshop</source>
          ,
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          September, Vienna, Austria,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14] Viktor Tr´on, P´eter Hal´acsy, P´eter Rebrus,
          <article-title>Andr´as Rung, P´eter Vajda, and Eszter Simon. Morphdb.hu: Hungarian lexical database and morphological grammar</article-title>
          .
          <source>In Proceedings of LREC 2006</source>
          , pages
          <fpage>1670</fpage>
          -
          <lpage>1673</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Viktor</given-names>
            <surname>Tr</surname>
          </string-name>
          <article-title>´on, Gy¨orgy Gyepesi, P´eter Hal´acsy, Andr´as Kornai, L´aszl´o N´emeth, and D´aniel Varga. Hunmorph: open source word analysis</article-title>
          .
          <source>In Proceeding of the ACL 2005 Workshop on Software</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>