<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised Morpheme Analysis Evaluation by a Comparison to a Linguistic Gold Standard - Morpho Challenge 2008</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mikko Kurimo</string-name>
          <email>Mikko.Kurimo@tkk.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matti Varjokallio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Adaptive Informatics Research Centre, Helsinki University of Technology</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Algorithms</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The goal of Morpho Challenge 2008 was to find and evaluate unsupervised algorithms that provide morpheme analyses for words in different languages. Especially in morphologically complex languages, such as Finnish, Turkish and Arabic, morpheme analysis is important for lexical modeling of words in speech recognition, information retrieval and machine translation. The evaluation in Morpho Challenge competitions consisted of both a linguistic and an application oriented performance analysis. This paper describes an evaluation where the competition entries were compared to a linguistic morpheme analysis gold standard. Because the morpheme labels in an unsupervised analysis can be arbitrary, the evaluation is based on matching the morpheme-sharing words between the proposed and the gold standard analyses. In addition to Finnish, Turkish, German and English evaluations performed in Morpho Challenge 2007, the competition this year had an additional evaluation in Arabic. The results in 2008 show that although the level of precision and recall varies substantially between the tasks in different languages, the best methods seem to manage all the tested languages quite well. The Morpho Challenge was part of the EU Network of Excellence PASCAL Challenge Program and organized in collaboration with CLEF.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The topic of the Morpho Challenge 2008 competition is to evaluate proposed unsupervised
machine learning algorithms in the task of morpheme analysis for words in different languages. The
Morpho Challenge evaluation consisted of both a linguistic and an application oriented
performance analysis. The linguistic evaluation described in this paper, Competition 1, is based on a
comparison of the suggested morpheme analysis to a linguistic morpheme analysis gold standard.
The practical application oriented evaluation described in the companion paper [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
Competition 2, contained information retrieval (IR) experiments from CLEF, where the all the words in
the queries and text corpus were replaced by their morpheme analyses.
      </p>
      <p>
        The Morpho Challenge 2008 tasks and training corpora were the same as in our previous
Morpho Challenge 2007 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], except that it involved one additional morphologically complex language,
Arabic. There was also an optional evaluation of the IR performance using the morpheme analysis
of word forms in their full text context. The difference to our first Morpho Challenge 2005 [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
which focused on just the segmentation of words into morphologically meaningful units, was that
the units should further be clustered into the abstract classes of morphemes. For example, this
analysis should find the link between the word forms “foot” and “feet”.
      </p>
      <p>
        Especially in morphologically complex languages, such as Finnish, Turkish and Arabic, the
morpheme analysis is important for lexical modeling of words in speech recognition [
        <xref ref-type="bibr" rid="ref1 ref9">1, 9</xref>
        ],
information retrieval [
        <xref ref-type="bibr" rid="ref13 ref7">13, 7</xref>
        ] and machine translation [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]. Due to the high level of agglutination,
inflection, and compounding, there are millions of different word forms, which is clearly too much
for building an effective vocabulary and training probabilistic models for the relations between
words. There also exist carefully constructed linguistic tools for morphological analysis, but only
for few languages. Even in these cases using statistical machine learning methods we may still
discover interesting alternatives that may rival even the most sophisticated linguistically designed
morphologies.
      </p>
      <p>The scientific objectives of the Morpho Challenge competitions are: to learn about the word
construction in natural languages, to advance machine learning methodology, and to discover
approaches that are suitable for many languages. The portability to different languages is very
important, because the language technology often needs to be quickly extended to various new
languages for which there are limited amount of resources available. Unsupervised learning is
then the most attractive approach for data analysis, because the majority of the available data is
unannotated and human annotation work is expensive.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Task and Data in Competition 1</title>
      <p>The task in the Morpho Challenge 2008 was to return the given list of words in each language
extended by the morpheme analysis of each word form. The morpheme analyses should be
obtained by an unsupervised learning algorithm that would preferably be as language independent
as possible. In each language, the participants were pointed to a training corpus in which all the
words occur (in a sentence), so that the algorithms may also utilize information about the word
context. The tasks were the same as in the Morpho Challenge 2007 last year with the addition of
one new language, Arabic.</p>
      <p>The training corpora were the same as in the Morpho Challenge 2007, except for Arabic: 3
million sentences for English, Finnish and German, and 1 million sentences for Turkish in plain
unannotated text files that were all downloadable from the Wortschatz collection1 at the
University of Leipzig (Germany). The corpora were specially preprocessed for the Morpho Challenge
(tokenized, lower-cased, some conversion of character encodings).</p>
      <p>
        The Arabic text data (135K sentences with 3.9M words) is the same as used by Habash and
Sadat [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Because this text data is unfortunately not freely available, only a list of word forms
was provided, so if the participants wanted to use typical word contexts in training their models
in Arabic, they had to find their own text corpus. All words in the Arabic data were presented
in Buckwalter transliteration2. In other languages the lists of word forms to be analyzed were
extracted from the Wortschatz corpora and included all the different word forms existing there
and their frequencies in the corpora. The total amount of word types were 2,206,719 (Finnish),
617,298 (Turkish), 1,266,159 (German), 384,903 (English), and 143,966 (Arabic).
1http://corpora.informatik.uni-leipzig.de/
2http://www.qamus.org/transliteration.htm
      </p>
      <p>
        The exact syntax of the word lists and the required output lists with the suggested morpheme
analyses were explained previously in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As the learning is unsupervised, the returned morpheme
labels may be arbitrary: e.g., ”foot”, ”morpheme42” or ”+PL”. The order in which the morpheme
labels appear after the word forms does not matter. Several interpretations for the same word can
also be supplied, and it was left to the participants to decide whether they would be useful in the
task, or not.
3
3.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Reference analysis</title>
      <sec id="sec-3-1">
        <title>Linguistic Gold Standard</title>
        <p>In Competition 1 the proposed unsupervised morpheme analyses were compared to the correct
grammatical morpheme analyses called here the linguistic gold standard. The gold standard
morpheme analyses were prepared in exactly the same format as the result file the participants
were asked to submit, alternative analyses separated by commas. See Table 1 for examples.</p>
        <p>
          The gold standard reference analyses were the same as in the Morpho Challenge 2007 [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ],
except in Arabic. The Arabic gold standard analyses are based on the representation of lexeme
and features used in the Aragen system (a wrapper using publicly available BAMA-1 databases)
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The first part of an analysis is a lexeme followed by a list of features. The original features
were here modified to connect the POS label to the root of the word, e.g. “Algbn = gabon POS:N
Al+ +SG”. In addition, the gender morphemes were removed (e.g. the German gold standard
doesn’t contain these either). This did not affect the ranking of the submissions, but made the
evaluation resemble more the other tested languages.
        </p>
        <p>In the word lists described in the previous section, the gold standard analyses were available for
650,169 (Finnish), 214,818 (Turkish), 125,641 (German), 63,225 (English), and 141,876 (Arabic)
word types.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Morfessor</title>
        <p>
          As baseline results for unsupervised morpheme analysis, the organizers provided morpheme
analysis by a publicly available unsupervised algorithm called “Morfessor Categories-MAP” developed
at Helsinki University of Technology [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] (or here “Morfessor catmap” or “Morfessor MAP”, for
short as in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]). Analysis by the original Morfessor [
          <xref ref-type="bibr" rid="ref2 ref4">2, 4</xref>
          ] (or here “Morfessor baseline”), which
provides only a surface-level segmentation, was also provided for reference.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Participants and their submissions</title>
      <p>By the submission DL at the end of June, 2008, four research groups had submitted nine
different algorithms which were then evaluated by the organizers. After the DL, more submissions
were received from another author (Goodman), which were evaluated separately outside the
Competition 1. One group (Can) decided not to submit the final wordlists that could be evaluated and
one (McNamee) wanted only to participate in Competition 2. Thus, the final amount of evaluated
algorithms was nine: six in Competition 1, one outside the competition, and two reference results
by Morfessor. The algorithm submissions and their authors are listed in Table 2.</p>
      <p>Some characteristics of morpheme analyses proposed by the unsupervised algorithms together
with the gold standard analyses are briefly presented in Tables 3 and 4. The statistics of each
submission include the average amount of alternative analyses per word, the average amount
of morphemes per analysis, and the total amount of morpheme types. The “Allomorfessor” is
an extension to the “Morfessor Baseline” that attempts to discover common baseforms for the
different surface forms that are likely to represent the same morpheme. The “ParaMor” is another
algorithm for segmenting words into morphemes which, after improvements from the previous
Morpho Challenge, was submitted also as a combination with the publicly available “Morfessor
CATMAP”. The “Zeman 1” is a resubmission from the previous Morpho Challenge which, after
attempts to include a new treatment of prefix, was submitted as the “Zeman 3”. It is interesting
to note that this year all the algorithms resulted in a very large lexicon, usually much larger than
the reference methods did.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Evaluation</title>
      <p>
        The evaluation of Competition 1 in Morpho Challenge 2008 was similar as in Morpho Challenge
2007 except that there was one new language, Arabic. The full description of the method to
compare the submitted unsupervised morpheme analyses were to the linguistic gold standard
analyses is in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In the current paper we just remind the main points and obtained performance
measures.
      </p>
      <p>Because the morpheme analysis candidates are achieved by unsupervised learning, the
morpheme labels can be arbitrary and different from the ones designed by linguists. The basis of the
evaluation is, thus, to compare whether any two word forms that contain the same morpheme
according to the participants’ algorithm also has a morpheme in common according to the gold
standard and vice versa. In practice, the evaluation is performed by randomly sampling a large
number of morpheme sharing word pairs from the compared analyses. Then the precision is
calculated as the proportion of morpheme sharing word pairs in the participant’s sample that really has</p>
      <p>Example word: popUlerliGini #a
popUler liGini 1
popUlerl +i +G +in +i 1
pop/STM +U/SUF +ler/SUF +liGini/SUF 1
pop/STM +U/SUF +ler/SUF +liGini/SUF,
popUlerl +i +G +in +i 2
popUlerliGin i, popUlerliGi ni 3.24
popU lerliGi ni, popU lerliGin i,
popU lerliGini, popUlerliGi ni, popUlerliGin i 1.14
popUler liGini 1
pop +U +ler +liGini 1
popUler +DER lHg +POS2S +ACC,
popUler +DER lHg +POS3 +ACC3 1.99
Example word: AlmtHdp
AlmtHd +p
+Al/PRE mtHd/STM +p/SUF
+Al/PRE mtHd/STM +p/SUF, AlmtHd +p
AlmtHdp, AlmtHd p, AlmtH dp
AlmtHdp
Al mtHdp
Al/PRE mtHd/STM p/SUF
mut aHidap POS:PN Al+ +SG,
mut aHid POS:AJ Al+ +SG
#a
1
1
2
2.24
1.23
1
1</p>
      <p>Example word: baby-sitters
baby- sitters
bab +y, sitt +er +s
+baby-/PRE sitter/STM +s/SUF
+baby-/PRE sitter/STM +s/SUF,
bab +y, sitt +er +s
baby-sitter s, baby-sitt ers
baby-sitt ers, baby-sitter s
baby- sitters
baby - sitters
baby N sit V er s +PL</p>
      <p>The results of the linguistic evaluation are presented in Tables 5 and 6. The tasks in
Competition 1 were the same as in Morpho Challenge 2007, so it is possible to directly compare the
improvements made over the previous algorithms. However, direct comparisons between the
evaluation measures in different languages are not valid, because the corpora and gold standards are
different. In all tasks except the English one, improvements were made in 2008 and the best
obtained F-measure was now higher. As clearly seen in Tables 5 and 6, this is mainly due to
the improved version of “Monson paramor+morfessor” that dominated all tasks. The difference is
especially clear in the recall statistics where the performance of the “Monson paramor+morfessor”
is superior. Behind Monson’s algorithms, the “Zeman 1” that is a re-submission from last year,
was better than the rest of the algorithms, which all suffered from a very low recall. It is worth
noting that the “Kohonen allomorfessor” algorithm achieved clearly the highest precision of all
algorithms in all tasks, but due to the low recall, or undersegmentation, it got rather low F-measure
values.</p>
      <p>
        From the Competition 1 in Morpho Challenge 2007 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], only the winner “best 2007” in each
task was chosen in Tables 5 and 6 for reference. The “Monson paramor+morfessor” was able to
clearly beat the publicly available reference methods “Morfessor baseline” and “Morfessor catmap”
in all tasks. It is interesting to note that the “Morfessor baseline”, which is the original simpler
Morfessor version and only attempts to split words into morphemes without any further analysis,
actually beats the more sophisticated “Morfessor catmap”, as well as “Monson morfessor” and
“Zeman 1”, in English and Arabic. Otherwise, the ranking between the different 2008 algorithms
remains the same in all tasks.
7
      </p>
    </sec>
    <sec id="sec-6">
      <title>Discussions and Conclusions</title>
      <p>The Morpho Challenge 2008 was a successful follow-up to our previous Morpho Challenges 2005
and 2007. Since the main tasks were unchanged, the participants of the previous challenges were
able to track improvements of their algorithms. It also gave a possibility for the new participants
and those who missed the previous deadlines to try more established benchmark tasks. This year
the evaluation was performed also in Arabic, and despite the relatively small wordlist and the
disability to distribute a relevant text corpus, this task was again successful in finding significant
differences between the submitted algorithms.</p>
      <p>The significance of the differences in F-measure was analyzed for all algorithm pairs in all
tasks using the t-test. The analysis was performed by splitting the data into several partitions
and comparing the results in each independent partition separately. The results of the tests show
that all differences were statistically significant, except “Zeman 1” vs “Morfessor catmap” in the
English task.</p>
      <p>As already noted in the previous section, the ranking of the algorithms would have been very
different, if only the precision measure was utilized. Some of the methods, especially “Kohonen
allomorfessor” undersegmented the word forms heavily, which produced high precision but low
recall. However, because it is difficult to estimate the relative weight of precision against recall
in different applications, it remains for the application based evaluations in different tasks to
show which algorithms are most useful. Many of the grammatical morphemes (such as +PL and
+PAST in Table 1) are very common and may not be very relevant in IR, for example, compared
to recognizing the right stem.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We thank all the participants for their submissions and enthusiasm. We owe great thanks as
well to the organizers of the PASCAL Challenge Program and CLEF who helped us organize
this challenge and the challenge workshop. We are most grateful to the University of Leipzig
for making the training data resources available to the Challenge, and in particular we thank
Stefan Bordag for his kind assistance. We are indebted to Ebru Arisoy for making the Turkish
gold standard available to us. We are most grateful to the Nizar Habash from the University
of Columbia for his kind assistance and making the Arabic word frequency list and reference
analyses available to the Challenge. Our work was supported by the Academy of Finland in
the projects Adaptive Informatics and New adaptive and learning methods in speech recognition.
This work was supported in part by the IST Programme of the European Community, under the
PASCAL Network of Excellence, IST-2002-506778. This publication only reflects the authors’
views. We acknowledge that access rights to data and other materials are restricted due to other
commitments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Jeff</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bilmes</surname>
            and
            <given-names>Katrin</given-names>
          </string-name>
          <string-name>
            <surname>Kirchhoff</surname>
          </string-name>
          .
          <article-title>Factored language models and generalized parallel backoff</article-title>
          .
          <source>In Proceedings of the Human Language Technology</source>
          ,
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL)</article-title>
          , pages
          <fpage>4</fpage>
          -
          <lpage>6</lpage>
          , Edmonton, Canada,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Creutz</surname>
          </string-name>
          and
          <string-name>
            <given-names>Krista</given-names>
            <surname>Lagus</surname>
          </string-name>
          .
          <article-title>Unsupervised discovery of morphemes</article-title>
          .
          <source>In Proceedings of the Workshop on Morphological and Phonological Learning of ACL-02</source>
          , pages
          <fpage>21</fpage>
          -
          <lpage>30</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Creutz</surname>
          </string-name>
          and
          <string-name>
            <given-names>Krista</given-names>
            <surname>Lagus</surname>
          </string-name>
          .
          <article-title>Inducing the morphological lexicon of a natural language from unannotated text</article-title>
          .
          <source>In Proceedings of the International and Interdisciplinary Conference on Adaptive Knowledge Representation and Reasoning (AKRR'05)</source>
          , pages
          <fpage>106</fpage>
          -
          <lpage>113</lpage>
          , Espoo, Finland,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Creutz</surname>
          </string-name>
          and
          <string-name>
            <given-names>Krista</given-names>
            <surname>Lagus</surname>
          </string-name>
          .
          <article-title>Unsupervised morpheme segmentation and morphology induction from text corpora using Morfessor</article-title>
          .
          <source>Technical Report A81</source>
          , Publications in Computer and Information Science, Helsinki University of Technology,
          <year>2005</year>
          . URL: http://www.cis.hut.fi/projects/morpho/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Nizar</given-names>
            <surname>Habash</surname>
          </string-name>
          .
          <article-title>Large scale lexeme based arabic morphological generation</article-title>
          .
          <source>In Proceedings of Traitement Automatique du Langage Naturel (TALN-04)</source>
          , Fez, Morocco,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Nizar</given-names>
            <surname>Habash</surname>
          </string-name>
          and
          <string-name>
            <given-names>Fatiha</given-names>
            <surname>Sadat</surname>
          </string-name>
          .
          <article-title>Arabic preprocessing schemes for statistical machine translation</article-title>
          .
          <source>In Proceedings of the Human Language Technology</source>
          ,
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL)</article-title>
          , New York, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Mikko</given-names>
            <surname>Kurimo</surname>
          </string-name>
          , Mathias Creutz, and
          <string-name>
            <given-names>Ville</given-names>
            <surname>Turunen</surname>
          </string-name>
          .
          <article-title>Unsupervised morpheme analysis evaluation by IR experiments - Morpho Challenge 2007</article-title>
          .
          <source>In Working Notes for the CLEF 2007 Workshop</source>
          , Budapest, Hungary,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Mikko</given-names>
            <surname>Kurimo</surname>
          </string-name>
          , Mathias Creutz, and
          <string-name>
            <given-names>Matti</given-names>
            <surname>Varjokallio</surname>
          </string-name>
          .
          <article-title>Unsupervised morpheme analysis evaluation by a comparison to a linguistic Gold Standard - Morpho Challenge 2007</article-title>
          .
          <source>In Working Notes for the CLEF 2007 Workshop</source>
          , Budapest, Hungary,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Mikko</given-names>
            <surname>Kurimo</surname>
          </string-name>
          , Mathias Creutz, Matti Varjokallio, Ebru Arisoy, and
          <string-name>
            <given-names>Murat</given-names>
            <surname>Saraclar</surname>
          </string-name>
          .
          <source>Unsupervised segmentation of words into morphemes - Challenge</source>
          <year>2005</year>
          ,
          <article-title>an introduction and evaluation report</article-title>
          .
          <source>In PASCAL Challenge Workshop on Unsupervised segmentation of words into morphemes</source>
          , Venice, Italy,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Mikko</given-names>
            <surname>Kurimo</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ville</given-names>
            <surname>Turunen</surname>
          </string-name>
          .
          <article-title>Unsupervised morpheme analysis evaluation by IR experiments - Morpho Challenge 2008</article-title>
          .
          <source>In Working Notes for the CLEF 2008 Workshop</source>
          , Aarhus, Denmark,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.-S.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>Morphological analysis for statistical machine translation</article-title>
          .
          <source>In Proceedings of the Human Language Technology</source>
          ,
          <article-title>Conference of the North American Chapter of the Association for Computational Linguistics (HLT-NAACL)</article-title>
          , Boston, MA, USA,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Sami</surname>
            <given-names>Virpioja</given-names>
          </string-name>
          , Jaakko J. V¨ayrynen, Mathias Creutz, and
          <string-name>
            <given-names>Markus</given-names>
            <surname>Sadeniemi</surname>
          </string-name>
          .
          <article-title>Morphologyaware statistical machine translation based on morphs induced in an unsupervised manner</article-title>
          .
          <source>In Proceedings of Machine Translation Summit XI</source>
          , Copenhagen, Denmark,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.L.</given-names>
            <surname>Zieman</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.L.</given-names>
            <surname>Bleich</surname>
          </string-name>
          .
          <article-title>Conceptual mapping of user's queries to medical subject headings</article-title>
          .
          <source>In Proceedings of the 1997 American Medical Informatics Association (AMIA) Annual Fall Symposium</source>
          ,
          <year>October 1997</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>