<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erik M. van Mulligen</string-name>
          <email>e.vanmulligen@erasmusmc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quoc-Chinh Bui</string-name>
          <email>q.bui@erasmusmc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jan A. Kors</string-name>
          <email>j.kors@erasmusmc.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Medical Informatics, Erasmus University Medical Center Rotterdam</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe how we applied a general- purpose machine translation tool for translating biomedical thesauri. We used corresponding terms in parallel corpora to check the validity of the translations. The advantage of this approach is that a single corresponding set of terms can be verified where techniques to retrieve translations from a parallel corpus do not exploit the knowledge contained in current state of the art machine translation software.</p>
      </abstract>
      <kwd-group>
        <kwd>machine translation</kwd>
        <kwd>multilingual</kwd>
        <kwd>concept annotation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>second translated thesaurus only based on the existing non-English terms in the
UMLS. This UMLS-translated thesaurus was included as a baseline thesaurus to
assess the improvement that could be obtained with the machine translation. We
applied this approach for two languages, Dutch and German.</p>
      <p>We used two parallel corpora: multilingual drug labels from the EMEA corpus, and
bilingual titles of scientific abstracts from Medline.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>We indexed the parallel corpora with the English thesaurus, the UMLS-translated
thesaurus, and the machine-translated thesaurus. We differentiated for the different
semantic types and the different parallel corpora we have (for German restricted to
the drugs labels from EMEA and MedLine titles The results for Dutch, German,
and French are show in table 1. Each table shows the results for finding concepts
using only the manual translation as contained in the UMLS and the results when
the machine translated terms are added. We provide not only the figures for the
translated terms that have correspondences in English and the translation language
(BOTH), but also terms that have only be found in English (ENGLISH) or only in
the translation language (DUTCH, GERMAN, or FRENCH). The tables show the
results for the EMEA drug label parallel corpus and for the Medline titles.</p>
      <p>Tabel 1. Results for the Dutch, German, and French EMEA drug labels and Medline titles.</p>
      <p>ENTRY</p>
      <p>OBJC</p>
      <p>GEOG</p>
      <p>CHEM</p>
      <p>DEVI</p>
      <p>PHEN</p>
      <p>DISO</p>
      <p>ANAT</p>
      <p>LIVB</p>
      <p>PHYS</p>
      <p>TOTAL
ENGLISH
BOTH
FRENCH
ENGLISH
BOTH
FRENCH
ENGLISH
BOTH
FRENCH
ENGLISH
BOTH
Language
Dutch
German
French
102
Tabel 2. Overall statistics for the different languages and corpora.</p>
      <p>EMEA</p>
      <p>Medline
Original</p>
      <p>Translated</p>
      <p>Original</p>
      <p>Translated
4</p>
    </sec>
    <sec id="sec-3">
      <title>Discussion</title>
      <p>The results show that machine translation can help to enrich a thesaurus. Compared
with the manual UMLS-translated thesaurus, the number of terms in the
machinetranslated thesaurus that are found in the parallel corpora, doubles for German,
French and almost triples for Dutch when only considering concepts that have been
found in English. This is consistent for both parallel corpora included in this
evaluation. The German and French manual translations are more extensive than
the Dutch one, which likely explains the difference in number of extra terms found.
The increase is largest for some semantic groups that have hardly been translated
(objects, devices, and chemicals). If one also looks at terms that have only been
found in the translated corpus and not in the original English corpus an additional
set of translated terms can be found. We will also evaluate this set of terms only
found in the translated corpus for correctness. We will extend this evaluation to
include Spanish as well.</p>
      <p>van der Plas L, Tiedemann J. Finding medical term variations using parallel corpora
and distributional similarity. In: Proc 6th Workshop on Ontologies and Lexical
Resources, 2010.</p>
      <p>Och FJ, Ney H. A systematic comparison of various statistical alignment models.
Comp Ling 2003;29:19-51.</p>
      <p>Schulz S. et al. Machine vs. human translation of SNOMED CT terms. Currently
under review for MEDINFO 2013
https://developers.google.com/translate/
Bodenreider and O, McCray, AT. Exploring semantic groups through visual
approaches, Journal of Biomedical Informatics 36(6):414–432, 2003.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>Methods Inf Med</source>
          . 1993 Aug;
          <volume>32</volume>
          (
          <issue>4</issue>
          ):
          <fpage>281</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          http://mantra-project.eu
          <string-name>
            <surname>Deleger</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merkel</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            <given-names>P</given-names>
          </string-name>
          .
          <article-title>Translating medical terminologies through word alignment in parallel text corpora</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <year>2009</year>
          ;
          <volume>42</volume>
          :
          <fpage>692</fpage>
          -
          <lpage>701</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>