<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multilingual Named-Entity Recognition from Parallel Corpora</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andreea Bodnari</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aurelie Neveol</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ozlem Uzuner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre Zweigenbaum</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Szolovits</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Studies, University at Albany, SUNY</institution>
          ,
          <addr-line>Albany, New York</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LIMSI-CNRS</institution>
          ,
          <addr-line>rue John von Neumann, F-91400 Orsay</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>MIT, CSAIL</institution>
          ,
          <addr-line>Cambridge, Massachusetts</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present a named-entity recognition (NER) system for parallel multilingual text. Our system handles three languages (i.e., English, French, and Spanish) and is tailored to the biomedical domain. For each language, we design a supervised knowledge-based CRF model with rich biomedical and general domain information. We use the sentence alignment of the parallel corpora, the word alignment generated by the GIZA++[8] tool, and Wikipedia-based word alignment in order to transfer system predictions made by individual language models to the remaining parallel languages. We re-train each individual language system using the transferred predictions and generate a nal enriched NER model for each language. The enriched system performs better than the initial system based on the predictions transferred from the other language systems. Each language model bene ts from the external knowledge extracted from biomedical and general domain resources.</p>
      </abstract>
      <kwd-group>
        <kwd>Named-entity recognition</kwd>
        <kwd>Medical NLP</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Natural Language Processing (NLP) technologies can help extract structured
information from written text, determine relationships between concepts, or
perform bilingual translation. In the biomedical and clinical[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] domain NLP has
multiple applications: computerized clinical decision support, personalized medicine,
automatic extraction of protein interactions and associations of proteins to
functional concepts, to name a few. Despite their value and extensive applicability,
biomedical and clinical NLP systems are scarcely developed for languages other
than English.
      </p>
      <p>
        In this study, we present a system for extracting structured information from
multilingual biomedical corpora. We de ne the structured information of interest
as noun phrases (NPs) that fall under nine Uni ed Medical Language System
(UMLS)[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] semantic categories. Our data consists of parallel multilingual corpora
in English (en), French (fr), and Spanish (es).
      </p>
      <p>We present a NP extraction system for multilingual biomedical corpora that
makes use of external knowledge sources and the common structure shared by
the parallel languages. We analyze the errors performed by the NER system, and
we discuss the contribution of the knowledge sources and the common language
structure to system performance.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Problem de nition</title>
      <p>
        We focus on NPs that fall under nine semantic categories: anatomy (ANAT),
chemicals and drugs (CHEM), devices (DEVI), disorders (DISO), geographic
areas (GEOG), living beings (LIVB), objects (OBJC), phenomena (PHEN), and
physiology (PHYS). These semantic categories are groupings of semantic types
in the UMLS semantic network.[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] Our goal is two-fold: for each of the three
languages, rst identify noun phrases that belong to the semantic categories of
interest and then map the identi ed noun phrases to an already existing Concept
Unique Identi er (CUI) in UMLS.
2
2.1
      </p>
      <sec id="sec-2-1">
        <title>Materials and Methods</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>
        The data used for the system development is made available by the
organizers of the 2013 CLEF-ER challenge[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and comes from two sources: the
European Medicines Agency (EMEA) and Medline. Each corpus contains
sentencedelimited plain text. Sentence alignment information is available for the language
pairs en-fr and en-es. The EMEA corpus contains the same number of sentences
for each language (140522) so that each English sentence has an equivalent
alignment in the other two languages. The Medline corpus contains approx. 1.5
million English sentences, 0.5 million French sentences and 0.25 million Spanish
sentences (see Table 1). In the Medline corpus, half of the English sentences do
not have an equivalent alignment in French or Spanish, while the other half have
an equivalent alignment in either French or Spanish.
We manually annotated all noun phrase instances of the nine semantic types
within a set of 385 sentences, for each language and for each corpus source. We
use di erent annotation processes for the EMEA and the Medline corpora. For
the EMEA corpus, we randomly selected 385 sentences from the automatically
generated gold standard provided by the challenge organizers. For each of the
selected sentences, we transferred the English annotations to the corresponding
French- and Spanish-aligned sentences based on the word alignment provided
by the GIZA++ software and Wikipedia as explained below. We manually
reviewed the transferred annotations and annotated additional noun phrases for
each language when appropriate.
      </p>
      <p>For the Medline corpus, we used the Medical Subject Heading (MeSH)
indexing and the UMLS Metathesaurus to automatically generate noun phrase
annotations. The MEDLINE citations corresponding to the sentences (i.e.,
titles) in the Medline corpus have been manually assigned a set of MeSH indexing
terms by professional indexers at the National Library of Medicine. MeSH main
headings can be derived from this set and mapped to UMLS CUIs. We
randomly selected 385 Medline sentences for each language and made use of the
MeSH indexing-UMLS CUI relationship in order to generate Medline reference
annotations. Speci cally, when a term associated to a CUI from the MeSH
indexing can be found in the corresponding Medline sentence, we automatically create
a reference annotation. We apply the same annotation process to the English,
French, and Spanish versions of the Medline corpus.</p>
      <p>The EMEA and Medline reference annotations were manually reviewed by a
sole annotator, due to time constraints and lack of additional resources.</p>
      <p>The number of unique annotations for EMEA and Medline are similar across
languages (368 English annotations for EMEA and 401 for Medline, 364 French
annotations for EMEA and 339 for Medline, and 368 Spanish annotations for
EMEA and 303 for Medline). The EMEA corpus contains a signi cantly larger
number of actual annotations due to numerous noun phrase duplicates (see Table
2).
2.3</p>
    </sec>
    <sec id="sec-4">
      <title>System design</title>
      <p>
        We perform word alignment using the GIZA++ software.[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] GIZA++ aligns
words based on statistical models. For source sentence sJ = s1s2::sJ and target
sentence tI = t1t2::tI , we de ne an alignment of the two sentences as A f(i; j) :
j = 1; ::; J ; i = 0; ::; Ig, where the case i = 0 represents source words that are not
aligned to any target words. The probability of a source sentence given a target
sentence is P (sjt) = PaJ P (sJ ; a1J jtI ), where a1J represents the sentence pair
alignment. The best sentence alignment is given by ^aJ = argmaxaJ P (sJ ; aJ jtI ).
      </p>
      <p>In addition to the GIZA++ alignment, we use word alignment from Wikipedia
metadata. We rely on the fact that some English Wikipedia articles have direct
correspondents in French and Spanish. We lter the English articles that have
titles consisting of a single word (e.g., \Food", \Disease", \Nausea") and nd
their corresponding foreign language article. The title of the English article and
the title of the correspondent foreign language article represent a word
alignment pair. Because Wikipedia word alignments are more precise, we overwrite
the GIZA++ generated word alignments with a score lower than 0.5 with the
Wikipedia generated word alignments.</p>
      <p>Annotation count</p>
      <p>English French Spanish
ANAT MEMedElinAe 26022 6301 3288
CHEM MEMedElinAe 158586 55159 55782
DEVI MEMedElinAe 4130 63 110
DISO MEMedElinAe 1166588 412870 416754
GEOG MEMedElinAe 3145 191 1232
LIVB MEMedElinAe 47922 16302 13357
OBJC MEMedElinAe 1195 347 328
PHEN MEMedElinAe 727 320 310</p>
      <p>PHYS MEMedElinAe 22792 1523 664</p>
      <p>Total NPs MEMedElinAe 4446911 1328661 1334704
Total Unique NPs EMEA 368 364 368</p>
      <p>Medline 401 339 303
Table 2. Annotations description per corpus type: annotation count, separated by
language and semantic category
Noun phrase identi cation Our system is designed based on the supervised
CRF framework. The CRF model includes lexical features (the normalized form
of the token), syntactic features (part-of-speech and parse tree information of
each token), and knowledge-based features (information extracted from UMLS
and Wikipedia). A nal set of features is generated based on the shared structure
of the parallel languages. All feature sets are used in the nal CRF model.</p>
      <p>
        In order to obtain the parse and POS information we use di erent resources
for each language: for English we use the Stanford parser,[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] while for French and
Spanish we use the Malt parser.[
        <xref ref-type="bibr" rid="ref3 ref6">3, 6</xref>
        ] We extract the knowledge-based features
from UMLS using MetaMap[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] for the English corpus and direct lexical mapping
for the French and Spanish corpora. The lexicons we used contained about 0.3
million terms corresponding to 0.15 million CUIs for French and 0.5 million
terms corresponding to 0.3 million CUIs for Spanish.
      </p>
      <p>The Wikipedia knowledge-based features are dependent on the language and
are extracted based on the respective language version of Wikipedia. We make
use of the fact that Wikipedia articles are tagged with speci c categories. We
map these Wikipedia categories to one of the nine UMLS categories of interest.
Then, we identify the n-grams (n 3) that represent titles of Wikipedia articles
and signal within the attribute vector whether an n-gram belongs to one of the
nine categories of interest based on its Wikipedia categorization.</p>
      <p>Using the lexical, syntactic, and knowledge-based features described above,
we create a CRF model (the Initial model) for each language and label the
training data using the relevant model. We further make use of the common structure
shared by the parallel corpora and transfer the NP predictions across aligned
bilingual sentence pairs. We hypothesize that some language models might
predict NPs that other language models would miss, and by using sentence- and
word-alignment information we can inform the other language models of the
missed predictions. The CRF model containing the transferred predictions is
referred to as the enriched model. The entire system design is depicted in Figure
1.
The NER system predicts the text of the NPs together with their semantic
categories. In order to obtain the CUI associated with each NP we use MetaMap
for English and direct lookup inside the UMLS database for the French and
Spanish languages. Because the French and Spanish versions of UMLS do not
contain the same number of concepts as the English UMLS, there are French and
Spanish concepts identi ed by our system that cannot be linked to a CUI via
direct UMLS lookup. For those concepts we rely on word alignment to determine
the English concept to which the foreign language concepts are aligned; once an
aligned English concept is found we transfer the CUI information to the foreign
language concept.
3</p>
      <sec id="sec-4-1">
        <title>Discussion</title>
        <p>The reference annotations we generated for this challenge relied partially on
automated annotation tools. These annotation tools generated an imperfect
annotation output and generally failed to identify a percentage of the noun-phrase
instances. For the Medline corpus, the annotation tools failed to identify NPs
in the sentences that belonged to oldest citations. These citations were only
assigned a couple of MeSH indexing terms, vs. about a dozen for more recent
citations. Also, because the MeSH indexing terms are assigned based on the
fulltext article content and not on the article title, they might be linked to CUIs
not present in the title. A similar issue arises when the Medline title contains
a shortened or altered version of the CUI string indexed by MeSH (e.g. the
string \hypertensive patient" occurring in the title of citation 1838917 could not
be reconciled with CUI C0020538 \hypertension"). For EMEA, the main
problem was the transfer of English annotations based on the word alignment: the
GIZA++ software does not generate a perfect word alignment, and even though
the Wikipedia word alignment manages to x some incorrect alignments, there
are still cases of incorrect or missing alignments.</p>
        <p>The problem of the noisy word alignment impacts both the generation of the
reference annotations and the performance of the enriched system. Speci cally,
the predictions transferred between the systems are usually incomplete (e.g.,
for the English noun phrase active substance, the word alignment could only
help transfer the labeling for the token \substance" and failed to transfer the
labeling for the token \active" as it could not nd a valid word French or Spanish
alignment for the word \active").</p>
        <p>Mapping the noun phrases to a CUI is a di cult problem for the French and
Spanish languages since the number of annotated CUIs is relatively small. A
large number of UMLS English concepts are not present in the foreign language
portions of UMLS; thus, even though the English translation of a French or
Spanish noun phrase is present in UMLS, the French or Spanish noun phrase
is not included. Our system has to rely on machine translation or noun phrase
alignment for mapping certain French and Spanish noun phrases to their CUIs.</p>
        <p>Based on manual revision, the noun phrases identi ed by the enriched model
present a higher precision and slightly lower recall. In general, the enriched model
predicts more complete noun phrases (e.g., \insu sance hepatique severe" vs.
\insu sance hepatique", \diminution des bruits respiratoires", vs. \bruits
respiratoires", \paresies des cordes vocales" vs. \des paresies des cordes", \tumor
principal" vs. \tumor", \movimientos involuntarios" vs. \involuntarios"). We
notice that the semantic categories of the noun phrases are better assigned when
the enriched model is used (e.g., \linfocitos" assigned a LIVB category by the
initial model and an ANAT category by the enriched model), and even
correctly adjusted together with the span of the noun phrase (e.g., noun phrase
\mand bula" classi ed as ANAT in the initial model is changed into
osteonecrosis de la mandbula classi ed as DISO in the enriched model).
4</p>
      </sec>
      <sec id="sec-4-2">
        <title>Conclusion</title>
        <p>We present a multilingual NER extraction system targeting the biomedical
domain. The NER system relies on a supervised learning framework enriched with
external knowledge and cross-linguistic information transferred based on
common structure between languages. We prepare reference annotations for English,
French, and Spanish that together with the NER system can be used for
developing additional multilingual resources. We illustrated the improvements brought
by our second round of training based on cross-language transfer of initial
annotations.
5</p>
      </sec>
      <sec id="sec-4-3">
        <title>Acknowledgments</title>
        <p>This work was funded in part by award number 2U54LM008748 from the
National Institutes of Health (NIH)/National Library of Medicine (NLM) and
cofounded by the National Heart, Lung and Blood Institute (NHLBI), and by
contract number 90TR0002 (SHARPd Secondary Use of Clinical Data) from the
O ce of the National Coordinator (ONC) for Health Information Technology.
The main author was funded by the Chateaubriand Science Fellowship 2012-13.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alan R Aronson and Francois-Michel Lang</surname>
          </string-name>
          .
          <article-title>An overview of metamap: historical perspective and recent advances</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <volume>17</volume>
          (
          <issue>3</issue>
          ):
          <volume>229</volume>
          {
          <fpage>236</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Olivier</given-names>
            <surname>Bodenreider</surname>
          </string-name>
          .
          <article-title>The uni ed medical language system (umls): integrating biomedical terminology</article-title>
          .
          <source>Nucleic acids research</source>
          ,
          <volume>32</volume>
          (
          <issue>suppl 1</issue>
          ):D267{
          <fpage>D270</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Marie</given-names>
            <surname>Candito</surname>
          </string-name>
          , Beno^t Crabbe,
          <string-name>
            <surname>Pascal Denis</surname>
          </string-name>
          , et al.
          <article-title>Statistical french dependency parsing: treebank conversion and rst results</article-title>
          .
          <source>In Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2010</year>
          ), pages
          <year>1840</year>
          {
          <year>1847</year>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Dina</given-names>
            <surname>Demner-Fushman</surname>
          </string-name>
          , Wendy W Chapman, and
          <string-name>
            <surname>Clement J McDonald</surname>
          </string-name>
          .
          <article-title>What can natural language processing do for clinical decision support? Journal of biomedical informatics</article-title>
          ,
          <volume>42</volume>
          (
          <issue>5</issue>
          ):
          <volume>760</volume>
          {
          <fpage>772</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Dan</given-names>
            <surname>Klein</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>Accurate unlexicalized parsing</article-title>
          .
          <source>In Proceedings of the 41st Annual Meeting on Association for Computational LinguisticsVolume 1</source>
          , pages
          <fpage>423</fpage>
          {
          <fpage>430</fpage>
          . Association for Computational Linguistics,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Montserrat</given-names>
            <surname>Marimon</surname>
          </string-name>
          , Nuria Bel, Sergio Espeja, and
          <string-name>
            <given-names>Natalia</given-names>
            <surname>Seghezzi</surname>
          </string-name>
          .
          <article-title>The spanish resource grammar: pre-processing strategy and lexical acquisition</article-title>
          .
          <source>In Proceedings of the Workshop on Deep Linguistic Processing</source>
          , pages
          <volume>105</volume>
          {
          <fpage>111</fpage>
          . Association for Computational Linguistics,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Alexa T McCray</surname>
          </string-name>
          ,
          <string-name>
            <surname>Anita Burgun</surname>
            ,
            <given-names>Olivier</given-names>
          </string-name>
          <string-name>
            <surname>Bodenreider</surname>
          </string-name>
          , et al.
          <article-title>Aggregating UMLS semantic types for reducing conceptual complexity</article-title>
          .
          <source>Studies in health technology and informatics, (1):</source>
          <volume>216</volume>
          {
          <fpage>220</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>FJ</given-names>
            <surname>Och</surname>
          </string-name>
          and
          <string-name>
            <given-names>H</given-names>
            <surname>Ney</surname>
          </string-name>
          .
          <article-title>Giza++: Training of statistical translation models</article-title>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Hanna</given-names>
            <surname>Suominen</surname>
          </string-name>
          , Sanna Salantera Sumithra Velupillai, Wendy W. Chapman, Guergana Savova, Noemie Elhadad, Sameer Pradhan, Brett R. South, Danielle Mowery,
          <string-name>
            <given-names>Gareth J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , Johannes Leveling, Liadh Kelly, Lorraine Goeuriot, David Martinez,
          <string-name>
            <given-names>and Guido</given-names>
            <surname>Zuccon</surname>
          </string-name>
          .
          <article-title>Overview of the ShARe/CLEF eHealth Evaluation Lab 2013</article-title>
          .
          <source>In Proceedings of CLEF</source>
          <year>2013</year>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>