<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Supervised Named Entity Recognition for Clinical Data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Devanshu Jain</string-name>
          <email>devanshu.jain919@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dhirubhai Ambani Institute of Information and Communication Technology</institution>
          ,
          <addr-line>Gandhinagar, Gujarat</addr-line>
          ,
          <country country="IN">India 382007</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Clinical Named Entity Recognition is a part of Task 1b, organised by CLEF eHealth organisation in 2015. The aim is to automatically identify clinically relevant entities in medical text in French. A supervised learning approach has been used for training the tagger. For the purpose of training, Conditional Random Fields(CRF) has been used. An extensive set of features was used for training. Precision, recall and F1 Score were used as evaluation metrics. Ten fold cross validation technique was used to evaluate the system. The best precision obtained was 0.91 and the best recall obtained was 0.66. After the test results were announced, the best F1 score obtained for exact matching was 0.67 and for relaxed case (i.e. inexact matching), it was 0.73.</p>
      </abstract>
      <kwd-group>
        <kwd>Clinical Data</kwd>
        <kwd>Named Entity Recognition</kwd>
        <kwd>UMLS</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>CRF</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>There is a huge amount of raw medical data available in the form of textual
information. The goal of Information Extraction, here, is to present the data
in a way that enhances user experience and allows better comprehension. The
major task involved in this process is the identi cation of named entities within
the document. This allows the user to have a better understanding of the jargons.
It also allows to identify important terms that may be helpful in summarising
the medical document.</p>
      <p>The Clinical Named Entity Recognition is di erent from other common
sequence tagging problems, like POS (Part of Speech) tagging. The major point of
di erence is the existence of ambiguity in the medical document. The span of an
entity may overlap with the span of another entity i.e. the same word can be a
part of multiple entities. Another issue with it is the presence of non contiguous
entities. That is, the span of the entity may be discontinuous over a sentence.
Moreover, there are ample resources such as thesaurus, but almost all of them
are English centric. The training data, being in French language, poses another
challenge.</p>
      <p>In order to annotate the clinical entities, UMLS (Uni ed Medical Language
System) is used. It is a compendium of vocabularies in biomedical science. There
are various semantic groups, under which an entity can lie.</p>
      <p>In order to tackle the problem, an extensive list of features were used.
CRFsuite software was used to train the tagger. The following sections explains in
greater details the methods and tools used for tackling the problem.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Approach</title>
      <sec id="sec-2-1">
        <title>Overview</title>
        <p>The available data was rst pre-processed by stemming all the words. We
have considered the NER task as a sequence tagging problem. We, therefore have
used CRF basic tagger to train the system. CRF is a state of the art method
used for the purpose of sequence tagging. We have used CRFsuite software for
this purpose.
1. Lexical Features: Uni-grams, bi-grams and tri-grams of words were used as
features within a window of 3 around the current word. Then the POS tags
of the words in the window were also chosen as features. Finally, we used
capitalisation of letters and presence of digit as features too. In addition,
pre xes and su xes of 3 letter length were also used as features for each
word.
2. UMLS Features: UMLS features were extracted using MetaMap. Since,
MetaMap is English-centric, therefore, all the French words in the training
data were translated into English, separately. Then, using MetaMap API,
semantic group of these words were obtained and used as features for
training.
3. Global Features: For every word, we calculated the position of the word
in that sentence. To do this, we treated the word as LEFT, when it lied
in the left quarter of the sentence. When it lied in the right quarter of the
sentence, it was treated as being RIGHT. Otherwise, it was treated as being
CENTER.</p>
        <p>In order to account for the case when the span of entities was over multiple
words in a contiguous manner, we used BIO format. For the starting of an entity,
its name was pre xed with -B. If it was an intermediate word, the entity name
was pre xed with -I. Otherwise, it was named as O. The system currently does
not handle the discontinous terms.
3</p>
        <p>Tools Used
1. Microsoft Bing translator was used for translating each french word
(separately) into English.
2. Snowball stemmer was used for the stemming during pre-processing step
3. CRFSuite tagger was used to train the model based on the traing data and
for tagging the test les.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Training Data</title>
      <p>The training data was provided by the CLEF organisation itself. The data
consisted of 833 MEDLINE documents, that contained single lines of medical data
in French language. An additional 11 EMEA documents were also provided.
Annotation les for each document were also given.</p>
      <p>Some statistics of the training data is as follows:</p>
      <p>Total Word Count 25,500</p>
      <p>Number of Annotations 5,690
Non Contiguous Annotations 40
Overlapping Spans of entities 797</p>
      <p>As can be seen, the non contiguous annotations account for just 0.7% of the
total number of annotations and hence don't a ect the validations, much.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>Precision, Recall and F1 measure were used as evaluation metrics to evaluate
the system. These are de ned as follows:</p>
      <p>T rueP ositives
P recision = T rueP ositives+F alsepositives</p>
      <p>T rueP ositives
Recall = T rueP ositives+F alseNegatives</p>
      <p>F 1Score = 2 P recision Recall</p>
      <p>P recision+Recall</p>
      <p>We used ten fold cross validation techniques. The available data was
partitioned into training data and validation data. The partition was done randomly,
i.e. random samples were taken from the data and used for validation. The
partition was done four times in the ratio of 60:40, 70:30, 80:20 and 90:10. Ten
di erent and random partitions for each ratio were created. Then, the average
precision and recall was calculated.</p>
      <p>When MetaMap was not used as an external source of information, following
results were obtained:
1. Run 1: Predictions were completely made by CRFsuite based on the model
generated using training data.
2. Run 2: Predictions made use of CRFsuite software as well as UMLS
Metathesaurus information obtained by MetaMap. Whenever the model
tagged a token as a non-entity, information from MetaMap was used to
specify the tag for it. It was done at the level of single word only.
The runs were submitted for entity identi cation only. We didn't participate in
entity normalisation.
6.2</p>
      <sec id="sec-4-1">
        <title>Results</title>
        <p>For the EMEA documents, following results were obtained:</p>
        <p>Run
Run 1
Run 2</p>
        <p>Exact Match Inexact Match
Precision Recall F1 Score Precision Recall F1 Score
We present a supervised Clinical named Entity Recognition system that can
detect the named entities from a French medical data, using an extensive list of
features, with an F1 Score of 0.68. It also uses UMLS Metathesaurus information
obtained by MetaMap, using the English translated version of French words.</p>
        <p>The surprising thing to observe was that although, we used extra UMLS
Metathesaurus information obtained by MetaMap in run2, although it improved
the precision, but reduced the recall drastically. This remains to be investigated.</p>
        <p>The system does not handle the non contiguous entities properly. We can use
mutual information as a technique to identify the relationship between di erent
entities, in order to detect that.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Clef e health:
          <source>Task</source>
          <volume>1b</volume>
          <year>2015</year>
          . https://sites.google.com/site/clefehealth2015/ task-1/task-1b.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Crfsuite</surname>
          </string-name>
          <article-title>: a fast implementation of conditional random elds (crfs)</article-title>
          . http://www. chokkan.org/software/crfsuite/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F. S. F. I. K. C.</given-names>
            <surname>Czajkowski</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>Conditional random elds: Probabilistic models for segmenting and labeling sequence data</article-title>
          .
          <source>In Proceedings of the Eighteenth International Conference on Machine Learning.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>A. A.</surname>
          </string-name>
          et. al.
          <article-title>E ective mapping of biomedical text to the umls metathesaurus: the metamap program</article-title>
          .
          <source>In Proceedings of the AMIA Symposium.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Souminen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hanlen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neveol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grouin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Zuccon</surname>
          </string-name>
          .
          <article-title>Overview of theclef ehealth evaluation lab 2015</article-title>
          .
          <source>In CLEF 2015 - 6th Conference and Labs of the Evaluation Forum. Lecture Notes in Computer Science (LNCS)</source>
          , Springer,
          <year>September 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A.</given-names>
            <surname>Neveol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Grouin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tannier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hamon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Zweigenbaum</surname>
          </string-name>
          .
          <article-title>CLEF eHealth evaluation lab 2015 task 1b: clinical named entity recognition</article-title>
          .
          <source>In CLEF 2015 Online Working Notes. CEUR-WS</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>