<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Clinical Name Entity Recognition using Conditional Random Field with Augmented Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dawei Geng (Intern at Philips Research China</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shanghai)</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>In this paper, We presents a Chinese medical term recognition system submitted to the competition held by China Conference on Knowledge Graph and Semantic Computing. I compare the performance of Linear Chain Conditional Random Field (CRF) with that of Bi-Directional Long Short Term Memory (LSTM) with Convolutional Neural Network (CNN) and CRF layers performance and find that CRF with augmented features performs best with F1 0.927 on the offline competition dataset using cross-validation. Hence, this system was built by using a conditional random field model with linguistic features such as character identity, N-gram, and external dictionary features.</p>
      </abstract>
      <kwd-group>
        <kwd>Linear-Chain Conditional Random Field</kwd>
        <kwd>Name Entity Recognition</kwd>
        <kwd>Long Short Term Memory</kwd>
        <kwd>Convolutional Neural Network</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>named entities do not follow any nomenclature, which makes rule-based methods
hard to be perfect. Besides, rule-based systems require domain experts, and they are
not flexible to other NE types and domains.</p>
      <p>
        Machine learning methods are more robust and they can identify potential
biomedical entities which are not previously included in standard dictionaries. More
and more machine learning methods are explored to solve the Bio-NER problem, such
as Hidden Markov Model (HMM), Support Vector Machine (SVM), Maximum
Entropy Markov Model (MEMM), and Conditional Random Fields (CRF)
        <xref ref-type="bibr" rid="ref1">(Lafferty et
al., 2001)</xref>
        .
      </p>
      <p>
        Further, many deep learning methods are employed to tag sequence data. For
example, Convolutional network based models (Collobert et al., 2011) have been
proposed to tackle sequence tagging problem. Such model consists of a convolutional
network and a CRF layer on the output. In speech language understanding community,
recurrent neural network
        <xref ref-type="bibr" rid="ref7 ref9">(Mesnil et al., 2013; Yao et al., 2014)</xref>
        and convolutional nets
        <xref ref-type="bibr" rid="ref6">(Xu and Sarikaya, 2013)</xref>
        based models have been recently proposed. Other relevant
works include
        <xref ref-type="bibr" rid="ref3 ref8">(Graves et al., 2005; Graves et al., 2013)</xref>
        which proposed a
bidirectional recurrent neural network for speech recognition.
      </p>
      <p>In this paper, I make a comparison between classical statistical machine learning
method- Conditional Random Fields and deep learning method – Bi-LSTM with CNN
and CRF layers in terms of their performance on the competition dataset.
3</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>We make comparison between the performance of Conditional Random Field with
designed linguistic features and that of Bi-directional LSTM with CNN and CRF
layer on the competition dataset. CRF experiment is carried out using Python
sklearncrfsuite 0.3 package, LSTM is conducted using Tensorflow r1.2.
3.1</p>
      <p>Labeling
In order to conduct supervised learning, we label the Chinese character sequence,
There are 5 kinds of entities, which are encoded into 0 to 4. Since we label data on
character level and entities usually consist of multiple characters, we label the
beginning character of the entity B- with corresponding coded category, the rest of the
entity character I- with corresponding coded category. If a character is not part of the
entity, we label this character O. For example:”右肩左季肋部” is an entity belongs to
Body Parts category. This entity is labeled as followed: 右 B-4, 肩I-4, 左I-4, 季I-4,
肋I-4, 部I-4
3.2</p>
      <p>Conditional Random Field
Conditional Random Fields (CRFs) are undirected statistical graphical models, a
special case of which is a linear chain that corresponds to a conditionally trained
finite-state machine. Such models are well suited to sequence analysis, and CRFs in
particular have been shown to be useful in part-of-speech tagging, shallow parsing,
and named entity recognition for newswire data.</p>
      <p>Let o=o1, o2 ,...on  be an sequence of observed words of length n. Let S be a set
of states in a finite state machine, each corresponding to a label l  L , Let
s  s1, s2 ,...sn  be the sequence of states in S that correspond to the labels assigned</p>
      <p>P(s | o) 
to words in the input sequence o. Linear chain CRFs define the conditional
probability of a state sequence given an input sequence to be:
1 n
exp(</p>
      <p>m
 j1 j f j (si1, si , o, i))</p>
      <p>Z0 i1
where Z0 is a normalization factor of all state sequences, f j (si1, si , o, i) is one of m
functions that describes a feature, and  j is a learned weight for each such feature
function. This paper considers the case of CRFs that use a first order Markov
independence assumption with binary feature functions.</p>
      <p></p>
      <p>Intuitively, the learned feature weight j for each feature f j should be positive for
features that are correlated with the target label, negative for features that are
anticorrelated with the label, and near zero for relatively uninformative features. These
weights are set to maximize the conditional log likelihood of labeled sequences in a
training set D  {o, l(1) ,..., o, l(n)} :</p>
      <p>n m
LL(D)   log(P(l(i) | o(i) ))  </p>
      <p>i1 j1 2 2</p>
      <p>When the training state sequences are fully labeled and unambiguous, objective
function is convex, thus the model is guaranteed to find the optimal weight settings in
terms of LL(D). Once these settings are found, the labeling for a new, unlabeled
sequence can be done using a modified Viterbi algorithm. CRFs are presented in more
complete detail by Lafferty et al. (2001).
 2
j
3.2</p>
      <p>Bi-directional LSTM with CNN and CRF layer</p>
      <sec id="sec-2-1">
        <title>3.2.1 Convolutional Neural Network layer</title>
        <p>Convolution is widely used in sentence modeling to extract features. Generally, let l
and d be the length of sentence and word vector, respectively. Let C  dl be the
sentence matrix. A convolution operation involves a convolutional
kernel H  dw which is applied to a window of w words to produce a new feature.
For instance, a feature ci is generated from a window of words C , i : i  w by
ci   ( (C , i : i  w H )  b)</p>
        <p>Here b  is a bias term and  is a non-linear function, normally tanh or ReLu.
is the Hadamard product between two matrices. The convolutional kernel is applied
to each possible window of words in the sentence to produce a feature
map. c  [c1, c2 ,..., clw1] with c  lw1 .</p>
        <p>Next, I apply pairwise max pooling operation over the feature map to
capture the most important feature. The pooling operation can be considered
as feature selection in natural language processing.</p>
        <p>Specifically, the output of convolution, the feature map c = [c1, c2 ,..., clw1]is the
input of the pooling operation. The adjacent two features in the feature map be
calculated as follows:
p  max(ci1, ci )</p>
        <p>i</p>
        <p>The output of the max pooling operation is p  [ p1, p2 ,..., pl ] , p  l pi captures
the neighborhood information around character i within a window of specified step
size. If I apply 100 different kernels, 2 different step sizes to extract features, then p
will become 200l . Then I concatenate convolutional features to their corresponding
original character features (word2vec features) to get feature sentence matrix.</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2.2 Bi-directional Long Short Term Memory</title>
        <p>Long Short Term Memory is a special kind of Recurrent Neural Network. It can
maintain a memory based on history information using purpose-built memory cells,
which enables the model to predict the current output conditioned on long distance
features. LSTM memory cell is implemented as the following:
it   (Wxi xi  Whi ht1  Wcict1  b i )
ft   (Wxf xt  Whf ht1  Wcf ct1  bf )
ct  ft ct 1  it tanh(Wxc xt  W hc ht 1  bc )
ot   (Wxo xt  Whoht1  Wcoct  bo )
ht  ot tanh(ct )</p>
        <p>Where sigma is logistic sigmoid function and i, f, o and c are the input gate,
forget gate, output gate, and cell vectors all of which are the same size as the hidden
vector h. The weight matrix subscripts have the meaning as the name suggests. For
example, Whi is the hidden-input gate matrix, Wxo is the input-output gate matrix etc.
The weight matrices from the cell to gate vectors (e.g. Wci) are diagonal, so element
m in each gate vector only receives input from element m of the cell vector. LSTM
cell’s structure is illustrated in Figure 1.</p>
        <p>Fig. 1 - LSTM CELL</p>
        <p>
          Here we use bi-directional LSTM as proposed in
          <xref ref-type="bibr" rid="ref8">(Graves et al., 2013)</xref>
          because
both past and future input features for a given time could be accessed. In doing so, we
can efficiently make use of past features (via forward states) and future features (via
backward states) for a specific time frame. (Below dashed boxes are the LSTM cells)
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>4.1</p>
      <p>Features Template For CRF
For convenience, features are generally organized into some groups called feature
templates. For example, a bigram feature template C1 stands for the next character
occurring in the corpus after each character.</p>
      <sec id="sec-3-1">
        <title>4.1.1 External Dictionary</title>
        <p>We use dictionary such as ICD10, ICD9, and other Medicine, Pathology dictionary
(72000 terms in total)to summarize common bigram, trigram, 4gram prefix and suffix.
For example, if a bigram from sentence appears in common prefix or suffix, we make
this feature 1 otherwise 0.
4.2</p>
        <p>Hyperparameter Tuning
In Conditional Random Field, we use Elastic Nets as regularizing term and set
optimization algorithm as LBFGS and maximum iteration as 500. After random
search to tune the regularization coefficients C1, C2, I get the best C1,C2 as 0.089 and
0.004.</p>
        <p>As for deep learning models, we set the parameters as below:
We use cross-validation to evaluate F1 performance across models, here list only one
validation result for your reference.
Models Precision
Conditional Random Field(only character 92.85%
features)
Conditional Random Field(all template 93.10%
features)
Bi-directional LSTM+CRF layer 84.19%
Bi-directional LSTM with CNN, CRF layers 90.31%</p>
        <p>Using the hyperparameters listed above, we could see with basic character
features such as unigram, bigram, trigram, CRF is able to perform better than
bidirectional LSTM with CRF layers. With other augmented features, CRF’s
performance could be improved further. Convolutional Neural Network layers could
help bi-directional LSTM extract features better and improve its performance, but still
not as good as CRF.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>Conditional Random Field with augmented features performs better in my experiment
compared to Bi-directional LSTM with CNN and CRF layers in terms of F1
performance. Hence I used CRF with augmented features model for the competition.
Future works will concentrate on hyperparameter tuning for deep learning models to
get a better sense of how good the model is.</p>
    </sec>
    <sec id="sec-5">
      <title>Special Thanks</title>
      <p>We want to give special thanks to Dr. Liang Tao’s guidance during this competition.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>John D. Lafferty</surname>
          </string-name>
          ,
          <string-name>
            <surname>Andrew McCallum</surname>
          </string-name>
          , and
          <string-name>
            <surname>Fernando</surname>
            <given-names>C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Pereira</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Conditional random fields: Probabilistic models for segmenting and labeling sequence data</article-title>
          .
          <source>In ICML '01: Proceedings of the Eighteenth International Conference on Machine Learning</source>
          , pages
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          , San Francisco, CA, USA. Morgan Kaufmann Publishers Inc.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Settles</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2004</year>
          ,
          <string-name>
            <surname>August.</surname>
          </string-name>
          <article-title>Biomedical named entity recognition using conditional random fields and rich feature sets</article-title>
          .
          <source>In Proceedings of the International Joint Workshop on Natural Language Processing in Biomedicine and its Applications</source>
          (pp.
          <fpage>104</fpage>
          -
          <lpage>107</lpage>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Framewise Phoneme Classification with Bidirectional LSTM and Other Neural Network Architectures.Neural Networks</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>R.</given-names>
            <surname>Collobert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bottou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Karlen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Kuksa</surname>
          </string-name>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <year>2011</year>
          .
          <article-title>Natural Language Processing (Almost) from Scratch</article-title>
          .
          <source>Journal of Machine Learning Research(JMLR)</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P.</given-names>
            <surname>Xu</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Sarikaya</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Convolutional neural network based triangular CRF for joint intent detection and slot filling</article-title>
          .
          <source>Proceedings of ASRU.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>G.</given-names>
            <surname>Mesnil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Deng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Investigation of recurrent neuralnetwork architectures and learning methods for language understanding</article-title>
          .
          <source>Proceedings of INTERSPEECH.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>A.</given-names>
            <surname>Graves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Speech Recognition with Deep Recurrent Neural Networks</article-title>
          . arxiv.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>K. S.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. L.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zweig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. L.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Gao</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Recurrent conditional random fields for language understanding</article-title>
          .
          <source>ICASSP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Bidirectional LSTM-CRF models for sequence tagging</article-title>
          .
          <source>arXiv preprint arXiv:1508</source>
          .
          <year>01991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Combination of convolutional and recurrent neural network for sentiment analysis of short texts</article-title>
          .
          <source>In: 26th International Conference on Computational Linguistics, COLING 2016, Proceedings of the Conference: Technical Papers</source>
          , Osaka, Japan,
          <fpage>11</fpage>
          -16
          <source>December</source>
          <year>2016</year>
          , pp.
          <fpage>2428</fpage>
          -
          <lpage>2437</lpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>