<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Clinical Named Entity Recognition: ECUST in the CCKS-2017 Shared Task 2</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuhang Xia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Qi Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>East China University of Science and Technology</institution>
          ,
          <addr-line>Shanghai, China 200237</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>Clinical named entity recognition aims to identify and classify clinical terms in electronic medical records, including diseases, symptoms, treatments, exams, and body parts. Challenges occur due to ambiguity in the boundary of Chinese words and the number limitation of annotated training data. In this paper, we propose a Bi-LSTM CRF model along with self-taught learning, active learning, and ensemble learning to recognize clinical named entities. The results achieved on CCKS-2017 Task 2 dataset with a F1-Measure of 89.88% ranks among the top systems.</p>
      </abstract>
      <kwd-group>
        <kwd>clinical named entity recognition</kwd>
        <kwd>Bi-LSTM CRF</kwd>
        <kwd>self-taught learning</kwd>
        <kwd>active learning</kwd>
        <kwd>ensemble learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Electronic medical record systems have been widely used in China. Many
tasks in clinical text mining rely on accurate clinical named entity recognition
(NER), the identi cation of text spans mentioning a concept of a speci c class,
including disease, symptom, exam, treatment, and body part. Challenges occur
due to ambiguity in the boundary of Chinese words and number limitation of
annotated training data.</p>
      <p>
        Traditionally, most of the e ective NER approaches are based on machine
learning techniques, such as Support Vector Machines (SVM) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Hidden Markov
Models (HMM) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Conditional Random Fields (CRF) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], Convolutional Neural
Network (CNN) based models [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and Recurrent Neural Network (RNN) based
models [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. For biomedical NER tasks, existing e orts include rule or dictionary
based methods [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], supervised methods [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], and distant supervision methods [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In this paper, we regard the clinical NER as a sequence labeling problem,
and use the Bi-LSTM CRF model which is similar to the one presented by
Huang et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to address the problem. Di erent from Huang et al., we exploit
character embedding rather than word embedding to deal with the ambiguity
in the boundary of Chinese words. In addition, self-taught learning and active
learning is introduced to enlarge the training set. Finally, ensemble learning is
used to obtain the best recognition performance for all ve types of clinical
named entities. The results achieved on CCKS-2017 Task 2 dataset with a
F1Measure of 89.88% ranks among the top systems.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Problem Formalism</title>
      <p>The clinical named entity recognition task is de ned as a sequence labeling
problem in this paper. Given a text sequence X =&lt; x1, ..., xn &gt;, the goal is to
label X with tag sequence Y =&lt; y1, ..., yn &gt;. We experiment with three di
erent tagging formats for the recognition, including BIO (Begin, Inside, Outside),
BIOS (Begin, Inside, Outside, Single), and BIEOS (Begin, Inside, End, Outside,
Single). Examples of the three tagging formats can be found in Table 1.</p>
      <p>
        This model is similar to the one presented by Huang et. al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. It combines the
framework of bidirectional LSTM layer [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] with linear chain CRF [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Di erent
from Huang et. al., we employ character embedding rather than word embedding
to deal with the ambiguity in the boundary of Chinese words.
      </p>
      <p>
        The raw natural language input sentence is processed into sequence of
characters X = [x]1T . The character sequence is fed into an embedding layer, which
produces dense vector representation of characters. The character vectors are
then fed into a bidirectional LSTM layer. The LSTM [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] incorporates a gated
memory-cell to capture long-range dependencies within the data. In the
bidirectional LSTM, for any given sequence, the network computes both a left, !ht , and
a right, ht , representations of the sequence context at every input, xt. The nal
!
representation is created by concatenating them as ht = [ht , ht ]. The
bidirectional LSTM along with the embedding layer is the main machinery responsible
for learning a good feature representation of the data.
      </p>
      <p>Then the network use sentence level tag information via a CRF layer fed by a
fully connected hidden layer. The CRF layer is represented by lines which connect
consecutive output layers, and has a state transition matrix as parameters. With
such a layer, we can e ciently use past and future tags to predict the current tag,
which is similar to the use of past and future input features via a bidirectional
LSTM network. We consider the matrix of scores f ([x]1T ) are output by the
network. The element [f ]i,t of the matrix is the score output by the network
with parameters , for the sentence [x]1T and for the i-th tag, at the t-th character.
We introduce a transition score [A]i,j to model the transition from i-th state to
j-th for a pair of consecutive time steps. Note that this transition matrix is
position independent. We now denote the new parameters for our network as
~ = [ f[A]i,j 8i, jg. The score of a sentence [x]1T along with a path of tags [i]1T
is then given by the sum of transition scores and network scores:
S([x]1T , [i]1T , ~) =</p>
      <p>
        T
X([A][i]t 1,[i]t + [f ][i]t,t)
t=1
The dynamic programming [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] can be used e ciently to compute [A]i,j and
optimal tag sequences for inference. See [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] for details.
3.2
      </p>
      <sec id="sec-2-1">
        <title>Add Training Data by Self-taught Learning</title>
        <p>Generally, the more training data we have, the better performance we can
get. Considering that we have some unlabeled sentences, a self-taught learning
algorithm is introduced to enlarge the training set. First, we train a Bi-LSTM
CRF model using the original training set. Second, we apply the trained model
to annotate the unlabeled sentences. Third, we choose some high-quality
annotation results and add them to the original training set. Speci cally, we de ne
an annotating con dence (AC) to evaluate the quality of the annotation results:
(1)
(2)
AC([x]1T , [y]1T , ~) =</p>
        <p>eS([x]1T ,[y]1T ,~)
P eS([x]1T ,[j]1T ,~)
j
where [y]1T is the tag sequence predicted by the trained model and [j]1T is the set
of all possible output sequences. We choose those annotation results whose AC
is greater than a threshold. In addition, considering that disease and treatment
entities are much fewer than symptom and exam entities, and the body part
entity recognition is not well performed, we only choose the annotation results
which must include disease, treatment, or body part entities.
3.3</p>
      </sec>
      <sec id="sec-2-2">
        <title>Add Training Data by Active Learning</title>
        <p>We also use active learning algorithm to improve the recognition
performance. First, we use the original training data along with the self-taught
highquality annotation results to train a Bi-LSTM CRF model. Second, we apply
the trained model to annotate the rest unlabeled sentences. Third, we manually
re-label a few low-quality annotation results whose AC is less than a threshold,
and then add them to the training set. Similar to the self-taught algorithm, we
only choose the annotation results which must include disease, treatment, or
body part entities to re-label manually.
3.4</p>
      </sec>
      <sec id="sec-2-3">
        <title>Improve Recognition Performance through Ensemble Learning</title>
        <p>The task has ve types of clinical named entities to be recognized. If we
train models for each type of entities respectively, a class imbalance problem
(too many O-tags in sequences) will exist, and the models cannot utilize the
information of other types of entities. Thus, we train models to annotate all
ve types of entities at the same time, but nally choose ve models in which
each model has the best performance for the corresponding type on validation
data, so that we can obtain the best recognition performance for all ve types
of entities. For example, for disease entities, we choose the model which has the
best performance in recognizing diseases, and only use its disease recognition
results as part of the nal recognition results.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <sec id="sec-3-1">
        <title>Dataset and Evaluation Metrics</title>
        <p>We use the CCKS-2017 Task 2 dataset to perform our experiments. The
dataset contains 10,420 unannotated instances and 1,596 annotated instances
with ve types of clinical named entities, including diseases, symptoms, exams,
treatments, and body parts. The annotated instances are already partitioned
into 1,198 training instances and 398 test instances. Each instance has one or
several sentences. We further partition the training sentences, take 70% of them
as training data, and the rest 30% as validation data. We score our methods by
using the CCKS-2017 Task 2 o cial metrics, which computes F1-Measure.
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Hyper-parameters and Training Details</title>
        <p>
          As for the Bi-LSTM CRF model, we initialize character embeddings via
word2vec[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] on both the annotated data and the unannotated data. Each
character embedding is 100-dimensional. We compare the results with word
embedding segmented by Jieba Chinese segmentation module, which is also
100dimensional. We set the size of LSTM hidden layer to 64, and apply dropout [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
to the output of the Bi-LSTM layer. The dropout rate is 0.2. The Bi-LSTM CRF
model is trained by AdaDelta [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and the batch size is 128. For self-taught
learning, we add the top 3,850 automatically annotated sentences to the training set.
For active learning, we re-label the worst 125 automatically annotated sentences
manually. We split training data into sentences by periods, and set the maximum
length of a sentence to 185. If a sentence is longer than 185 characters, it will be
further split into clauses. Then we take all the labeled sentences and clauses as
a whole data set after de-duplication. We also try to split all the training data
into clauses by periods, commas and semicolons to train di erent models.
4.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Results and Discussion</title>
        <p>To explore the impact of di erent settings, we train Bi-LSTM CRF models
using di erent tagging formats on training data both in clauses and sentences, and
test them on validation data. The results are shown in Table 2. First,
characterlevel models outperform word-level models. It is because word-level approaches
may have segmentation error. What's more, the word set is much bigger than
the character set. This means the corpus is not big enough to learn word
embeddings e ectively. Second, the recognition of diseases, treatments, and body
parts is not well performed, which indicates the necessity of self-taught learning
and active learning. Third, di erent types of entities are recognized the best by
di erent settings, respectively. This indicates the necessity of ensemble learning.</p>
        <p>To dissect the e ectiveness of self-taught learning, active learning, and
ensemble learning, we train Bi-LSTM CRF models in character level only, and test
them on test data. Table 3 shows that the Bi-LSTM CRF model along with
self-taught learning, active learning, and ensemble learning achieves the best
performance in clinical NER, with a F1-Measure of 89.88%.</p>
        <p>In this paper, we propose a Bi-LSTM CRF model along with self-taught
learning, active learning, and ensemble learning to recognize clinical named
entities. We exploit character embedding to deal with the ambiguity in the boundary
of Chinese words, and employ self-taught learning and active learning to increase
training data. After comparing di erent tagging schemes, we use ensemble
learning to obtain the best recognition performance for all ve types of entities. The
results achieved on CCKS-2017 Task 2 dataset ranks among the top systems.
Acknowledgements. This work is supported by the 863 Program funded by
China Ministry of Science and Technology (Program No.2015AA020107).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Asahara</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsumoto</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Japanese named entity extraction with redundant morphological analysis</article-title>
          .
          <source>In: Proceedings of the 2003 Conference of the North American Chapter of the Association for Computational Linguistics on Human Language Technology-Volume</source>
          <volume>1</volume>
          , Association for Computational Linguistics (
          <year>2003</year>
          )
          <volume>8</volume>
          {
          <fpage>15</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bikel</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weischedel</surname>
          </string-name>
          , R.:
          <article-title>Nymble: a high-performance learning name- nder</article-title>
          .
          <source>In: Proceedings of the fth conference on Applied natural language processing</source>
          ,
          <source>Association for Computational Linguistics</source>
          (
          <year>1997</year>
          )
          <volume>194</volume>
          {
          <fpage>201</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Mccallum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Early results for named entity recognition with conditional random elds, feature induction and web-enhanced lexicons</article-title>
          .
          <source>In: Conference on Natural Language Learning at Hlt-Naacl</source>
          .
          <article-title>(</article-title>
          <year>2003</year>
          )
          <volume>188</volume>
          {
          <fpage>191</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Collobert</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karlen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kavukcuoglu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuksa</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Natural language processing (almost) from scratch</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ) (
          <year>2011</year>
          )
          <volume>2493</volume>
          {
          <fpage>2537</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bidirectional lstm-crf models for sequence tagging</article-title>
          .
          <source>Computer Science</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Kipper-Schuler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaggal</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Masanz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ogren</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>System evaluation on a named entity corpus from clinical notes</article-title>
          .
          <source>In: International Conference on Language Resources and Evaluation</source>
          <year>2008</year>
          . (
          <year>2008</year>
          )
          <volume>3007</volume>
          {
          <fpage>3011</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Supervised methods for symptom name recognition in free-text clinical records of traditional chinese medicine: An empirical study</article-title>
          .
          <source>Journal of biomedical informatics 47</source>
          (
          <year>2014</year>
          )
          <volume>91</volume>
          {
          <fpage>104</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Bing</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ling</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>W.W.</given-names>
          </string-name>
          :
          <article-title>Distant ie by bootstrapping using lists and document structure</article-title>
          .
          <source>arXiv preprint arXiv:1601.00620</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Graves</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Framewise phoneme classi cation with bidirectional lstm and other neural network architectures</article-title>
          .
          <source>Neural Networks the O cial Journal of the International Neural Network Society</source>
          <volume>18</volume>
          (
          <issue>5</issue>
          ) (
          <year>2005</year>
          )
          <volume>602</volume>
          {
          <fpage>610</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>La</surname>
            <given-names>erty</given-names>
          </string-name>
          , John,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>McCallum</surname>
          </string-name>
          , Andrew, Pereira, Fernando,
          <string-name>
            <surname>C.N.</surname>
          </string-name>
          :
          <article-title>Conditional Random Fields: Probabilistic Models for Segmenting and Labeling Sequence Data</article-title>
          . (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hochreiter</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmidhuber</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Long short-term memory</article-title>
          .
          <source>Neural computation 9(8)</source>
          (
          <year>1997</year>
          )
          <volume>1735</volume>
          {
          <fpage>1780</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Rabiner</surname>
            ,
            <given-names>L.R.:</given-names>
          </string-name>
          <article-title>A tutorial on hidden markov models and selected applications in speech recognition</article-title>
          .
          <source>Readings in Speech Recognition</source>
          <volume>77</volume>
          (
          <issue>2</issue>
          ) (
          <year>1990</year>
          )
          <volume>267</volume>
          {
          <fpage>296</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient estimation of word representations in vector space</article-title>
          .
          <source>Computer Science</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Hinton</surname>
            ,
            <given-names>G.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>R.R.</given-names>
          </string-name>
          :
          <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>
          .
          <source>Computer Science</source>
          <volume>3</volume>
          (
          <issue>4</issue>
          )
          <article-title>(2012) pags</article-title>
          .
          <volume>212</volume>
          {
          <fpage>223</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Zeiler</surname>
          </string-name>
          , M.D.:
          <article-title>Adadelta: An adaptive learning rate method</article-title>
          .
          <source>Computer Science</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>