<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Performance of a multi-class biomedical tagger on clinical records</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>S. V. Ramanan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shereen Broido</string-name>
          <email>shereen.broido@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P. Senthil Nathan</string-name>
          <email>senthil@npjoint.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>RelAgent Pvt Ltd</institution>
          ,
          <addr-line>Chennai</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vidhodaya Schools</institution>
          ,
          <addr-line>Chennai</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We tested the performance of Cocoa, an existing dictionary/rule based entity tagger that tags multiple semantic types in biomedical domain including diseases, on disease/sign/symptom detection in clinical records in the ShARe/CLEF eHealth task. Initial analysis showed that the precision was high ( 90%), but recall was low ( 50%) due to (a) phrases peculiar to clinical notes (b) disambiguation of common words and (c) the large number of unde ned acronyms. We extended the system to handle these cases by reference to the local intrasentential context as derived from the training set. A small module was also added for event-based detection of annotated sentence fragments containing verbs/gerunds; an example is `LV systolic function appears depressed'. The event detection system had about 30 rules. With these modi cations, the f-score was 0.75 on the test set. In a second run, we added about 70 frequently occurring acronyms as well 15 phrases which were all in caps. The nal results on the test set (f = 0:78) show that a multi-class tagger can work reasonably well on clinical records.</p>
      </abstract>
      <kwd-group>
        <kwd>rule-based tagger</kwd>
        <kwd>multiple entity types</kwd>
        <kwd>clinical notes</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Automatically tagging and normalizing mentions of diseases, signs and
symptoms in clinical records is a useful addition even when these records have already
been manually assigned ICD codes. Automatically assigned tags may help
uncover unexpected correlations between symptoms [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], and may also be useful in
checking the accuracy of the manual annotations.
      </p>
      <p>
        Previous shared tasks in the clinical domain have addressed subsets of records
that are typically seen by a medical practitioner, such as radiology reports [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
and discharge summaries [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The current ShARe/CLEF eHealth task covers
annotation of diseases, signs and symptoms in a mixed bag of documents, including
discharge summaries and echo, radiology and ECG reports [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However, the task
does not cover GP notes, a very challenging category [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Cocoa [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is an existing named entity tagger for published literature in the
biomedical domain. Cocoa tags entities across a variety of semantic classes,
including chemicals, proteins, cellular parts, anatomical parts and diseases. We
wished to explore how well such a system would perform on clinical notes, which
is a domain slightly di erent in scope and context from published biomedical
literature. For example, common terms and phrases, such as `mass' and `e usion',
refer exclusively to signs/symptoms when used in clinical notes. We explored
whether sentence-level disambiguation is su cient to resolve such ambiguous
phrases. Additionally, acronyms are well-understood in the clinical context, and
therefore used without an associated expansion in discharge summaries for
example. With a small subset of pre-de ned long acronyms, context-sensitive short
acronyms, and with sentence-level disambiguation, the system gave a precision
of 0:90 and a recall of 0:69 for a f-score of 0:78. While much less than the
topranked score (f = 0:87), the results show that multi-class recognition systems
can produce reasonable performance in the clinical domain.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>System pipeline</title>
      <p>A schematic of the system is given in Figure 1. As mentioned above, the system
tags entities across a number of semantic classes. We restrict our discussion to
entities relevant to the present task, namely diseases, signs and symptoms, along
with the anatomical parts that they a ect.</p>
      <p>Sentence</p>
      <p>Splitter
Coordination
module
Event
Detector</p>
      <p>Acronym
Detector
Multi-word</p>
      <p>Entity
Detector</p>
      <p>POS tagger
/Chunker</p>
      <p>Entity
tagger
Fig. 1. Block-level pipeline of modules in the Cocoa NER system
(a) Sentence splitter. In the biomedical literature, sentence boundaries present
a challenge as sentences can span multiple lines. Further, sentences can begin
with lower case letters (e.g. `cAMP') and numbers, as many biological entities
have complex, but widely recognized, orthography. However, we observed in the
CLEF training set that, both in clinical notes and in lab reports, sentences often
did not have a trailing period (`full stop'). Moreover, sentence fragments were
often used to describe mental states or conditions for example. We therefore used
newlines to mark sentence boundaries.</p>
      <p>(b) Acronym detector. The system detects acronyms through a dynamic
programming methodology. Even though we did not note any acronym de nitions
in the training set, this module was not disabled.</p>
      <p>(c) POS tagging and chunking. These were done by TBL methods, with a
Brill POS tagger followed by a fnTBL-based chunker. Both are already heavily
modi ed for the biomedical domain in the existing system, and we did not make
any substantial changes for this task.</p>
      <p>(d) Entity tagging. Both anatomical parts and diseases were tagged at a
word level with the help of word-level dictionaries (e.g., `Parkinsonism') and
dictionaries of morphological pre xes and su xes (e.g. `cephalon' for anatomy
and `oglossia' for diseases). False positive dictionaries were maintained for
entities detected by morpheme-based methods. Tags were used in a limited way to
correct chunking errors, primarily in VP chunks.</p>
      <p>(e) Multi-word entities. Adjectives such as `aberrant', `ruptured' and
`enlarged' followed by an anatomical part are tagged as a symptom. We also tagged
other adjectives connected with time (e.g. `postictal') preceding diseases,
disease postpositions (`progressiva'), as well as a host of domain-dependent word
combinations (`prominent ear', `wasting disease'). These multi-word
combinations were derived from an exhaustive analysis of UMLS and ICL de nitions
of diseases, signs and symptoms. Multi-word combinations involving
anatomical parts were derived from a number of sources, including Gray's Anatomy.
For the CLEF task, we extended this module to disambiguate common words
as signs/symptoms with appropriate context (`negative drift', `negative masses',
`bilateral e usion', `adventitious movements'). A narrow context/trigger based
tagging was also added for certain acronyms (`negative for DVT', `moderate
MR', `depressed LVEF', `without r/w/w'). A few acronyms were also marked
up for the second run of the system against the test set when they were long
and seemed to have no other association in the biomedical literature (`ARDS',
`NTND') or were extensively used in the training data (`MR', `TR', `AS').
Entity tagging is case sensitive, thus we marked up certain phrases which were all
in capital letters and were commonly observed in the training set (`ARTERY
DISEASE', `PNEUMONIA'). Markup of a few acronyms without a surrounding
context, and markup of a few all-caps phrases, constitute the only di erence
between 1st and 2nd runs of the the system.</p>
      <p>(f) Coordination module. This modules marks up noun phrases that are
in coordination. This can occur through placement of commas and functional
words (`and', `or'), or through compatible tags in head entities in putative
coordinated phrases. Anatomical entities followed by disease tags are united as a
single disease/symptom entity. Further anatomical parts in coordinated phrases
and followed by a disease tag are also marked up as diseases (`breast, ovarian and
prostate cancer'). Certain disease pre xes are also merged at this stage (`acute',
`lethal'), and bodypart-disease coordination is repeated to detect phrases such
as `ovarian and early-onset breast cancer'. Certain organism-disease
combinations are also detected here (`viral infection'). We did not make any substantial
changes in this module for the CLEF task.</p>
      <p>(g) An experimental event detector for the clinical domain. Verbs, gerunds
and nominals de ne `trigger words' which take NP's as arguments, and de ne
some of the extended annotations in this task. Examples are `LV systolic function
appears depressed' and `ascending aorta is moderately dilated'. Such extended
annotations seemed primarily to correspond to signs and symptoms of disease.
We wrote a small module to detect some of these trigger-based sign/symptom
events in the testing set. Altogether, about 30 rules were added to detect such
events for this task.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        We rst tested the performance of the system as available online [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] against
the development sets, and considered only entities marked up as `Disease' or
`Diseased bodypart'. In the relaxed evaluation mode, precision was 0:91, while
recall was 0:51. Accordingly, we modi ed the system as described in the section
above to better capture additional entities in the clinical domain, but without
a ecting performance in the various other semantic classes detected by the Cocoa
tagger.
      </p>
      <p>The major changes that were e ected are (a) disambiguation of common
words (`mass') when they resolve to signs/symptoms in a clinical document
and (b) resolution of acronyms that occur commonly in clinical records
without an associated expansion. Disambiguation of common words and resolution
of acronyms were done in a intrasentential context-sensitive manner based on
manual examination of the training set data and appropriate framing of the rules.
A small event detection module with about 30 rules was added to detect sentence
chunks which corresponded primarily to signs and symptoms (`hematocrit had
not increased'). On the test set, this approach (Run marked `TeamRelAgent.1')
yielded a precision of 0:91 and a recall of 0:64, with a f-measure of 0:75.</p>
      <p>A number of common short acronyms for diseases/symptoms remained
undetected by this approach in the training set. Moreover, there were words and
phrases marked all in capital letters in the text that were also left untagged by
the system, which is case-sensitive. We added about 70 acronyms that occurred
frequently in the training corpus (`AS' `MVP', `PVD') as well as 15 all-caps
phrases and tried a second run (`TeamRelAgent.2') on the test set. The
precision lowered by 0:01 to 0:90, but the recall increased far more, from 0:64 to 0:69,
with a f-score of 0:78.</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>
        We re ned an existing multi-class entity tagger for the biomedical domain
(Cocoa) against the test set. The existing tagger already has reasonable performance
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] against a UMLS-based disease corpus (the Arizona disease corpus, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). The
challenge in extending the system to clinical records was in keeping the precision
high while increasing the recall, and yet not compromise performance against
other entity classes. With fairly minor improvements, the system achieved a
f-score of 0:78 against the test set as compared to the best score of 0:87.
      </p>
      <p>We did not address the problem of increasing precision during this task, apart
from addressing obvious errors, such as demarking `Allergies' when it is a section
heading or a department name. Recall was increased by manually analyzing the
test for untagged or wrongly tagged entities. Many of these arose from mistagging
of common words, such as `mass' and `drift', which are symptoms in the clinical
context. Acronyms are used frequently in discharge summaries and lab reports
without any expansion, and are another source of low recall. Even taking these
into account, we could achieve a recall of 0:69 at best in the relaxed evaluation.
By comparison, the best-performing system had a recall of 0:83.</p>
      <p>Low recall arises for a number of reasons. We did not mark up words such
as `agitated', `lethargic', `uncooperative' and `mass' without an intra-sentential
context for disambiguation. We also did not mark up sentence fragments such as
`temperature decreased', as they may not refer to symptoms in other contexts,
such as in biochemistry. Rarer acronyms also remain untagged as diseases, as
they may refer to chemical or protein names in a general biological context.
Even with these constraints, we were able to get a reasonable recall of 0:79 in
the training set. However, recall dropped to 0:69 in the test set, for reasons that
we have not yet analysed. However, given the small number of modi cations that
we made to the existing system to increase recall to reasonable gures, we felt
that the system is capable of better performance with added e ort.</p>
      <p>In summary, we have shown that a multi-class entity detection system is
capable of achieving reasonable performance in the clinical domain without
compromising performance in other classes (data not shown). Clinical documents
often contain chemical and protein names and associated quantitative values (e.g.
dosage, serum concentrations). A multi-class NER system may thus be useful
in correlating multiple entity classes as well as quantitative information with
disease occurrence in clinical records. Such correlations would be of relevance in
hypothesis-based discovery, such as in cohort analysis, but also in hypothesis-free
analysis of large datasets.</p>
      <p>Acknowledgments. We thank Dr. K. E. Ravikumar for suggestions and
discussions. The ShARe/CLEF eHealth shared task was made possible by NIH grant
R01GM090187 to the task organizers.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Koeling</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carroll</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tate</surname>
            <given-names>A. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nicholson</surname>
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Annotating a corpus of clinical text records for learning to recognize symptoms automatically</article-title>
          .
          <source>2011. Proceedings of Louhi '11.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Leaman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>G:</given-names>
          </string-name>
          <article-title>Enabling Recognition of Diseases in Biomedical Text with Machine Learning:</article-title>
          <source>Corpus and Benchmark</source>
          .
          <source>2009. Symposium on Languages in Biology and Medicine</source>
          ,
          <volume>82</volume>
          -
          <fpage>89</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Pestian</surname>
            <given-names>J. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brew</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matykiewicz</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovermale</surname>
            <given-names>D. J.</given-names>
          </string-name>
          , Johnson N.,
          <string-name>
            <surname>Cohen K. B.: A Shared Task</surname>
          </string-name>
          <article-title>Involving Multi-label Classi cation of Clinical Free Text</article-title>
          .
          <year>2007</year>
          .
          <article-title>Association for Computational Linguistics (ACL</article-title>
          ),
          <year>2007</year>
          :
          <fpage>97104</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>4. RelAgent Pvt Ltd.: Cocoa. http://npjoint.com/CocoaEval.html</mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Roque</surname>
            ,
            <given-names>F. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jensen</surname>
            <given-names>P. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmock</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dalgaard</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Andreatta</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soeby</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bredkjor</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juul</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Werge</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jensen</surname>
            <given-names>L. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brunak</surname>
            ,
            <given-names>S:</given-names>
          </string-name>
          <article-title>Using electronic patient records to discover disease correlations and stratify patient cohorts</article-title>
          .
          <year>2011</year>
          . PLoS Comp. Bio.
          <volume>7</volume>
          (
          <issue>8</issue>
          ):
          <fpage>e1002141</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Suominen</surname>
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salantera</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai S</surname>
          </string-name>
          . et al.:
          <source>Three Shared Tasks on Clinical Natural Language Processing. Proceedings of CLEF 2013</source>
          . To appear.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Uzuner</surname>
            <given-names>O.</given-names>
          </string-name>
          :
          <year>2011</year>
          i2b2/
          <article-title>VA co-reference annotation guidelines for the clinical domain</article-title>
          . Available from: https://www.i2b2.org/NLP/Coreference/assets/CoreferenceGuidelines.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>