<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Annotating a corpus of clinical text records for learning to recognize symptoms automatically</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rob Koeling</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Carroll</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Rosemary Tate</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amanda Nicholson</string-name>
          <email>A.C.Nicholson@bsms.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Brighton and Sussex Medical School</institution>
          ,
          <addr-line>Brighton</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Informatics, University of Sussex</institution>
          ,
          <addr-line>Brighton</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>43</fpage>
      <lpage>50</lpage>
      <abstract>
        <p>We report on a research effort to create a corpus of clinical free text records enriched with annotation for symptoms of a particular disease (ovarian cancer). We describe the original data, the annotation procedure and the resulting corpus. The data (approximately 192K words) was annotated by three clinicians and a procedure was devised to resolve disagreements. We are using the corpus to investigate the amount of symptom-related information in clinical records that is not coded, and to develop techniques for recognizing these symptoms automatically in unseen text.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        UK primary care databases provide a valuable source of information for research
into disease epidemiology, drug safety and adverse drug reactions. Analyses of
existing large-scale electronic patient records held in the form of lar ge primary
care datasets such as the General Practice Research Database have almost
exclusively exploited coded data. Such data are readily accessible to the classical
methods of epidemiological analysis, once the complexities of defining and
selecting a patient cohort have been overcome. However, since clinicians can choose
to what extent they code a consultation, an unknown amount of clinical data is
not coded, and ‘hidden’ in free text. Free text records often con tain important
information on the severity of symptoms or on additional symptoms which have
not been coded [
        <xref ref-type="bibr" rid="ref3 ref6">6, 3</xref>
        ]. The degree to which clinical information is coded and
how this varies between by practitioner, practice, or type of clinical problem
is currently unknown, as is the impact on public health research results of not
using information in free text. The aims of our work are to quantify how much
additional information is in the free text and to explore methods for extracting
it.
      </p>
      <p>Automatic extraction of complex information from notes written by general
practitioners, which may be ungrammatical and often contains ambiguous terms,
misspellings and abbreviations, is a very challenging natural language processing
(NLP) task. The text is much less uniform than data typically analysed by the
NLP research community, and issues of confidentiality make it difficult to gain
access to significant amounts of data.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Data</title>
      <p>
        This study builds on previous work [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], in which we used coded records from
the General Practice Research database (GPRD [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]) of 344 patients between 40
and 80 years of age (inclusive) diagnosed with ovarian cancer between 1 June
2002 and 31 May 2007. The records use the Read coding system, which was
originally developed in the 1980s and is used throughout the United Kingdom
for coding clinical events in primary care. Each Read code has an associated
textual description e.g. ‘Abdominal pain’, ‘Right iliac fossa pain’, ‘Const ipation’,
which are available on GP systems as an aid for recording the correct code. In
the current study we obtained manually anonymized free text records of all 344
patients for the period 12 months prior to the date of definite diagnosis.
      </p>
      <p>The free text records contain information from a variety of different sources.
Mostly they consist of notes typed by the GP during or after a consultation,
communication with secondary care (for example referral letters and discharge
summaries), and sometimes test results. However, about 90% of the records
contain notes typed by the GP, which turn out to be the most challenging category.
A typical example is:
5 day Hx of umbilical Dx. Smelly Dx. Red inflammed lump within
umbilicus. Swab sent. Try fluclox and review. Also ?mass felt left lower abdo.</p>
      <p>No weight loss, bowels reg. No Jaccol.</p>
      <p>Some differences between this data and standard English are: (1) inconsistent
use of capitalization and punctuation; (2) spelling errors and unusual
abbreviations, acronyms, and named entities; (3) anomalous tokenization (e.g. missing
spaces); and (4) ambiguous use of question marks. These characteristics make
it difficult to process the data automatically. They also impact on readability,
making human annotation more time-consuming, and therefore cos tly.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The Annotation Process</title>
      <p>
        In order to annotate symptoms associated with ovarian cancer in the free text
fields of patient records, we first identified the most commonly experienced
ovarian cancer symptoms. Table 1 lists these symptoms, which were taken from a
recent paper by Hamilton [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        To facilitate the work of the annotators, we created an easy to use
interactive annotation system. We used the Visual Tagging Tool (VTT), part of the
SPECIALIST NLP Tools [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. VTT allowed us to create an environment in which
an annotator can highlight a phrase in the text and choose, from a pull down
menu, the most appropriate tag to describe the symptom they had highlighted.
A screenshot of the annotation workbench is shown in Figure 1.
      </p>
      <p>
        We drafted a detailed set of annotation guidelines. Over several iterations we
refined the guidelines to minimize potential disagreement between the annotators
— learning from others’ experiences in defining a methodology for an notating
clinical data [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The guidelines ask annotators to:
identify all the different expressions (words or strings of words) in the notes
that represent a symptom from the list of pre-defined relevan t symptoms for
ovarian cancer. These expressions can be complaints expressed by patient,
signs detected on examination or by investigation or findings at operation. All
of these should be marked as long as they refer to one of the symptoms in the
pre-defined list.
      </p>
      <p>The main supporting instructions are:
– Do not infer a symptom from the text : only annotate symptoms if found as
such in the text.
– Presence or absence : e.g. in the case of T‘here was no evidence of abdominal
distension,’ a‘bdominal distension’ should be marked up.
– Annotate the bare minimum : only annotate as much text as you need to
identify the symptom. E.g. .‘..who has had right upper quadr ant abdominal pain
now for some weeks,’ only a‘bdominal pain’ should be marked u p.</p>
      <p>– Annotate every occurrence of a symptom in a record .</p>
      <p>Five individuals with a medical background were involved in developing the
guidelines, iteratively refining them on the basic of exploratory annotation
sessions. Once the final version of the guidelines was produced, three people each
annotated the whole dataset independently. One of the annotators was also
involved with developing the guidelines. All the annotators have a medical
background, either as a General Practitioner, a researcher with GP training, or a
final-year medical student. The annotators were given a copy of t he annotation
software and worked with it at their convenience. The data was presented to
them in batches of about 800 records each (but there was no requirement to
finish a batch in a single session).
4</p>
    </sec>
    <sec id="sec-4">
      <title>The Corpus</title>
      <p>We started with 6141 records, determined by the amount of data available for
the 344 patients for the period 12 months prior to the date of definite diagnosis
(Section 2). After these records were annotated by the three annotators, we
created a gold standard corpus. During the development of the guidelines and
exploratory annotation sessions we noticed that there were two main areas of
disagreement between annotators. Firstly, annotation is a mentally strenuous
task, and as a result annotators occasionally miss symptoms, especially when
a record contains a large number of them. Secondly, disagreements arise from
the fact that annotators are free to decide the points at which each marked-up
symptom starts and ends.</p>
      <p>
        Inter-annotator agreement [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is an important quality measure. T he standard
metric for inter-annotator agreement for categorization tasks with two
annotators is the kappa statistic, defined as k=P (a) − P (e)/1 − P (e), where P (a) is the
measured probability of agreement between annotators, and P (e) is the
probability that agreement is due to chance. However, a complicating factor for our
annotation task is that categorization is only one element of the task. The other
element is the choice of the boundaries of the expression that is associated with
a certain class. The set of possible expressions to annotate is extremely large,
which makes it impossible to estimate P (e). This issue arises whenever the set
of elements that has to be marked up is not fixed. This has also been noted in
annotation efforts that included similar tasks, such as [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        In our setup it is difficult to calculate an exact figure for agreement other
than the most basic one. The most straightforward measure is strict agreement,
defined as the proportion of all cases where all three annotators were in full
agreement (those cases where the annotated string starts and ends at exactly
the same position in the text and the string is assigned the same label). We
found three-way strict agreement in about 62% of cases. This figu re is difficult
to compare with other work reported in the literature. The closest comparison we
are aware of is the inter-annotator agreement for finding ‘Signs an d Symptoms’
in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. They report an F-measure (a combination of precision and reca ll) of
0.61 for double annotated text. However, on the one hand their task was more
complex (the text was annotated for several aspects at the same time), but on
the other hand the text we annotated is less like standard English, and we had it
triple annotated. Considering these factors we were pleasantly surprised by the
annotation agreement figure, especially since inspection of the remaining cases
suggested that many could be resolved without much effort.
      </p>
      <p>
        Both [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] propose second, less strict measure of annotation accuracy
that allows for partial matching of annotated strings. This is well-de fined for
double annotation, but difficult to adapt to triple annotated data. We therefore
decided to stick to the strict agreement measure for our data.
      </p>
      <p>
        In order to maximize the quality of the gold standard, we had to decide
on a method for combining the data individually created by the annotators. A
typical way of resolving disagreement between annotators is to have the data
double annotated and appoint a third annotator to choose between the two
conflicting opinions. However, this method does not allow for discussing cases
that are inherently difficult. We therefore decided to keep the triple annotation,
and employ a variant of the Delphi method [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to generate consensus. In the
Delphi method each annotator works individually, gets feedback on the cases
where there is disagreement, and is asked to revisit those cases. We produced an
overview of the disagreements and asked the three annotators to come back to
discuss and resolve them.
      </p>
      <p>In order to use the annotators’ time efficiently, we created five ca tegories
from the cases of disagreement, each of which could be approached differently.
The categories were:
1. Trivial difference: e.g. A‘bdominal pain’ vs. A‘bdominal pai’
2. Added modifier: e.g. s‘lightly bloated abdomen’ vs. b‘loated abdomen’
3. 2 agree/1 missed out: one annotator may have overlooked a symptom
4. 2 agree/1 disagree: with respect to the span of the string or associated label
5. Rest: all other instances
The first two categories, accounting for around 20% of the disagreements, were
easy to resolve. The third and fourth categories had to be checked one by one
by the annotators, but were mostly easily resolved. The Rest category was the
most challenging. Many cases in this category had multiple disagreements (over
both the span of the string and the associated tag) and often the annotators had
to view the full context of the annotated string to reach agreement.</p>
      <p>The resulting annotated corpus consists of a total of 6141 records, containing
about 192K words. The total number of annotated symptoms is 3955. Table 2
summarizes the symptoms labeled in the corpus. Even though the average
number of occurrences of each distinct expression is low, the distribution is very
skewed (see Figure 2). For example, ‘abdominal swelling’ is an expres sion that
the annotators label as ‘Abdominal Distension’; there are ten occu rrences
(‘tokens’) of this expression (‘type’) in the corpus. However, there is g reat variation
in the expressions used to describe the same symptom. The more formal ones are
used frequently, whereas informal expressions (e.g. ‘tummy is str ikingly larger’
or ‘abdo looks sl swollen’) often occur only once.</p>
      <p>In one experiment we looked at a subset of symptoms, covering about 3281
annotated expressions. Approximately 1200 of these occur only once in the
corpus. Figure 2 shows that the most frequent 100 types account for almost 65%
of the tokens. This observation is important for designing an automatic system
that finds this information in the text. If recall is not an absolute priority, then
it is possible to get a long way by concentrating on the high frequency types.
Moreover, the most frequently occurring expressions are generally more formal
and concise, which would also be of assistance. However, a system would also
need to integrate relevant contextual factors, in particular whether indications
of symptoms are negated or are attributed to someone other than the patient.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>The corpus we have produced is a rich resource, which we are using for two
purposes. Firstly, we are investigating the amount of symptom-relate d information
available in the free text fields of primary care patient records. By comparing
symptom annotations with coded information in the same records we are
exploring the hypothesis that a significant amount of information is missed when the
contents of free text fields are not taken account of in epidemiological research.
Secondly, we are starting to use the data for creating and evaluating techniques
for automatic recognition of symptoms in free text. Although we are ultimately
interested in developing machine learning-based models that precise ly capture
as wide a variety of symptoms as possible, the annotated corpus also allows us
to estimate the utility of unsophisticated techniques such as approximate string
matching and thesaurus-based expansion. If significant amounts of information
can be uncovered with such methods, then the epidemiological research
community would not require specialised natural language processing expertise to
be able to exploit free text resources. Automatic processing of information in
free text fields also opens up opportunities to work with un-anonym ized data.
Restricted access to textual data is a major hurdle in research using electronic
patient records. The possibility of retrieving information without the need for
manual anonymization would open up many new opportunities.</p>
      <p>The process of establishing consensus between annotators has given us new
insights into the nature of the data and has highlighted issues that need to be
addressed when processing this data and when defining further annotation tasks.
One of the main issues relates to symptoms recorded as a result of a complaint by
the patient versus the outcome of an examination or a test result. From a clinical
point of view, a complaint has a different status to an examination result. Even
though many potential disagreements of this type were resolved in the course of
developing the annotation guidelines, new issues will crop up when annotating
large amounts of text. Decisions about how to deal with these cases might need
to be informed by the research question being addressed.</p>
      <p>There is evidence that important information is missed when epidemiological
research using patient records relies only on coded information. The research
reported here is a step towards quantifying how much additional information is
potentially available, and is a prerequisite for research into automatic retrieval
of this information and making it available to the research community.
Acknowledgments This work was supported by the Wellcome Trust [086105]
(The ergonomics of electronic patient records: an interdisciplinary development
of methodologies for understanding and exploiting free text to enhance the utility
of primary care electronic patient records (“PREP”) ). The funder had no role
in study design, data collection and analysis, decision to publish, or preparation
of the manuscript.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. GPRD: GPRD.
          <article-title>Excellence in public health research</article-title>
          . http://www.gprd.
          <source>com)</source>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Group, L.S.: http://lexsrv3.nlm.nih.gov/LexSysGroup/Home/ (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>T.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bankhead</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharp</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Risk of ovarian cancer in women with symptoms in primary care: population based case-control study</article-title>
          .
          <source>British Medical J. 339 (AUG 25</source>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hripcsak</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rothschild</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          :
          <article-title>Agreement, the f-measur e, and reliability in information retrieval</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>12</volume>
          ,
          <fpage>2962</fpage>
          -
          <lpage>98</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Hripcsak</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilcox</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Reference standards, judges, and comparison subjects</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>9</volume>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Johansen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scholl</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hasvold</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ellingsen</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bellika</surname>
          </string-name>
          , J.G.: G“arbage In, Garbage Out”
          <article-title>- Extracting Disease Surveillance Data fro m EPR Systems in Primary Care</article-title>
          . pp.
          <fpage>5255</fpage>
          -
          <lpage>34</lpage>
          . ACM;
          <string-name>
            <surname>ACM</surname>
            <given-names>SIGCHI</given-names>
          </string-name>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2008</year>
          ), ACM C onference on Computer Supported Cooperative Work, San Diego, CA, NOV
          <volume>08</volume>
          -
          <issue>12</issue>
          ,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Pyysalo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ginter</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heimonen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bjrne</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boberg</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jrvinen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salakoski</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Bioinfer: a corpus for information extraction in the biomedical domain</article-title>
          .
          <source>BMC BioInformatics 8</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaizauskas</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hepple</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demetriou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guo</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roberts</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Setzer</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Building a semantically annotated corpus of clinical texts</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>42</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>South</surname>
            ,
            <given-names>B.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garvin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Samore</surname>
            ,
            <given-names>M.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gundlapalli</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          :
          <article-title>Developing a manually annotated clinical document corpus to identify phenotypic information for inflammatory bowel disease</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>10</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Tate</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>A.G.R.</given-names>
          </string-name>
          , Murray-Thomas,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Anderso n</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.R.</given-names>
            ,
            <surname>Cassell</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.A.</surname>
          </string-name>
          :
          <article-title>Determining the date of diagnosis - is it a simple matter? The impact of different approaches to dating diagnosis on estimates of delayed care for ovarian cancer in UK primary care</article-title>
          .
          <source>BMC Medical Research Methodology</source>
          <volume>9</volume>
          (
          <year>June 2009</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>