<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deriving an English Biomedical Silver Standard Corpus for CLEF-ER</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ian Lewin</string-name>
          <email>ian.lewin@linguamatics.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon Clematide</string-name>
          <email>simon.clematide@uzh.ch</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Linguamatics Ltd</institution>
          ,
          <addr-line>324 Science Park, Milton Road, Cambridge CB4 0WG</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Zurich</institution>
          ,
          <addr-line>Binzmuhlestr. 14, 8050 Zurich</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We describe the automatic harmonization method used for building the English Silver Standard annotation supplied as a data source for the multilingual CLEF-ER named entity recognition challenge. The use of an automatic Silver Standard is designed to remove the need for a costly and time-consuming expert annotation. The nal voting threshold of 3 for the harmonization of 6 di erent annotations from the project partners kept 45% of all available concept centroids. On average, 19% (SD 14%) of the original annotations are removed. 97.8% of the partner annotations that go into the Silver Standard Corpus have exactly the same boundaries as their harmonized representations.</p>
      </abstract>
      <kwd-group>
        <kwd>annotation</kwd>
        <kwd>silver standard</kwd>
        <kwd>challenge preparation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The English markup supplied for the CLEF-ER challenge was generated
automatically using a Silver Standard methodology in order to harmonize six
di erent annotations from the project partners. This paper explains the process
used to generate the Silver Standard and the issues raised during its construction.
In the following sections, we rst outline the task requirements for the
CLEFER English silver standard annotation. Next, we discuss how the Silver Standard
methodology from the predecessor project \Collaborative Annotation of a Large
Scale Biomedical Corpus" (CALBC [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]) was adapted to the new scenario. Finally
we present some results and our conclusions.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>CLEF-ER requirements</title>
      <p>The CLEF-ER task scenario imposes a number of requirements which makes
the collection of manual and/or Gold Standard annotations (i.e. veri ed by a
reliable procedure incorporating expert opinion) particularly onerous.
1. The concepts to be annotated are
(a) very large in number
(b) highly specialist
(c) highly diverse in nature, even though they all t under the broad heading
of `biomedical'. It is unlikely any one individual will be an expert in all
of the included concepts.
2. The document set to be marked up is
(a) very large in size
(b) specialist
(c) highly diverse, ranging from extracts of scienti c papers, to drug labels
to claims in patent documents.
3. The task requires assignment of concept identi ers (sometimes called
\normalization" or \grounding") and not just the identi cation of names in text,
perhaps with an indication of their semantic type.</p>
      <p>The concepts are taken from UMLS and are (all) the members of the
following speci ed semantic groups: anatomy, chemicals, drugs, devices, disorders,
geographic areas, living beings, objects, phenomena and physiology. The sources
of the concepts are MeSH (Medical Subject Headings), MedDRA (Medical
Dictionary for Regulatory Activities) and SNOMED-CT (Systemized
Nomenclature of Human and Veterinary Medicine). There are 531,466 concepts in total.
The concept identi ers to be assigned are UMLS Concept Unique Identi ers (or
CUIs).</p>
      <p>The document set includes nearly 1.6m sentences (15.7m words) from
scienti c articles (source: titles of Medline abstracts)6, 364k sentences (5m words)
from drug label documentaion (source: European Medicines Agency)7 and 120k
claims (6m words) from patent documents (source: IFI claims)8.
6 http://www.ncbi.nlm.nih.gov
7 http://www.ema.europa.eu/ema
8 http://ificlaims.com
endogenous TGF-beta-specific complex</p>
      <p>TGF-beta
endogenous TGF-beta</p>
      <p>TGF-beta-specific complex
. . . t h e e n d o g e n o u s T G F b e t a s p e c i f i c c o . . .
. . . 0 0 0 2 2 2 2 2 2 2 2 2 2 4 4 4 4 4 4 2 2 2 2 2 2 2 2 2 2 . . .
One additional requirement which clearly distinguishes the CLEF-ER entity
recognition task from other similar challenges is that the recognition is to be
performed in a di erent language from that of the supplied marked up data.
This has the advantage, for current purposes, that the precise boundaries of the
entities supplied in the English source data are not of critical importance. No
challenge submissions will be evaluated against them. Nevertheless, the
assignment of reasonable boundaries still remains important to the challenge not least
as a possible locus for the use of machine translation technology in nding
correlates in other languages. Furthermore, it is important that good precision is
obtained in the assignment of CUIs.</p>
      <p>These requirements all clearly speak to the use of an automatically derived
silver standard for the English annotation. Automatic annotation enables one to
cover large amounts of data and, by suitably combining the results of several
systems, to avoid precision errors introduced by the inevitable idiosyncracies of any
one system. Consequently, we decided to adapt and re-use the centroid
harmonization methodology, deployed in a previous large-scale annotation challenge:
the CALBC challenge (Collaborative Annotation of a Large Scale Biomedical
Corpus).
3</p>
    </sec>
    <sec id="sec-3">
      <title>The centroid harmonization of alternative annotations</title>
      <p>Figure 1 shows four human expert (BioCreative Gold Standard) annotations
over the string endogenous TGF-beta-speci c complex.</p>
      <p>The centroid annotation algorithm, developed initially for the CALBC project,
reads in markup from several di erent annotators and generates a single
harmonized annotation which represents both the common heart of the set of
annotations and its boundary distribution. The inputs can be Gold Standard as in
Figure 1 or imperfect automatic annotations.</p>
      <p>First, text is tokenized at the character level and (ignoring spaces) votes are
counted over pairs of adjacent inter-entity characters in the mark-up. Figure
1 also shows the inter-entity character counts. For example, only two of the
markups consider that the transition e-n at the start of endogenous falls within
an entity. All four consider that T -G do. The focus on inter-entity pairs, rather
than single characters, simply ensures that boundaries are valued when two
Visceral &lt;e b='l:0:2,l:9:3,r:0:5'&gt; adipose tissue&lt;/e&gt;is particularly responsive to
somatropin
di erent entity names happen to be immediately adjacent to each other in a
markup. Over the course of a whole text, the number of votes will mostly be
zero, punctuated by occasional bursts of wave-like variation. The centroids are
the substrings over character pairs that are peaks (or local maxima) in a burst
of votes. In Figure 1, TGF-beta is the centroid.</p>
      <p>The inter-entity character votes also de ne the boundary distribution around
the centroid. We de ne a boundary whenever the number of votes changes. Its
value is the di erence in votes. So, the centroid TGF-beta has a possible left
boundary before T and receives a value of 2 (the di erence between 4 and 2)
and another before endogenous, which also receives 2 (the di erence between 0
and 2). Therefore, these alternative boundaries are equally preferred. There is
however no boundary immediately after speci c.</p>
      <p>
        With respect to Gold Standard inputs, the advantages of the centroid
representation lie rst in its perspicuous representation of variability and secondly
in its use as something which candidate annotations can be evaluated against
(see [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for details) with a clear semantics and scoring method in which equally
preferred alternatives are equally scored and more preferred alternatives score
more highly than less preferred ones.
      </p>
      <p>When the inputs are less than perfect (the result of automatic annotators
rather than human experts), the result is a harmonized Silver Standard. Figure
2 shows one centroid uncovered by applying the algorithm to the annotations of
CLEF-ER data generated by the ve Mantra project partners. In this case, the
centroid itself is simply the most highly voted common substring adipose tissue
but the boundary distribution shows the alternative boundaries.</p>
      <p>`l:0:2" represents the left boundary at exactly the centroid left boundary
(position 0) for which two votes were cast. \l:9:3" represents the boundary 9
characters to the left with three votes. \r:0:5" shows that all ve votes agreed
that the centroid's right boundary is exactly (position 0) the right boundary of
the entity.</p>
      <p>
        One advantage of a centroid Silver Standard is that one easily tailors it
for precision or recall (for example, by throwing away centroids or boundaries
with very few votes). It should be noted that the centroid heart, adipose
tissue in no way represents the correct annotation, or even the best annotation.
In the evaluation scheme of [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] it is simply the string which a candidate
annotation must at least cover in order to be true-positive; and a candidate which
also included Visceral would in fact score proportionately more because more
standards-contributing annotators agree that it should.
      </p>
    </sec>
    <sec id="sec-4">
      <title>Adapting centroids for CLEF-ER</title>
      <p>Although we considered giving CLEF-ER participants the maximally
informative centroid representation, it was at least equally desirable to provide the
simplest possible representation. We did not wish to discourage participation.
There was also no guarantee that the distributional information could be
usefully exploited by participants, nor were we ourselves planning to exploit it in
evaluation. Therefore, for clarity and simplicity, we decided to o er individual
entity markup in a classical format, rather than a distributional format.</p>
      <p>First, we apply a threshold to centroids so that only centroids with at least
three votes percolate through to the Silver Standard. Higher thresholds would
generate a higher precision harmonization but experiment showed that a three
vote threshold gave good recall without sacri cing too much in precision.</p>
      <p>Secondly, in the context of CLEF-ER, we decided that, where harmonization
o ered a distribution over boundaries, it would be more useful to participants
to o er the widest possible boundary, again subject to a threshold, rather than
the most popular boundary. In this way, the greatest amount of lexical content
would be included in the marked up entities. Consequently, for each centroid,
we calculate an extended centroid, or e-centroid, which has the greatest leftmost
(rightmost) boundary with votes above the boundary threshold. The e-centroid
for Figure 2 would therefore have a leftmost boundary at position 9, since the
leftmost boundary receives the most votes of any left boundary. In this case, the
boundary also coincides with the most popular boundary.</p>
      <p>Thirdly, we considered the assignment of CUIs to e-centroids. Since e-centroids
result from a harmonization of several voting systems, this is not entirely
trivial. In the vast majority of cases, the text stretch of the e-centroid does match
the text stretch of at least one of the voting systems, in which case those CUIs
are assigned. In cases where the string extent of no contributing annotation
exactly matched an e-centroid, we experimented with a voting system over CUIs. It
turned out however that such cases were nearly always the result of unusual
combinations of errors from di erent contributing systems. Therefore, rather than
try to assign a CUI, we used the absence of agreement between the e-centroid
and any voting system as a lter on the silver standard.
4.1</p>
      <p>The granularity of harmonization
Finally we re-considered the issue of which annotations to harmonize.</p>
      <p>In the CALBC challenge, the task had been to determine typed mentions,
i.e. the string extents of entities of certain speci ed semantic groups. Thus,
when evaluating challenge submissions over the title Bacterial eye infection in
neonates, a prospective study in a neonatal unit looking for disease and anatomy
markup, it makes sense to evaluate against one centroid centred on infection,
but preferably extending to the left and another centred on eye and preferably
not extending at all. In order to generate these centroids, we run the centroid
algorithm once for all annotations of type disease and independently for all
annotations of type anatomy.
&lt;e grp='ANAT' cui='C1563740'&gt;Visceral adipose tissue&lt;/e&gt;is
particularly &lt;e grp='DISO' cui='C1273518'&gt;responsive to&lt;/e&gt;&lt;e grp='CHEM'
cui='C0376560'&gt;somatropin&lt;/e&gt;</p>
      <p>The Silver Standard CLEF-ER English data is supplied as \hints" on what
entities might be found in the non-English text and, since we are not supplying a
distribution, it could be misleading to suggest that eye infection might be found,
when a di erent disease bacterial eye infection can also be found.</p>
      <p>The centroid algorithm in itself is perfectly neutral over the range of its
inputs, though its outputs will make semantic sense only if the inputs are all
annotations of the same semantic sort. The sort however can be semantic groups,
or types or even individual CUIs, with the deciding factor being only a) the
intended use of the output centroids b) the possibility of data sparsity in a too
ne-grained harmonization. For example, if one harmonizes at the level of CUI
and some systems annotate only the longest match and others only the shortest
match, there is a clear danger that the minimum thresholds will not be met in
either case.</p>
      <p>We therefore experimented with both types of harmonization. The result
was that the gain in the number of entities obtained outweighed the (only small)
sparsity issues that resulted, especially as our thresholding was being set to low
values (3 votes or more). Therefore, harmonization by CUI was our preferred
option and nal choice.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results and Examples</title>
      <p>To illustrate the sort of work that harmonization carries out, Figure 3 shows
the annotations generated by one of the contributing Mantra project partners.
One of the other four partners agreed exactly with the annotation over Visceral
adipose tissue, and another also agreed with the string extent. Consequently,
this annotation is propagated to the silver standard, and in fact it would be
regardless of whether harmonization were carried out by group or by CUI. Two
of the four partners did not support this annotation although they did
annotate the substring adipose tissue and agreed amongst themselves on the CUI for
it. Consequently, in harmonization by CUI, this annotation also propagates to
the silver standard. In harmonization by group, this annotation would be lost, at
least if the \widest boundaries which meet the threshold" is used as the criterion
for selecting from the distribution. No other partner agreed with the DISO
annotation of Figure 3 so this annotation is not replicated in the silver standard.
All partners agreed with the string extent of the annotation over somatropin
but the three other CUI-assigning annotations supported a di erent CUI. In
harmonization by CUI, therefore, this annotation does not propagate, although</p>
      <p>Partner
another annotation with the same string extent (and with the other CUI) does
propagate.</p>
      <p>In Table 1 we give a quantitative evaluation of all 6 partner annotations and
show how many of these annotations nd their way into a harmonized corpus
of a given voting threshold. As explained in the last section, there are a few
annotations that are lost even at a voting threshold of 1. Partner P2 and P4
annotated relatively few CUIs, however, most of these annotations had a broad
support and contributed substantially more CUIs for the harmonized corpora.
Table 1 shows also that our nal voting threshold of 3 removes on average 19%
(SD 14%) of all partner annotations.</p>
      <p>A slightly di erent view on the e ects of voting thresholds on our data is
presented in Table 2. For each corpus, we evaluate the loss of CUI annotations
against a voting threshold of 1. In total, this amounts to 16 million centroids.
The stock of centroids having support from all partners (voting threshold 6)
is quite low, i.e. the common set of annotations is 17% of the whole set of
annotations available from all partners that go into the SSC. For the preparation
of the nal SSC for the challenge, we inspected in detail the most frequent
annotations from the voting thresholds 2 and 3. Raising the threshold from 2 to
3 ltered a substantial amount of noise. Raising the threshold to 4 would have
been viable too. However, threshold 3 was more inline with the strategy to stick
to a reasonably high recall harmonization.</p>
      <p>In Table 3 we investigate the e ect of boundary thresholds with respect to the
original partner annotations for the nal voting threshold of 3. Two questions
are addressed here. How many annotations from partners that are covered by a
e-centroid retain their original boundaries? How many of them get shortened or
extended by di erent e-centroid boundary thresholds? Between 96.4% and 97.8%
of the partner annotations are represented by e-centroids with exactly the same
boundaries. For obvious reasons, a boundary threshold of 1 can only extend
partner annotations. A threshold of 3 leads to a shorter e-centroid representation for
the majority of inexact boundary matches. A boundary threshold of 2 represents
a balanced strategy and was the setting for the nal Silver Standard Corpus.
The CLEF-ER challenge is an interestingly new variant on the classic \named
entity recognition" task. In order to generate good quality English marked up
data to be provided as part of the challenge, an automatic annotation method
was required. The centroid harmonization method has proved a good basis for
building a Silver Standard suitable for the CLEF-ER challenge, even in the
circumstance where its distributional nature is not to be exploited in a direct
evaluation.</p>
      <p>On average, about 81% of all partner annotations are represented by the
harmonized annotations using the nal voting threshold of 3. This amounts to
45% of all centroids that could have been built by a voting threshold of 1.</p>
      <p>For a voting threshold of 3, 97.8% of the partner annotations which go into
the Silver Standard Corpus have exactly the same boundaries as their e-centroid
representations. A boundary threshold of 2 extends a reasonable amount of the
remaining cases with divergent boundaries.</p>
      <p>This work is funded by EU Research Grant agreement 296410.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The Uni ed Medical Language System (UMLS): integrating biomedical terminology</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>32</volume>
          (
          <string-name>
            <surname>Database-Issue</surname>
            <given-names>)</given-names>
          </string-name>
          ,
          <volume>267</volume>
          {
          <fpage>270</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D.</given-names>
            <surname>Rebholz-Schuhmann</surname>
          </string-name>
          et al.:
          <article-title>Assessment of NER solutions against the rst and second CALBC Silver Standard Corpus</article-title>
          .
          <source>Journal of Biomedical Semantics</source>
          <volume>2</volume>
          (
          <issue>11</issue>
          /
          <year>2011</year>
          2011), http://www.jbiomedsem.com/content/2/S5/S11
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lewin</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kafkas</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rebholz-Schuhmann</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Centroids:
          <article-title>Gold standards with distributional variation</article-title>
          .
          <source>In: Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12)</source>
          .
          <source>European Language Resources Association (ELRA)</source>
          , Istanbul (May
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>