<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Gold-Standard Ontology-Based Annotation of Concepts in Biomedical Text in the CRAFT Corpus: Updates and Extensions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael Bada</string-name>
          <email>mike.bada@ucdenver.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lawrence Hunter</string-name>
          <email>larry.hunter@ucdenver.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nicole Vasilevsky</institution>
          ,
          <addr-line>Melissa Haendel</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ontology Development Group, Library Oregon Health &amp; Science University Portland</institution>
          ,
          <addr-line>OR</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Colorado School of Medicine Aurora</institution>
          ,
          <addr-line>CO</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>-Ontologies are increasingly used for semantic integration across disparate curated biomedical resources, while gold-standard annotated corpora are needed for accurate training and evaluation of text-mining tools. Bringing together the respective power of these, we created the Colorado Richly Annotated Full-Text (CRAFT) Corpus, a collection of full-length, open-access biomedical journal articles that have been manually annotated both syntactically and semantically with select Open Biomedical Ontologies (OBOs), the first release of which includes ~100,000 annotations of concepts mentioned in the text of 67 articles and mapped to the classes of eight prominent OBOs. Here we present our continuing work on the corpus, including updated versions of these annotations with newer versions of the ontologies, new annotations made with two additional OBOs, annotations made with newly created extension classes defined in terms of existing classes of the ontologies, and new annotations of roots of prefixed and suffixed words.</p>
      </abstract>
      <kwd-group>
        <kwd>annotation</kwd>
        <kwd>corpus</kwd>
        <kwd>markup</kwd>
        <kwd>ontology</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        With the ever-rising amount of biomedical literature, it is
increasingly difficult for scientists to keep up with the
published work in their fields of research, much less related
ones. The use of natural language processing (NLP) tools can
make the literature more accessible by aiding concept
recognition and information extraction. As NLP-based
approaches have been increasingly used for biocuration, so too
have biomedical ontologies, whose use enables semantic
integration across disparate curated resources, and millions of
biomedical entities have been annotated with them. Particularly
important are the Open Biomedical Ontologies (OBOs), a set
of open, orthogonal, interoperable ontologies formally
representing knowledge over a wide range of biology,
medicine, and related disciplines [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Manually annotated document corpora have become critical
gold-standard resources for the training and testing of
biomedical NLP systems. This was the motivation for the
creation of the Colorado Richly Annotated Full-Text (CRAFT)
Corpus, a collection of 97 full-length, open-access journal
articles from the biomedical literature [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. Within these
articles, each mention of the concepts explicitly represented in
      </p>
      <p>
        All continuing work on the concept annotations of the
CRAFT Corpus was performed in Knowtator, a plugin to
Protégé-Frames [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. (as was done for the v1.0 concept
annotations). The lead annotator (MB) made updates to the
v1.0 concept annotations using newer versions of the
ontologies that had been used to mark up the articles by
removing annotations of obsoleted classes, editing previously
made annotations, and creating new annotations for new
classes. A list of approximately 20 prefixes and suffixes was
compiled, and roots of words with these affixes were
annotated as their unaffixed analogs would be. As the
updating progressed with each ontology, corresponding
extension classes were created to use for further annotation.
      </p>
      <p>
        Annotation of the corpus with the Molecular Process
Ontology (MOP) and Uberon was performed in one primary
round (by NV) followed by a review (by MB) using the
original concept annotation guidelines [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Roots of words with
aforementioned affixes were also annotated, and extension
classes were also created and used for additional annotation.
The articles were annotated with a single ontology at a time
and a batch at a time (8 articles per batch for the MOP and 4
articles per batch for Uberon), and interannotator agreement
(IAA) was calculated for each batch using Knowtator’s built-in
IAA calculation functionality. The curators strove for IAA ≥
90% for each annotation batch.
      </p>
    </sec>
    <sec id="sec-2">
      <title>III. RESULTS AND DISCUSSION</title>
      <p>So as to remain current and relevant, the v1.0 concept
annotations of the corpus are being reviewed and updated by
addition, editing, and deletion of annotations as appropriate,
relying on newer versions of the eight OBOs previously used.
Updating with four of these has been completed.</p>
      <p>The extension of annotation of specific affixed root words
is largely for consistency: In the v1.0 corpus, any whitespace
or punctuation character could serve as an annotation
delimiter; thus, “chromatin” of “anti-chromatin” would be
annotated with the Gene Ontology class for chromatin
(GO:0000785), but it could not be annotated within
“antichromatin”, as there is no delimiter. The rendering of
such affixes is variable in that they can be nondelimited from
their root words or delimited by whitespace or punctuation, so
with this updating, the markup of such affixed words is now
more consistent; furthermore, additional knowledge is
captured. A specific list of such affixes to consider has been
compiled and will be provided with the next release.</p>
      <p>While creating the concept annotations for the v1.0 corpus,
we encountered a variety of difficulties with annotating
exclusively with explicitly represented OBO classes, including
class ambiguity, lack of sufficiently generic classes, lack of
classes for words consisting of combinations of multiple
ontology classes, representation of the same concept in
multiple ontologies and incompleteness of ontologies. To
ameliorate these issues, we have been creating and using
specific extension classes for concept annotations for the
corpus update. All of these are formally defined in terms of
explicitly represented OBO classes, and we intend to make
these definitions available in OWL files in the next release.
However, we also intend to release the annotations in sets both
including and excluding these extension classes for users who
respectively do and do not wish to make use of annotations
with such classes in their work.</p>
      <p>
        Finally, for the purpose of capturing additional types of
biomedically relevant concepts, annotations have been created
for the articles of the corpus using the classes of the MOP
ontology of chemical processes [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] and the Uberon anatomical
ontology [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Tables 1 and 2 display relevant statistics for the
67 articles of the public set, excluding and including use of
extension classes, and IAA statistics are presented in Figure 1.
ontology
ontology
total # annotations
total #
unique
concepts
19 / 20
850 / 898
      </p>
      <p>average #
unique concepts
per article</p>
      <p>median #
unique concepts
per article</p>
    </sec>
    <sec id="sec-3">
      <title>IV. CONCLUSIONS</title>
      <p>We have presented our continuing work on the
goldstandard concept annotations of the CRAFT Corpus, including
updated versions of the annotations with newer versions of
ontologies, new annotations made with additional OBOs,
annotations made with newly created ontology extension
classes, and new annotations of roots of prefixed and suffixed
words. We intend to soon release these updated annotations in
future versions of the corpus, and we also have longer-term
plans for further development of the corpus.</p>
    </sec>
    <sec id="sec-4">
      <title>ACKNOWLEDGMENT</title>
      <p>This work was supported by grant DARPA-BAA-14-14.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>[1] http://www.obofoundry.org</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Bada</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eckert</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shipley</surname>
            <given-names>K</given-names>
          </string-name>
          , et al. (
          <year>2012</year>
          )
          <article-title>Concept Annotation in the CRAFT Corpus</article-title>
          .
          <source>BMC Bioinform</source>
          <volume>13</volume>
          :
          <fpage>161</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Verspoor</surname>
            <given-names>K</given-names>
          </string-name>
          et al. (
          <year>2012</year>
          )
          <article-title>A corpus of full-text journal articles is a robust evaluation tool for revealing differences in performance of biomedical natural language processing tools</article-title>
          .
          <source>BMC Bioinform</source>
          <volume>13</volume>
          :
          <fpage>207</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Funk</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumgartner</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roeder</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bada</surname>
            <given-names>M.</given-names>
          </string-name>
          et al. (
          <year>2014</year>
          )
          <article-title>Large-scale biomedical concept recognition: an evaluation of current automatic annotators and their parameters</article-title>
          .
          <source>BMC Bionform</source>
          <volume>15</volume>
          :
          <fpage>59</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Liu</surname>
            <given-names>H</given-names>
          </string-name>
          et al. (
          <year>2012</year>
          )
          <article-title>BioLemmatizer: a lemmatization tool for morphological processing of biomedical text</article-title>
          .
          <source>J Biomed Semantics</source>
          <volume>3</volume>
          :
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nunes</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campos</surname>
            <given-names>D</given-names>
          </string-name>
          et al. (
          <year>2013</year>
          )
          <article-title>BeCAS: biomedical concept recognition services and visualization</article-title>
          .
          <source>Bioinform</source>
          <volume>29</volume>
          (
          <issue>15</issue>
          ),
          <fpage>1915</fpage>
          -
          <lpage>1916</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Ogren</surname>
            <given-names>P.V.</given-names>
          </string-name>
          (
          <year>2006</year>
          )
          <article-title>Knowtator: a Protégé plug-in for annotated corpus construction</article-title>
          .
          <source>Proc Hum Lang Tech Conf N Am Chap Assoc Comp Ling.</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bada</surname>
            <given-names>M et</given-names>
          </string-name>
          <year>al</year>
          . (
          <year>2010</year>
          )
          <article-title>An overview of the CRAFT concept annotation guidelines</article-title>
          .
          <source>Proc 4th Ling Annot Wkshp, Assoc Comp Ling</source>
          ,
          <fpage>207</fpage>
          -
          <lpage>211</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>[9] http://obofoundry.org/ontology/mop.html</mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mungall</surname>
            <given-names>CJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torniai</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gkoutos</surname>
            <given-names>GV</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lewis</surname>
            <given-names>SE</given-names>
          </string-name>
          et al. (
          <year>2011</year>
          )
          <article-title>Uberon, an integrative multi-species anatomy ontology</article-title>
          .
          <source>Genome Biol</source>
          <volume>13</volume>
          :
          <fpage>R5</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>