<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Task 1: ShARe/CLEF eHealth Evaluation Lab 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sameer Pradhan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noemie Elhadad</string-name>
          <email>noemie@dbmi.columbia.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brett R. South</string-name>
          <email>brett.south@hsc.utah.edu</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Martinez</string-name>
          <email>david.martinez@nicta.com.au</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lee Christensen</string-name>
          <email>leenlp@q.com</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amy Vogel</string-name>
          <email>amy.vogel@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hanna Suominen</string-name>
          <email>hanna.suominen@nicta.com.au</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wendy W. Chapman</string-name>
          <email>wwchapman@ucsd.edu</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guergana Savova</string-name>
          <email>guergana.savovag@childrens.harvard.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Boston Chlidren's Hospital and Harvard Medical School</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Columbia University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>NICTA and The Australian National University</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>NICTA and The University of Melbourne</institution>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of California San Diego</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>University of Utah</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This report outlines the Task 1 of the ShARe/CLEF eHealth evaluation lab pilot. This task focused on identi cation (1a) and normalization (1b) of diseases and disorders in clinical reports. It used annotations from the ShARe corpus. A total of 22 teams competed in Task 1a and 17 of them also participated Task 1b. The best systems had an F1 score of 0.75 (0.80 Precision, 0.71 Recall) in Task 1a and an accuracy of 0.59 in Task 1b. The organizers have made the text corpora, annotations, and evaluation tools available for future research and development.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Processing</kwd>
        <kwd>Text Normalization</kwd>
        <kwd>Medical Informatics</kwd>
        <kwd>Reference Standard Generation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A large amount of very useful information { both for the medical researchers and
the patients { is present in the form of unstructured text within the clinical notes
and discharge summaries that form a patient's medical history. Adapting and
extending NLP techniques to mine this information can open doors to better,
novel, clinical studies on one hand, and help patients understand the contents of
their clinical records on the other. Organization of this shared task helps establish
state-of-the-art baselines and paves way to further explorations. The shared task
was one of three shared tasks organized at the CLEF eHealth Evaluation Labs [
        <xref ref-type="bibr" rid="ref1 ref2">1,
2</xref>
        ]
? WWC, BRS, and DLM led the task; WWC, BRS, DLM, NE, SP, and GS de ned
the task; GS and NE led the annotation e ort; AV provided coordination and
management of the annotations; HS co-chaired the lab; DLM, BRS and LC processed
and distributed the dataset; and DM and WWC led result evaluations
      </p>
    </sec>
    <sec id="sec-2">
      <title>Data</title>
      <p>The ShARe corpus7 comprises of annotations over de-identi ed clinical reports
from from US intensive care (version 2.5 of the MIMIC II database8.) The
corpus consisted of discharge summaries and electrocardiogram, echocardiogram,
and radiology reports. Although the clinical reports were de-identi ed, they still
needed to be treated with appropriate care and respect. Hence, all participants
were required to register to the lab, obtain a US human subjects training
certi cate9, create an account to a password-protected site on the Internet, specify
the purpose of data usage, accept the data use agreement, and get their account
approved.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Task Description</title>
      <p>
        The annotation of disorder mentions in clinical reports was carried out as part of
the ongoing ShARe project10. For this task in the evaluation lab, the focus was
on the annotation of disorder mentions only. As such, there were two parts to
the annotation: identifying a span of text as a disorder mention and mapping the
span to a UMLS [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] CUI. Each note was annotated by two professional coders
trained for this task, followed by an open adjudication step. UMLS11 represented
over 130 lexicons/thesauri with terms from a variety of languages. It integrated
resources used world-wide in clinical care, public health, and epidemiology. It
also provided a semantic network in which every concept is represented by its
CUI and is semantically typed [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A disorder mention was de ned as any span of
text which can be mapped to a concept in SNOMED-CT and which belongs to
the Disorder semantic group12. A concept was in the Disorder semantic group if
it belonged to one of the following UMLS semantic types: Congenital
Abnormality; Acquired Abnormality; Injury or Poisoning; Pathologic Function; Disease or
Syndrome; Mental or Behavioral Dysfunction; Cell or Molecular Dysfunction;
Experimental Model of Disease; Anatomical Abnormality; Neoplastic Process;
and Signs and Symptoms. The annotations covered about 181,000 words.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Methods</title>
      <p>
        The following evaluation criteria were used:
7 https://www.clinicalnlpannotation.org
8 Multiparameter Intelligent Monitoring in Intensive Care http://mimic.physionet.org
9 The course was available free of charge on the Internet, for
example, via the CITI Collaborative Institutional Training Initiative at
https://www.citiprogram.org/Default.asp or the US National Institutes of Health
(NIH) at http://phrp.nihtraining.com/users/login.php.
10 https://www.clinicalnlpannotation.org
11 https://uts.nlm.nih.gov/home.html
12 Note that this de nition of Disorder semantic group did not include the Findings
semantic type, and as such di ered from the one of UMLS Semantic Groups, available
at http://semanticnetwork.nlm.nih.gov/SemGroups
1a. correctness in identi cation of the character spans of disorders,
1b. correctness in mapping disorders to SNOMED-CT codes,
In Tasks 1a and 1b each participating team was permitted to upload the outputs
of up to two systems. Task 1b was optional for Task 1 participants. Teams
were allowed to use additional annotations in their systems, but this counted
towards the permitted systems; systems that used annotations outside of those
provided were evaluated separately. The evaluation for all tasks was conducted
using the blind, withheld test data. The participants were provided a training
set containing clinical text as well as pre-annotated spans and named entities
for disorders (Tasks 1a and 1b). For Task 1a, participants were instructed to
develop a system that predicts the spans for disorder named entities. For Tasks
1b, participants were instructed to develop a system that predicts the
SNOMEDCT code. The outputs needed to follow the annotation format. The corpus of
reports was split into 200 training and 100 testing. The system performance was
evaluated agaist the criteria by using the F1 score in Task 1a and Accuracy
in Tasks 1b. We relied on non-parametric statistical signi cance tests called
random shu ing [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] to better compare the measure values for the systems and
benchmarks. In Task 1a, the F1 score was de ned as the harmonic mean of
Precision (P) and Recall (R); P as nT P =(nT P + nF P ); R as nT P =(nT P + nF N );
nT P as the number of instances, where the spans identi ed by the system and
gold standard were the same; nF P as the number of spurious spans by the
system; and nF N as the number of missing spans by the system. We referred to
the Exact (Relaxed) F1-score if the system span is identical to (overlaps) the
gold standard span. In Tasks 1b the Accuracy was de ned as the number of
preannotated spans with correctly generated code divided by the total number of
pre-annotated spans. In both tasks, the Exact Accuracy and Relaxed Accuracy
were measured. In the Exact Accuracy for Task 1b, total was de ned as the total
number of gold standard named entities. In this case, the system was penalised
for incorrect code assignment for annotations that were not detected by the
system. In the Relaxed Accuracy for Task 1b, total was de ned as the total
number of named entities with strictly correct span generated by the system. In
this case, the system was only evaluated on annotations that were detected by
the system.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>System Results</title>
      <p>A total of 22 teams competed in Task 1a and 16 of them also participated Task
1b. The performance of these systems is detailed in Tables 1 and 2. The best
systems had an F1 score of 0.75 (0.80 Precision, 0.71 Recall) in Task 1a and an
accuracy of 0.59 in Task 1b.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion</title>
      <p>We have created a reference standard with high inter-annotator agreement and
evaluated systems on the task of identi cation and normalization of diseases and</p>
      <p>System ID (fteamg.fsystemg)
disorders apprearing in clinical reports. The results have demonstrated that an
NLP system can complete this task with reasonably high accuracy. We plan to
annotate more data and perform another evaluation in the near future.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We greatly appreciate the hard work and feedback of our program committee
members and annotators, including, but not limited to David Harris, Glenn
Zaramba, Erika Siirala, Qing Zeng, Tyler Forbush, Jianwei Leng, Maricel
Angel, Erikka Siirala, Helja Lundgren-Laine, Jenni Lahdenmaa, Laura Maria
Murtola, Marita Ritmala-Castren, Riitta Danielsson-Ojala, Saija Heikkinen, and Sini
Koivula. This shared task was partially supported by Shared Annotated
Resources (ShARe) project NIH 5R01GM090187, VAHSR&amp;D HIR 08-374, NICTA
(National Information and Communications Technology Australia), funded by
the Australian Government as represented by the Department of Broadband,
Communications and the Digital Economy and the Australian Research Council
through the ICT Centre of Excellence program), and NLM 5T15LM007059.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Salantera,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Velupillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.W.</given-names>
            ,
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Pradhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>South</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.R.</given-names>
            ,
            <surname>Mowery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.L.</given-names>
            ,
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.J.F.</given-names>
            ,
            <surname>Leveling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Zuccon</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          :
          <article-title>Overview of the share/clef ehealth evaluation lab 2013</article-title>
          . In: Proceedings of ShARe/CLEF eHealth Evaluation Labs. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Mowery</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>South</surname>
            ,
            <given-names>B.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Christensen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murtola</surname>
            ,
            <given-names>L.M.</given-names>
          </string-name>
          , Salantera,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Martinez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Elhadad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Pradhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            , ,
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <surname>W.W.</surname>
          </string-name>
          :
          <article-title>Task 2: Share/clef ehealth evaluation lab 2013</article-title>
          . In: Proceedings of ShARe/CLEF eHealth Evaluation Labs. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliver</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Shortli e, E.:
          <article-title>The Uni ed Medical Language System: Towards a collaborative approach for solving terminologic problems</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ) (
          <year>1998</year>
          )
          <volume>12</volume>
          {
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCray</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Exploring semantic groups through visual approaches</article-title>
          .
          <source>Journal of Biomedical Informatics</source>
          <volume>36</volume>
          (
          <year>2003</year>
          )
          <volume>414</volume>
          {
          <fpage>432</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Yeh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>More accurate tests for the statistical signi cance of result di erences</article-title>
          .
          <source>In: Proceedings of the 18th Conference on Computational Linguistics (COLING)</source>
          , Saarbrucken,
          <string-name>
            <surname>Germany</surname>
          </string-name>
          (
          <year>2000</year>
          )
          <volume>947</volume>
          {
          <fpage>953</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>