<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Temporal Anno-
tation in the Clinical Domain. Transac-
tions of the Association for Computational
Linguistics</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>The E3C Pro ject: European Clinical Case Corpus</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>El proyecto E</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>C: European Clinical Case Corpus</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bernardo Magnini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Begon~a Altuna</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alberto Lavelli</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manuela Speranza</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Zanoli</string-name>
          <email>zanolig@fbk.eu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fondazione Bruno Kessler</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad del Pa s Vasco/Euskal Herriko Unibertsitatea</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>[Styler et al.2014a] Styler</institution>
          ,
          <addr-line>W., G. Savova, M. Palmer, J. Pustejovsky, T. O'Gorman, and P. C.</addr-line>
          <institution>deGroen. 2014a. THYME Annotation Guidelines. Technical report, University of Colorado</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <volume>2</volume>
      <issue>143</issue>
      <fpage>17</fpage>
      <lpage>20</lpage>
      <abstract>
        <p>The European Clinical Case Corpus (E3C) project aims at collecting and annotating a large corpus of clinical documents in ve European languages (Spanish, Basque, English, French and Italian), which will be freely distributed. Annotations include temporal information, to allow temporal reasoning on chronologies, and information about clinical entities based on medical taxonomies, to be used for semantic reasoning.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The core of the corpus is a manually
annotated dataset of clinical cases. A clinical
case reports statements of a clinical practice,
presenting the reason for a clinical visit, the
description of physical exams, and the
assessment of the patient's situation (Box 1).</p>
    </sec>
    <sec id="sec-2">
      <title>1https://www.european-language-grid.eu/</title>
      <p>open-calls/
A 25-year-old man with a history
of Klippel-Trenaunay syndrome
presented to the hospital with
mucopurulent bloody stool and epigastric
persistent colic pain for 2 wk. Colonoscopy
showed continuous super cial ulcers and
bleeding. Subsequent gastroscopy
revealed mucosa with di use edema,
ulcers, errhysis, and granular and friable
changes in the stomach and duodenal
bulb. A diagnosis of GDUC was
considered. The patient hesitated about iv
corticosteroids, so he was treated with
pentasa 3.2 g/d. After 0.5 mo of treatment,
the symptoms achieved complete
remission. Follow-up examinations showed no
evidence of recurrence for 26 mo.</p>
      <p>Box 1: Sample clinical case.
2</p>
      <sec id="sec-2-1">
        <title>Motivation and Related Work</title>
        <p>The main motivation of the project is
creating a clinical document corpus that can be
freely redistributable and that contains
temporal information and clinical entity
annotations. The annotation of temporal
informaCopyright © 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
tion is the core e ort of the project, so we
expect the E3C corpus to be useful in tasks such
as event ordering and chronology generation.</p>
        <p>These decisions are justi ed by three
common issues on Natural Language Processing
(NLP) for the clinical domain, listed below.</p>
        <p>First, the E3C corpus is formed of already
published documents, mostly clinical cases in
journals so as to overcome the patient privacy
issues that Electronic Health Records (EHR)
often convey. We have opted for selecting
documents that already allow redistribution
to ensure the E3C corpus will be easily usable
for the research community.</p>
        <p>
          Secondly, as E3C is a multilingual
corpus, it helps reducing the gap between
English and other languages in terms of
available data for research. A large dataset of
clinical cases in Spanish is already available,
SPACCC
          <xref ref-type="bibr" rid="ref4">(Intxaurrondo et al., 2018)</xref>
          , and we
have expanded and enriched it. For a
lowresourced language as Basque, instead, E3C
is the rst clinical narrative corpus. In
addition, the project will also provide the NLP
community with the rst Creative Commons
clinical case corpora for French and Italian.
        </p>
        <p>
          Thirdly, E3C is centered on temporal
information in clinical narratives, which has
not been often targeted by scholars. Most
of the attention has been focused on clinical
entity extraction and classi cation
          <xref ref-type="bibr" rid="ref1 ref1 ref3 ref3 ref6 ref6">(Schulz et
al., 2020; Grabar et al., 2019; Dreisbach et
al., 2019; Luo et al., 2017)</xref>
          and only a few key
initiatives have addressed temporal
information processing, e.g. the THYME annotation
scheme (Styler et al., 2014b). THYME and
other o -spins have been used to annotate
clinical case corpora such as the i2b2
temporal relation corpus (Sun, Rumshisky, and
Uzuner, 2013), and have been used in clinical
narratives processing challenges, e.g., CLEF
eHealth
          <xref ref-type="bibr" rid="ref1 ref3 ref6">(Kelly et al., 2019)</xref>
          . This information
could be then merged with the information
on structured data collections, e.g. MIMIC
III
          <xref ref-type="bibr" rid="ref5">(Johnson et al., 2016)</xref>
          , enriching it.
3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Data Collection and</title>
      </sec>
      <sec id="sec-2-3">
        <title>Distribution</title>
        <p>The E3C corpus is a collection of both
existing corpora (e.g., the SPACCC corpus)
and published texts extracted from di
erent sources, such as PubMed2 (journal
abstracts), SciELO3 and the PanAfrican
Medi</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2https://pubmed.ncbi.nlm.nih.gov/</title>
    </sec>
    <sec id="sec-4">
      <title>3https://scielo.org/</title>
      <p>L2
162
113
171
168
174
Layer 1: about 25K tokens per language
(around 80 documents) of clinical narratives
with full manual annotation of clinical
entities, temporal information and factuality, for
benchmarking and linguistic analysis.</p>
      <p>Layer 1 is the core of the E3C corpus and
special attention has been awarded to
creating a balanced document set in terms of
size. Short (&lt;200 tokens), medium (200{
400 tokens) and long (400{600 tokens) texts
have been selected, as we presumed that text
length would directly a ect the temporal
information in text; the longer the text, the
more complex the temporal graph.</p>
      <p>Layer 2: 50{100K tokens per language of
clinical narratives with automatic annotation
of clinical entities and manual check of the
annotation of a small sample (about 10%).
Layer 3: about 1M tokens per language of
non-annotated medical documents (not
necessarily clinical narratives) to be exploited by
semi-supervised approaches.</p>
      <p>All the layers are covered for Spanish,
English, Italian and French. For Basque, Layer
2 (14K token) and 3 (600K token) are
covered only partially. In Table 1 we summarize
the distribution of documents per layer and
language. In the case of L1, the amount of
texts provides the information on how many
distinct temporal graphs or chronologies we
will be able to build from the dataset.
3.2</p>
      <sec id="sec-4-1">
        <title>Corpus Distribution</title>
        <p>The nal E3C corpus will be available for
download from the ELG platform
repository5. All documents will be released under</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4https://www.panafrican-med-journal.com/</title>
    </sec>
    <sec id="sec-6">
      <title>5https://live.european-language-grid.eu</title>
      <p>Creative Commons licenses. This is possible
as a large part of the texts in the corpus were
already released under Creative Commons
licenses and as permission for free distribution
has been requested and obtained from the
original owners of some of the documents.
4</p>
      <sec id="sec-6-1">
        <title>Corpus Annotation</title>
        <p>The E3C corpus will contain two types of
annotations: (i) temporal information and
factuality, and (ii) annotation of clinical entities,
speci cally disorders.</p>
        <p>Manual annotation has been performed
by a team of NLP students and researchers.
More precisely, temporal information and
clinical entity annotation guidelines have
been de ned by two teams of three and four
experts respectively. For the manual
annotation e ort, eight people have been trained
and are completing the annotation of
temporal information, while clinical entity
annotation is being conducted by ve people.
4.1</p>
        <sec id="sec-6-1-1">
          <title>Temporal Information</title>
        </sec>
        <sec id="sec-6-1-2">
          <title>Annotation</title>
          <p>Temporal information annotation is
performed following the THYME annotation
guidelines (Styler et al., 2014a) with some
minor adaptations (Magnini et al., 2020). This
scheme provides tags for events, time
expressions, temporal relations between events
and/or time expressions, and aspectual
relations between events. For each tag, a set of
attribute-value pairs allow to make the
relevant features explicit. In order to mark
information that further contributes to the clinical
history of a patient, we have added three
categories: measurements and test results,
actors (for the patient itself, health
professionals and other participants), body parts.</p>
          <p>In Figure 1 a simpli ed temporal
information annotation is displayed. Events are in
dark blue and their contextual modality
(ACTUAL), document time relation (BEFORE)
and polarity (POS) are highlighted. The 2
wk time expression (in gray) is classi ed as a
duration and the information about the actor
(the patient) is represented in light blue.
4.2</p>
        </sec>
        <sec id="sec-6-1-3">
          <title>Clinical Entity Annotation</title>
          <p>Clinical entity annotation focuses on
disorders. Following UMLS6 a disorder is de ned
as \a de nite pathologic process with a
characteristic set of signs and symptoms".</p>
          <p>
            In E3C, we mark disorders mentioned in
the text and assign them an UMLS concept
unique identi er (CUI). Disorder identi
cation and coding is performed following an
adaptation of the ShARe annotation
guidelines
            <xref ref-type="bibr" rid="ref2">(Elhadad et al., 2012)</xref>
            . In concept
selection, we restrict to the UMLS semantic
group Disorder, which includes the Finding
semantic type in addition to those proposed
by ShARe. Figure 1 shows disorders, marked
in red, and their UMLS codes.
5
          </p>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>Conclusions and Future Work</title>
        <p>The E3C project aims at reducing the lack of
available resources for clinical NLP,
gathering a large number of clinical narratives and
focusing on languages other than English.
After completing the project, the E3C
corpus and the associated resources (baselines,
scorers, etc.) will be available for research
under a Creative Commons license, which
will facilitate its acquisition and reusability.
More speci cally, since the corpus contains
information for temporal reasoning as well as
clinical entity mentions, it will be useful for
works on semantic interpretation of clinical
texts. The fact that the corpus contains texts
in ve languages will allow linguistic
comparisons as well as experimentation on transfer
learning. We also consider that the E3C
corpus is a resource that could be employed in a
series of evaluation challenges due to its
atypical contents and types of annotations.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Acknowledgements</title>
        <p>This work was partially supported by the
European Language Grid project through its</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6https://uts.nlm.nih.gov/uts/umls/home</title>
      <p>open call for pilot projects (EU grant no.
825627), and by the Basque Government
post-doctoral grant POS 2020 2 0026.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Dreisbach et al.2019]
          <string-name>
            <surname>Dreisbach</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Koleck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Bourne</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bakken</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A systematic review of natural language processing and text mining of symptoms from electronic patientauthored text data</article-title>
          .
          <source>International Journal of Medical Informatics</source>
          ,
          <volume>125</volume>
          :
          <fpage>37</fpage>
          {
          <fpage>46</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Elhadad et al.2012]
          <string-name>
            <surname>Elhadad</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Savova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zaramba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Harris</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Vogel</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>ShARe Guidelines for the Annotation of Modi ers for Disorders in Clinical Notes</article-title>
          .
          <source>Technical report</source>
          , Columbia University.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Grabar et al.2019]
          <string-name>
            <surname>Grabar</surname>
            , N.,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Hamon</surname>
            , and
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Claveau</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Recherche et extraction d'information dans des cas cliniques. Presentation de la campagne d'evaluation DEFT 2019</article-title>
          . In Actes du De Fouille de Textes 2019, pages
          <fpage>7</fpage>
          {
          <fpage>16</fpage>
          , Toulouse, France.
          <source>Actes DEFT</source>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Intxaurrondo et al.2018]
          <string-name>
            <surname>Intxaurrondo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marimon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gonzalez-Agirre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Lopez-Mart n</surname>
          </string-name>
          , H. Rodr guez, J. Santamar a, M. Villegas, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Krallinger</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Finding Mentions of Abbreviations and Their De nitions in Spanish Clinical Cases: The BARR2 Shared Task Evaluation Results</article-title>
          .
          <source>In Proceedings of the Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval</source>
          <year>2018</year>
          ), pages
          <fpage>280</fpage>
          {
          <fpage>289</fpage>
          ,
          <string-name>
            <surname>Seville</surname>
          </string-name>
          , Spain.
          <source>Spanish Society for Natural Language Processing.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Johnson et al.2016] Johnson, A. E.,
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Pollard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-w. H.</given-names>
            <surname>Lehman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ghassemi</surname>
          </string-name>
          , B. Moody, P. Szolovits,
          <string-name>
            <given-names>L. Anthony</given-names>
            <surname>Celi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R. G.</given-names>
            <surname>Mark</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>MIMIC-III, a freely accessible critical care database</article-title>
          .
          <source>Scienti c Data</source>
          ,
          <volume>3</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Kelly et al.2019]
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Suominen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Neves</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kanoulas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Azzopardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Spijker</surname>
          </string-name>
          , G. Zuccon,
          <string-name>
            <given-names>H.</given-names>
            <surname>Scells</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Palotti</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2019</article-title>
          . In F. Crestani,
          <string-name>
            <given-names>M.</given-names>
            <surname>Braschler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Savoy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rauber</surname>
          </string-name>
          , H. Muller, D. E.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>