<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Clinical Information Extraction at the CLEF eHealth Evaluation lab 2016</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aurelie Neveol</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K. Bretonnel Cohen</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cyril Grouin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thierry Hamon</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Lavergne</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Liadh Kelly</string-name>
          <email>liadh.kelly@tcd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorraine Goeuriot</string-name>
          <email>lorraine.goeuriot@imag.fr</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gregoire Rey</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aude Robert</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xavier Tannier</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre Zweigenbaum</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ADAPT Centre, Trinity College</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INSERM-CepiDC</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIMSI, CNRS, Universite Paris-Saclay</institution>
          ,
          <addr-line>Orsay</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Univ. Paris-Sud</institution>
          ,
          <addr-line>Orsay</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Universite Grenoble Alpes</institution>
          ,
          <addr-line>Grenoble</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Universite Paris Nord</institution>
          ,
          <addr-line>Villetaneuse</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>University of Colorado</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper reports on Task 2 of the 2016 CLEF eHealth evaluation lab which extended the previous information extraction tasks of ShARe/CLEF eHealth evaluation labs. The task continued with named entity recognition and normalization in French narratives, as o ered in CLEF eHealth 2015. Named entity recognition involved ten types of entities including disorders that were de ned according to Semantic Groups in the Uni ed Medical Language System R (UMLS R ), which was also used for normalizing the entities. In addition, we introduced a largescale classi cation task in French death certi cates, which consisted of extracting causes of death as coded in the International Classi cation of Diseases, tenth revision (ICD10). Participant systems were evaluated against a blind reference standard of 832 titles of scienti c articles indexed in MEDLINE, 4 drug monographs published by the European Medicines Agency (EMEA) and 27,850 death certi cates using Precision, Recall and F-measure. In total, seven teams participated, including ve in the entity recognition and normalization task, and ve in the death certi cate coding task. Three teams submitted their systems to our newly o ered reproducibility track. For entity recognition, the highest performance was achieved on the EMEA corpus, with an overall F-measure of 0.702 for plain entities recognition and 0.529 for normalized entity recognition. For entity normalization, the highest performance was achieved on the MEDLINE corpus, with an overall F-measure of 0.552. For death certi cate coding, the highest performance was 0.848 F-measure.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Processing</kwd>
        <kwd>Named Entity Recognition</kwd>
        <kwd>Entity Linking</kwd>
        <kwd>Text Classi cation</kwd>
        <kwd>UMLS</kwd>
        <kwd>French</kwd>
        <kwd>Biomedical Text</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        This paper describes an investigation of information extraction and
normalization (also called \entity linking") from French-language health documents. The
methodology applied is the shared task model. In shared tasks, multiple groups
agree on a \shared" task de nition, a shared data set, and a shared evaluation
metric. The idea is to allow evaluation of multiple approaches to a problem
while minimizing avoidable di erences related to the task de nition, the data
used, and the gure of merit applied [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ].
      </p>
      <p>
        Over the past three years, CLEF eHealth o ered challenges addressing several
aspects of clinical information extraction (IE) including named entity
recognition, normalization [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] and attribute extraction [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Initially, the focus was
on a widely studied type of corpus, namely written English clinical text [
        <xref ref-type="bibr" rid="ref3 ref5">3, 5</xref>
        ].
Starting in 2015, the lab's IE challenge evolved to address lesser studied corpora,
including biomedical texts in a language other than English i.e., French [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. This
year, we continue to o er a shared task based on a large set of gold standard
annotated corpora in French. In addition to named entity extraction and
entity normalization already o ered in 2015 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], we introduced a coding task that
required normalized entity extraction at the sentence level.
      </p>
      <p>
        The signi cance of this work comes from the observation that challenges and
shared tasks have had a signi cant role in advancing Natural Language
Processing (NLP) research in the clinical and biomedical domains [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ], especially for the
extraction of named entities of clinical interest [9{12], and entity normalization
[11, 13{16].
      </p>
      <p>
        One of the goals for this shared task is to foster the development of NLP
tools for French in spite of the known discrepancies in language resources
available for French and other languages in the biomedical domain, compared to
English [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Findings of last year's lab were that while there was a sustained
interest in addressing French from teams all over the world, results were very
heterogenous depending on methods and resources used, as well as technical issues
encountered [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This year's lab suggests increased maturity of the task as major
technical problems are now tackled, performance increases, and reproducibility
is introduced as an additional goal.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Material and Methods</title>
      <p>In the CLEF eHealth 2016 Evaluation Lab Task 2, two datasets were used.
The QUAERO French Medical corpus was used for named entity extraction and
normalization. The CepiDC corpus was used for coding. Further details on the
datasets, tasks and evaluation metrics are given below.
2.1</p>
      <sec id="sec-2-1">
        <title>Datasets</title>
      </sec>
      <sec id="sec-2-2">
        <title>The QUAERO French Medical corpus The QUAERO French Medical</title>
        <p>
          Corpus [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] was used for named entity extraction and normalization in CLEF
eHealth 2015 (task 1b) and CLEF eHealth 2016 (task 2). The dataset will be
shared freely with the community after the challenge results have been
announced. For a detailed description of the QUAERO corpus, we refer interested
readers to the corpus website http://quaerofrenchmed.limsi.fr/ and to the
2016 task 1b report [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], which include a detailed description of the annotation
guidelines and excerpts of the corpus. Table 1 presents statistics for the speci c
sets provided to participants in CLEF eHealth 2016. The training set released in
the CLEF eHealth 2016 Task 2 challenge corresponds to the training set provided
in the CLEF eHealth 2015 Task 1b challenge, the development set corresponds
to the test set provided in the CLEF eHealth 2015 Task 1b challenge, and the
test set was previously unreleased. EMEA documents were divided into several
les for readability through the BRAT interface.
The CepiDC corpus The CepiDC Corpus was provided by the French
institute for health and medical research (INSERM) for the task of ICD10 coding in
CLEF eHealth 2016 (task2). It consists of free text death certi cates collected
from physicians and hospitals in France over the period of 2006{2013.
        </p>
        <p>Table 2 presents statistics for the speci c sets provided to participants. The
training set covered the 2006{2012 period, and the test set covered the 2013
period. This time-oriented construction of the datasets re ects the practical
use case of coding death certi cates, where historical data is available to train
systems that can then be applied to current data to assist with new document
curation.</p>
        <p>
          CepiDC Dataset excerpts Death certi cates are standardized documents
lled by physicians to report the death of a patient. The content of the
medical information reported in a death certi cate and subsequent coding for public
health statistics follows complex rules described in a document that was supplied
to participants [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Table 3 presents an excerpt of the CepiDC corpus that
illustrates the heterogeneity of the data that participants had to deal with. While
some of the text lines were short and contained a term that could be directly
linked to a single ICD10 code (e.g., \Detresse respiratoire"), other lines could
be run-on (e.g., \Maladie de Parkinson ..."), contain non-diacritized text (e.g.,
\DENUTRITION" missing the diacritic on the \E"), a mix of cases and
diacritized text (\DEMENCE MIXTE EVOLUEE (stade severe)"), abbreviations
(e.g., \membre sup" instead of \membre superieur") and so on.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Named entity recognition (QUAERO Corpus). The task of named entity</title>
        <p>recognition consisted of analyzing plain text documents in order to mark the ten
types of entities of clinical interest de ned in the lab (Anatomy, Chemical and
Drugs, Devices, Disorders, Geographic Areas, Living Beings, Objects,
Phenomena, Physiology, Procedures). Participants could mark either plain entities (i.e.,
1 Cardio-respiratory arrest
2 Acute respiratory failure
3 Type 1 spinal muscular atrophy
4 Malnutrition dehydration
5 Advanced mixed dementia (late stage)
6 Idiopathic Parkinson's disease Recent angioedema of upper extremities w/o CT
exploration (no known drug cause)
mark the text mentions referring to an entity of interest) or normalized entities
(i.e., supply UMLS Concept Unique Identi ers corresponding to the entities in
addition to marking mentions).</p>
      </sec>
      <sec id="sec-2-4">
        <title>Entity normalization (QUAERO corpus). The task of entity normalization</title>
        <p>consisted of mapping entities of clinical interest marked in biomedical text to a
relevant UMLS CUI.</p>
        <p>ICD10 coding (CepiDC corpus). The task of coding consisted of mapping
sentences in the death certi cates to one or more relevant codes from the
International Classi cation of Diseases, tenth revision (ICD10).</p>
        <p>Replication. The replication task invited lab participants to submit a system
used to generate one or more of their submitted runs, along with instructions
to install and use the system. Then, two of the organizers independently worked
with the submitted material to replicate the results submitted by the teams as
their o cial runs.
2.3</p>
      </sec>
      <sec id="sec-2-5">
        <title>Evaluation metrics</title>
        <p>System performance was assessed by the usual metrics of information extraction:
precision (Formula 1), recall (Formula 2) and F-measure (Formula 3; speci cally,
we used =1.) for named entity recognition and entity normalization.</p>
        <p>Precision =</p>
        <p>true positives
true positives + false positives
Recall =</p>
        <p>true positives
true positives + false negatives
F-measure =
(1 +
2
2)</p>
        <p>
          precision recall
precision + recall
(1)
(2)
(3)
Performance measures were computed at the document level and micro-averaged
over the entire corpus. We determined system performance by comparing
participating system outputs against reference standard annotations on the test set.
For the QUAERO corpus, results were computed using the brateval program
initially developed by Verspoor et al. [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ], which we extended to cover the
evaluation of normalized entities. For the CepiDC corpus, results were computed
using a perl program. The evaluation tools were supplied to task participants
along with the training data.
        </p>
        <p>For plain entity recognition, an exact match (true positive) was counted
when the system's entity type and span matched the reference. A false positive
was counted if the system's entity type and span did not exactly match the
reference.</p>
        <p>For normalized entity recognition, an exact match (true positive) was
counted when the system's entity type, span and CUIs matched the reference.
Partial credits were given when only a subset of the expected CUIs were supplied
by the system for a given entity.</p>
        <p>For entity normalization, matches (true positives) were counted for each
CUI supplied with an entity. As a result, if either the system or the reference
supplied a list of CUIs associated with an entity, partial credit was awarded
if the reference and system lists were not identical but a subset of the lists
matched. However, system CUIs absent from the reference lists were counted as
false positives.</p>
        <p>For coding, matches (true positives) were counted for each ICD10 full code
supplied that matched the reference for the associated document line.</p>
        <p>The evaluation of the submissions to the replication task was essentially
qualitative: we used a scoring grid to record the ease of installing and running
the systems, the time spent to obtain results with the systems (analysts were
committed to spend at most one working day - or 8 hours - to work with each
system), and whether we managed to obtain the exact same results submitted
as o cial runs.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <p>Participating teams included between two and eight team members and resided
in France (teams ERIC-ECSTRA, LIMSI, LITL and SIBM), the Netherlands
(team Erasmus), Switzerland (Team BITEM) and Spain (Team UPF). Teams
often comprised members with a variety of backgrounds and drew from computer
science, informatics, statistics, information and library science, clinical practice.
It can be noted that one team (LITL) participated in the challenge as a
masterlevel class project.</p>
      <p>For the plain entity recognition task, ve teams submitted a total of 9 runs for
each of the corpora, EMEA and MEDLINE (18 runs in total). For the normalized
entity recognition task, three teams submitted a total of 5 runs for each of the
corpora (10 runs in total). For the normalization task, two teams submitted a
total of 3 runs for each of the corpora (6 runs in total). For the coding task, ve
teams submitted a total of 7 runs.</p>
      <p>Three systems were submitted, allowing us to attempt replicating a total of
seven runs.
3.1</p>
      <sec id="sec-3-1">
        <title>Methods implemented in the participants' systems</title>
        <p>Participants used a variety of methods, many of which relied on lexical sources
(medical terminologies and ontologies). Interestingly, some of these
knowledgebased methods relied on the training data supplied in the challenge as an
additional knowledge source. Some groups relied on statistical machine translation
to address the limitation of French coverage in the lexical sources available to
them. For each corpus, 3 teams out of 5 solely relied on knowledge-based sources,
and did not use machine learning for the speci c task of entity recognition and
normalization. The knowledge resources were used in combination with string
matching or indexing methods that were sometimes guided by linguistic
principles to identify entities and concepts in the challenge corpus.</p>
        <p>Machine-learning methods were still used by 2 teams out of 5 for each corpus.
They relied on Conditional Random Fields (CRFs), Latent Dirichlet Analysis
(LDA), Support Vector Machines (SVMs), and statistical information retrieval
models. They often used lexical resources as features.</p>
        <p>Participants who worked with the QUAERO and the CepiDC corpus did not
use the exact same systems to address both corpora.</p>
        <p>
          BITEM The BITEM team participated in the entity recognition and coding
tasks [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] using a di erent method for each task. Entity recognition in the
QUAERO corpus relied on a categorizer using the French UMLS to suggests
a ranked list of candidate entities potentially denoted by each text unit in the
corpus. Then, a second module anchored these candidates in the text, and
normalized the entities that could be anchored. For the coding task in the CepiDC
corpus, an ad hoc solution was developped based on pattern matching. This
method prioritizes exact matches that t the whole text. Failing that, the longest
match is then selected.
        </p>
        <p>
          ERIC-ECSTRA The ERIC-ECSTRA team participated in the coding subtask [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
Their rst run is based on the probabilistic topic model approach. It relies on
a supervised extension of the LDA model, called Labeled-LDA, that builds on
the latent topical structures to predict a category. The idea is that knowledge of
document topics can help predict the associated outputs (here the ICD10 codes).
Their second run is based on an SVM classi er with a bag-of-word data
representation. Their results suggest that Labeled-LDA and SVM both achieve
competitive results. It is interesting to note that one advantage of the LabeledLDA
method is that the classi er results are easier to understand for humans.
Erasmus MC The Erasmus MC team participated in the entity recognition and
the ICD-10 coding tasks [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ]. For both tasks a dictionary-based approach was
followed. For entity recognition and normalization, the system that had been
developed for the same task in the CLEF eHealth 2015 challenge [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], was tuned
on the 2016 training data. Brie y, a locally developed tagger, Peregrine, used a
dictionary consisting of French terminologies from the UMLS supplemented with
automatically translated English UMLS terms to index the QUAERO corpus.
Several post-processing steps were implemented to reduce the number of false
positive detections, including ltering based on precision scores that were derived
from the training data. For the coding task, two ICD-10 terminologies were
constructed based on the training material that was supplied by the challenge
organizers. The Solr text tagger was used with these terminologies to index the
death certi cates and generate codes. Again, precision-score ltering was applied
to improve precision.
        </p>
        <p>
          LITL The LITL team participated in the plain entity recognition task [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ].
The LITL team system was speci cally designed by master's students (LITL
programme, university of Toulouse) and their teachers for the challenge. The
system used is mainly based on supervised machine learning, through the use of
a CRF classi er (CRF++ 0.58) based on a varierty of linguistic features
(PartOf-Speech tags, generic word lists and syntactic parsing). Training and test data
have been POS-tagged and parsed by the Talismane toolkit [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ], and external
resources were used to tag the tokens (generic lists of su xes and pre xes, word
lists from SNOMED and from VIDAL database). The output of the CRF was
completed by a custom-made rule-based system which identi es syntactic
patterns in order to extract more complex entities.
        </p>
        <p>
          LIMSI The LIMSI team participated in the coding task [
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. Their system
offered a classi er with humanly-interpretable output, based on IR-style ranking
of candidate ICD10 diagnoses. A tf.idf-weighted bag-of-feature vector was built
for each training set code by merging all the statements found for this code in
the training data. Given a new statement, candidate codes were ranked with
Cosine similarity. Features included meta-information and n-grams of normalized
tokens. An ICD chapter classi er was also prepared with the same method and
it was used to rerank the top-k codes (k=2) returned by the code classi er. The
development phase focused on mono-code statements. Good precision could be
obtained using the top code and a signi cant performance gain was yielded by
chapter reranking. Accordingly, on test data, the system was set to return one
code for each statement, leaving multiple code assignment for future work.
SIBM The SIBM team participated in all tasks [
          <xref ref-type="bibr" rid="ref28">28</xref>
          ]. They approached entity
extraction from the provided QUAERO dataset as an indexing task relying on
multiple knowledge organization systems (KOS) partially or totally translated
into French. The extraction method, ECMT (Extracting Concepts with Multiple
Terminologies), performs bag of words concept matching at the sentence level.
It was originally designed to extract clinical concepts from Electronic Health
Records. They addressed the identi cation of relevant clinical entities within the
International Classi cation of Diseases version 10 in the CepiDC dataset with the
CIMIND system based on natural language processing and approximate string
matching methods.
        </p>
        <p>
          UPF The UPF team participated in the plain entity recognition and the
normalization tasks [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. They proposed two di erent systems for solving each phase.
For Phase I (entity recognition), a basic system uses a distant learning approach
based on a set SVM classi ers (one for each class) followed by a voting scheme for
choosing the best result. A second run was also submitted combining the result
of the basic system (run 1) with some symbolic processing for improving entity
classi cation. In Phase II (entity normalization), the system obtains
normalization information from public resources after obtaining the English translation of
each medical term.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>System performance on entity recognition</title>
        <p>Tables 4 and 5 present system performance on the plain entity recognition task.
Tables 6 and 7 present system performance on the normalized entity recognition
task. Team Erasmus had the best performance in terms of F-measure for both
the EMEA and MEDLINE corpora with their o cial runs. However, an uno cial
run (shown in italic font) submitted after the challenge deadline outperformed
the o cial runs by using the Solr indexing method instead of Peregrine. This
suggests that for knowledge-based methods, the speci c method used for
matching lexical resources carries a signi cant weight, in addition to the coverage of
these resources. Team LITL reports performing some corrective pre-processing
of the text to address extraneous spaces ocurring around punctuation marks,
which may cause issues with entity or concept recognition. However, they do not
report on the impact of the corrective step on their system performance.
Compared to last year, this year's performance show that all technical di culties
linked to the corpus format and annotation format seem to have been resolved.</p>
        <p>A t-test comparing all pairs of runs at entity level showed that all di erences
between runs were signi cant (p &lt; 0:001), with the exception of the two runs
from LILT (p = 0:28 on EMEA, exact match, p = 0:73 on MEDLINE, exact
match).</p>
      </sec>
      <sec id="sec-3-3">
        <title>System performance on entity normalization</title>
        <p>Tables 8 and 9 present system performance on the entity normalization task.
Team SIBM had the best performance in terms of F-measure for both the EMEA
and MEDLINE corpora, using a combination of knowledges resources dedicated
to French, compared to team UPF which relied on matching a translation of the
terms into English to English resources.</p>
      </sec>
      <sec id="sec-3-4">
        <title>System performance on death certi cate coding</title>
        <p>Table 10 presents system performance on the ICD10 coding task. Team Erasmus
had the best performance in terms of F-measure. Overall, systems performed
high on the coding task. It is interesting to note that participants addressed this
task independently from the entity recognition and normalization tasks o ered
on the QUAERO corpus. Since ICD10 is one of the terminologies aggregated
within the UMLS, a reasonable approach might have been to extract UMLS
concepts from the text of death certi cates, and then restrict the results to ICD10
in order to produce coding recommendations. However, none of the participating
teams chose this approach. The results show that both knowledge-based and
statistical methods can perform well on the task, as the best performance is
obtained from a knowledge-based method, while the second best is obtained
with statistical methods (Team ERIC-ECSTRA), followed by another knowledge
based method (team SIBM). The results are very encouraging from a practical
perspective and indicate that a coding assistance system could prove very useful
for the e ective processing of death certi cates.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Replication track and replicability of the results</title>
        <p>Three teams submitted systems to our replication track: one system covered
both QUAERO and CepiDC data, and two systems only processed CepiDC data.
Two teams expressed interest in submitting a system but eventually reported
that they did not have time to make the system ready for submission. One team
reported that they were reserving the distribution of their system to commercial
use and one team did not provide a reason for not participating to the track.</p>
        <p>The system submitted for replicating QUAERO results was in fact
incomplete as the submission included the results of pre-processing the corpus with a
tool that the team did not share as part of the replication track. Between the
two analysts working with each system, we were able to replicate exactly the
results submitted by 6 of the target runs (the QUAERO runs and two CepiDC
runs): the precision, recall and F-measure obtained from running the systems
were identical to that of the runs submitted by participants. For one run
adressing the CepiDC corpus, only one analyst was able obtain results from the system,
and the results obtained showed a 0.02 di erence in F-measure, which was
statistically signi cant. The analysts experienced varying degrees of di culty to
install and run the systems. Di erences were mainly due to the technical set-up
of the computers used to replicate the experiments. Analysts also report that
additional information on system requirements, installation procedure and
practical use would be useful for all the systems submitted. Overall, this indicates
that replication is achievable. However, it is not as straight-forward as one would
hope. More detailed communication about the systems could be an important
step towards making replication an e ortless reality.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>We released a new portion of the QUAERO French Medical corpus through
Task 2 of the CLEFeHealth 2016 Evaluation Lab. This corpus contains entity
annotations for ten entities of clinical interest, with normalization to UMLS CUIs.
In the evaluation lab, we evaluated systems on the task of plain or normalized
entity recognition as well as on the task of assigning CUIs to pre-identi ed
entities (normalization). In addition, we also released a large corpus of French death
certi cates to evaluate systems on the task of ICD10 coding. This is the second
edition of a biomedical NLP challenge that provides large gold-standard
annotated corpora in French. Results show that high performance can be achieved
by NLP systems on the tasks of entity recognition, normalization and coding
for French biomedical text. The corpus used and the participating team system
results are an important contribution to the research community and the focus
on a language other than English (French) remains a rare initiative.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>We want to thank all participating teams for their e ort in addressing new and
challenging tasks. We also want to thank Jan Kors from team Erasmus for his
contribution to the CepiDC evaluation script. The organization work for CLEF
eHealth 2016 task 2 was supported by the Agence Nationale pour la Recherche
(French National Research Agency) under grant number
ANR-13-JCJC-SIMI2CABeRneT.</p>
      <p>The CLEF eHealth 2016 evaluation lab has been supported in part by (in
alphabetical order) PhysioNetWorks Workspaces; the CLEF Initiative;</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Jones</surname>
            <given-names>KS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Galliers</surname>
            <given-names>JR</given-names>
          </string-name>
          .
          <source>Evaluating natural language processing systems: An analysis and review</source>
          .
          <source>1995</source>
          . Springer Science &amp; Business Media:
          <volume>1083</volume>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Voorhees</surname>
            <given-names>EM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harman</surname>
            <given-names>DK</given-names>
          </string-name>
          <article-title>and others</article-title>
          . TREC:
          <article-title>Experiment and evaluation in information retrieval</article-title>
          , vol
          <volume>1</volume>
          .
          <year>2005</year>
          . MIT press Cambridge.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Suominen</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salantera</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            <given-names>WW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elhadad</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pradhan</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>South</surname>
            <given-names>BR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mowery</surname>
            <given-names>DL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            <given-names>GJF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leveling</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Overview of the ShARe/CLEF eHealth Evaluation Lab 2013</article-title>
          . In: Forner P, Muller H,
          <string-name>
            <surname>Paredes</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            <given-names>P</given-names>
          </string-name>
          , Stein B (eds),
          <source>Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Visualization.
          <source>LNCS</source>
          (vol.
          <volume>8138</volume>
          ):
          <fpage>212</fpage>
          -
          <lpage>231</lpage>
          . Springer,
          <year>2013</year>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Goeuriot</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            <given-names>G</given-names>
          </string-name>
          .
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2015</article-title>
          . In: Information Access Evaluation. Multilinguality, Multimodality, and Interaction. Springer,
          <year>2015</year>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Kelly</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schreck</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leroy</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mowery</surname>
            <given-names>DL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velupillai</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            <given-names>WW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Overview of the ShARe/CLEF eHealth Evaluation Lab 2014</article-title>
          . In:
          <string-name>
            <surname>Kanoulas</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lupu</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            <given-names>A</given-names>
          </string-name>
          , Toms E (eds),
          <source>Information Access Evaluation</source>
          . Multilinguality, Multimodality, and Interaction.
          <source>LNCS</source>
          (vol.
          <volume>8685</volume>
          ):
          <fpage>172</fpage>
          -
          <lpage>191</lpage>
          . Springer,
          <year>2014</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Neveol</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tannier</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamon</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            <given-names>P</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>CLEF eHealth Evaluation Lab 2015 Task 1b: clinical named entity recognition</article-title>
          .
          <source>CLEF</source>
          <year>2015</year>
          , Online Working Notes,
          <source>CEUR-WS 1391.</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Chapman</surname>
            <given-names>WW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nadkarni</surname>
            <given-names>PM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschman</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>D'Avolio</surname>
            <given-names>LW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            <given-names>GK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uzuner</surname>
            <given-names>O</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Overcoming barriers to NLP for clinical text: the role of shared tasks and the need for additional creative solutions</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          ,
          <volume>18</volume>
          (
          <issue>5</issue>
          ):
          <fpage>540</fpage>
          -
          <lpage>3</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Huang</surname>
            <given-names>CC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            <given-names>Z</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Community challenges in biomedical text mining over 10 years: success, failure and the future</article-title>
          .
          <source>Brief Bioinform</source>
          ,
          <year>2015</year>
          <article-title>May 1</article-title>
          . pii: bbv024.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Yeh</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morgan</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colosimo</surname>
            <given-names>M</given-names>
          </string-name>
          , Hirschman L.
          <article-title>BioCreAtIvE task 1A: gene mention nding evaluation</article-title>
          .
          <source>BMC Bioinformatics</source>
          .
          <year>2005</year>
          ;
          <volume>6</volume>
          <issue>Suppl 1</issue>
          :
          <fpage>S2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Smith</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tanabe</surname>
            <given-names>LK</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ando</surname>
            <given-names>RJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuo</surname>
            <given-names>CJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chung</surname>
            <given-names>IF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            <given-names>CN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            <given-names>YS</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klinger</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Friedrich</surname>
            <given-names>CM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganchev</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torii</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haddow</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Struble</surname>
            <given-names>CA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Povinelli</surname>
            <given-names>RJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vlachos</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumgartner</surname>
            <given-names>WA</given-names>
          </string-name>
          Jr,
          <string-name>
            <surname>Hunter</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carpenter</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsai</surname>
            <given-names>RT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            <given-names>HJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katrenko</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Adriaans</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Blaschke</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torres</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neves</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakov</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Divoli</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Man~</surname>
          </string-name>
          a
          <string-name>
            <surname>-Lopez</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mata</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilbur</surname>
            <given-names>WJ</given-names>
          </string-name>
          .
          <article-title>Overview of BioCreative II gene mention recognition</article-title>
          .
          <source>Genome Biol</source>
          .
          <year>2008</year>
          ;
          <volume>9</volume>
          <issue>Suppl 2</issue>
          :
          <fpage>S2</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Arighi</surname>
            <given-names>CN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
            <given-names>CH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            <given-names>KB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschman</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krallinger</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valencia</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilbur</surname>
            <given-names>JW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegers</surname>
            <given-names>TC</given-names>
          </string-name>
          .
          <article-title>BioCreative-IV virtual issue</article-title>
          .
          <source>Database (Oxford)</source>
          .
          <source>2014 May</source>
          <volume>22</volume>
          ;
          <year>2014</year>
          . pii: bau039.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Uzuner</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>South</surname>
            <given-names>BR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <source>DuVall SL</source>
          .
          <year>2010</year>
          i2b2/
          <article-title>VA challenge on concepts, assertions, and relations in clinical text</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          .
          <year>2011</year>
          SepOct;
          <volume>18</volume>
          (
          <issue>5</issue>
          ):
          <fpage>552</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hirschman</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colosimo</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morgan</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeh</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Overview of BioCreAtIvE task 1B: normalized gene lists</article-title>
          .
          <source>BMC Bioinformatics</source>
          .
          <year>2005</year>
          ;
          <volume>6</volume>
          <issue>Suppl 1</issue>
          :
          <fpage>S11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. ,
          <string-name>
            <surname>Morgan</surname>
            <given-names>AA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            <given-names>AM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fluck</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruch</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Divoli</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fundel</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leaman</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hakenberg</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>HH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torres</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krauthammer</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lau</surname>
            <given-names>WW</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            <given-names>CN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schuemie</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            <given-names>KB</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hirschman</surname>
            <given-names>L</given-names>
          </string-name>
          .
          <article-title>Overview of BioCreative II gene normalization</article-title>
          .
          <source>Genome Biol</source>
          .
          <year>2008</year>
          ;
          <volume>9</volume>
          <issue>Suppl 2</issue>
          :
          <fpage>S3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lu</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kao</surname>
            <given-names>HY</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            <given-names>CH</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuo</surname>
            <given-names>CJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsu</surname>
            <given-names>CN</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsai</surname>
            <given-names>RT</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            <given-names>HJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okazaki</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cho</surname>
            <given-names>HC</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gerner</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solt</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agarwal</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vishnyakova</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruch</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romacker</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rinaldi</surname>
            <given-names>F</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhattacharya</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivasan</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torii</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matos</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campos</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verspoor</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Livingston</surname>
            <given-names>KM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilbur</surname>
            <given-names>WJ</given-names>
          </string-name>
          .
          <article-title>The gene normalization task in BioCreative III</article-title>
          .
          <source>BMC Bioinformatics</source>
          .
          <source>2011 Oct</source>
          <volume>3</volume>
          ;
          <issue>12 Suppl 8</issue>
          :
          <fpage>S2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Uzuner</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>South</surname>
            <given-names>BR</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <source>DuVall SL</source>
          .
          <year>2010</year>
          i2b2/
          <article-title>VA challenge on concepts, assertions, and relations in clinical text</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          .
          <year>2011</year>
          SepOct;
          <volume>18</volume>
          (
          <issue>5</issue>
          ):
          <fpage>552</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Neveol</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            <given-names>SJ</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            <given-names>P</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Language Resources for French in the Biomedical Domain</article-title>
          .
          <source>In: Proc of LREC</source>
          , p.
          <fpage>2146</fpage>
          -
          <lpage>2151</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Neveol</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leixa</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosset</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            <given-names>P</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>The QUAERO French Medical Corpus: A Ressource for Medical Entity Recognition and Normalization</article-title>
          .
          <source>In: Proc of Bio TextM</source>
          , p.
          <fpage>24</fpage>
          -
          <lpage>30</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Pavillon</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laurent</surname>
            <given-names>F</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Certi cation et codi cation des causes medicales de deces</article-title>
          .
          <source>Bulletin Epidemiologique</source>
          Hebdomadaire - BEH:
          <fpage>134</fpage>
          -
          <lpage>138</lpage>
          . http://opac. invs.sante.fr/doc_num.php?explnum_id=2065 (accessed:
          <fpage>2016</fpage>
          -06-06)
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Verspoor</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jimeno Yepes</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavedon</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McIntosh</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herten-Crabb</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plazzer</surname>
            <given-names>JP</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Annotating the Biomedical Literature for the Human Variome. Database (Oxford), virtual issue for BioCuration 2013 meeting</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Mottin</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gobeill</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mottaz</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasche</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaudinat</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruch</surname>
            <given-names>P</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <source>BiTeM at CLEF eHealth Evaluation Lab 2016 Task</source>
          <volume>2</volume>
          :
          <string-name>
            <given-names>Multilingual</given-names>
            <surname>Information Extraction CLEF 2016 Online Working</surname>
          </string-name>
          <article-title>Notes</article-title>
          . CEUR-WS
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Dermouche</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Looten</surname>
            <given-names>V</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Flicoteaux</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chevret</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Velcin</surname>
            <given-names>J</given-names>
          </string-name>
          and
          <string-name>
            <surname>Taright N</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>ECSTRA-INSERM @ CLEF eHealth2016-task 2: ICD10 Code Extraction from Death Certi cates</article-title>
          .
          <source>CLEF 2016 Online Working Notes. CEUR-WS</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Van Mulligen</surname>
            <given-names>E</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Afzal</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akhondi</surname>
            <given-names>SA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vo</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kors</surname>
            <given-names>JA</given-names>
          </string-name>
          <article-title>(</article-title>
          <year>2016</year>
          ).
          <source>Erasmus MC at CLEF eHealth</source>
          <year>2016</year>
          :
          <article-title>Concept Recognition and Coding in French Texts</article-title>
          .
          <source>CLEF 2016 Online Working Notes</source>
          , CEUR-WS
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Afzal</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akhondi</surname>
            <given-names>SA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>van Haagen</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Mulligen</surname>
            <given-names>E</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kors</surname>
            <given-names>JA</given-names>
          </string-name>
          <article-title>(</article-title>
          <year>2015</year>
          ).
          <article-title>Biomedical Concept Recognition in French Text Using Automatic Translation of English Terms</article-title>
          .
          <source>CLEF 2015 Online Working Notes. CEUR-WS</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Ho-Dac</surname>
            <given-names>LM</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tanguy</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grauby</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hnub</surname>
            <given-names>N</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heu Mby</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Malosse</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riviere</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veltz-Mauclair</surname>
            <given-names>A</given-names>
          </string-name>
          and
          <string-name>
            <surname>Wauquier</surname>
            <given-names>M</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>LITL at CLEF eHealth2016: recognizing entities in French biomedical documents</article-title>
          .
          <source>CLEF 2016 Online Working Notes. CEUR-WS</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Urieli</surname>
            <given-names>A</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Robust French syntax analysis: reconciling statistical methods and linguistic knowledge in the Talismane toolkit</article-title>
          .
          <source>PhD thesis</source>
          . Universite de Toulouse II-Le Mirail
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Zweigenbaum</surname>
            <given-names>P</given-names>
          </string-name>
          and
          <string-name>
            <surname>Lavergne</surname>
            <given-names>T</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>LIMSI ICD10 coding experiments on CepiDC death certi cate statements</article-title>
          .
          <source>CLEF 2016 Online Working Notes</source>
          . CEURWS
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Cabot</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soualmia</surname>
            <given-names>LF</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dahamna</surname>
            <given-names>B</given-names>
          </string-name>
          and
          <string-name>
            <surname>Darmoni SJ</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <source>SIBM at CLEF eHealth Evaluation Lab</source>
          <year>2016</year>
          :
          <article-title>Extracting Concepts in French Medical Texts with ECMT and CIMIND</article-title>
          .
          <source>CLEF 2016 Online Working Notes. CEUR-WS</source>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Vivaldi</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodriguez</surname>
            <given-names>H</given-names>
          </string-name>
          and
          <string-name>
            <surname>Cotik</surname>
            <given-names>V</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Semantic tagging and normalization of French medical entities</article-title>
          .
          <source>CLEF 2016 Online Working Notes. CEUR-WS</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>