<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ICD-10 coding of death certi cates with the NCBO and SIFR Annotators at CLEF eHealth 2017</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andon Tchechmedjiev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amine Abdaoui</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincent Emonet</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Clement Jonquet</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>(1) Laboratory of Informatics, Robotics and Microelectronics of Montpellier (LIRMM) University of Montpellier &amp; CNRS, France (2) Center for BioMedical Informatics Research (BMIR) Stanford University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The SIFR BioPortal is an open platform to host French biomedical ontologies and terminologies based on the technology developed by the US National Center for Biomedical Ontology (NCBO). The portal facilitates the use and fostering of terminologies and ontologies by o ering a set of services including semantic annotation. The SIFR Annotator (http://bioportal.lirmm.fr/annotator) is a publicly accessible, easily usable ontology-based annotation tool to process French text data and facilitate semantic indexing. The web service relies on the ontology content (preferred labels and synonyms) as well as on the semantics of the ontologies (is-a hierarchies) and their mappings. The SIFR BioPortal also o ers the possibility of querying the original NCBO Annotator for English text via a dedicated proxy that extends the original functionality. In this paper, we present a preliminary performance evaluation of the generic annotation web service (i.e., not speci cally customized) for coding death certi cates i.e., annotating with ICD-10 codes. This evaluation is performed against the CepiDC/CDC CLEF eHealth 2017 task 1 manually annotated corpus. For this purpose, we have built custom SKOS vocabularies from the CePIDC/CDC dictionaries as well as training and development corpora, for all three tasks using a most frequent code heuristic to assign ambiguous labels. We then submitted the vocabularies to the NCBO and SIFR BioPortal and ran the annotation services on the task 1 datasets. We obtained, for our best runs on each corpus the following results: English raw corpus (69.08% P, 51.37% R, 58,92% F1); French raw corpus (54.11% P, 48.00% R, 50,87% F1); French aligned corpus (50.63% P, 52.97% R, 51.77% F1).</p>
      </abstract>
      <kwd-group>
        <kwd>Semantic annotation</kwd>
        <kwd>SIFR Annotator</kwd>
        <kwd>NCBO Annotator</kwd>
        <kwd>ICD-10 coding</kwd>
        <kwd>Biomedical ontologies</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Biomedical data integration and semantic interoperability are necessary to enable
translational research. The biomedical community has turned to ontologies and
terminologies to describe their data and turn them into structured and formalized
knowledge [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ]. Ontologies help to address the data integration problem by
playing the role of common denominator. One way of using ontologies is by means
of creating semantic annotations. An annotation is a link from an ontology term
to a data element, indicating that the data element (e.g., article, experiment,
clinical trial, medical record) refers to the term [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In ontology-based indexing, we
use these annotations to \bring together" the data elements from the resources.
      </p>
      <p>
        The community has turned toward ontologies to design semantic indexes of
data that leverage the medical knowledge for better information mining and
retrieval. Despite a large adoption of English in science, a signi cant quantity
of biomedical data uses the French language. Besides the existence of various
English tools, there are considerably less terminologies and ontologies available
in French [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and there is a strong lack of related tools and services to exploit
them. This lack does not match the huge amount of biomedical data produced in
French, especially in the clinical world (e.g., electronic health records).
      </p>
      <p>
        In the context of the Semantic Indexing of French Biomedical Data Resources
(SIFR) project, we have developed the SIFR BioPortal (http://bioportal.
lirmm.fr) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], an open platform to host French biomedical ontologies and
terminologies based on the technology developed by the US National Center for
Biomedical Ontology [
        <xref ref-type="bibr" rid="ref6 ref7">6,7</xref>
        ]. The portal facilitates the use and fostering of
ontologies by o ering a set of services such as search and browsing, mapping hosting
and generation, metadata edition, versioning, visualization, recommendation,
community feedback, etc. As of today, the portal contains 24 public ontologies
and terminologies (+ 6 private ones) that cover multiple areas of biomedicine,
such as the French versions of MeSH, MedDRA, ATC, ICD-10, or WHO-ART but
also multilingual ontologies (for which only the French content is parsed) such as
Rare Human Disease Ontology, OntoPneumo or Ontology of Nuclear Toxicity.
      </p>
      <p>
        The SIFR BioPortal includes the SIFR Annotator1 a publicly accessible and
easily usable ontology-based annotation tool to process text data in French. This
service is originally based on the NCBO Annotator [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], a web service allowing
scientists to utilize available biomedical ontologies for annotating their datasets
automatically, but was signi cantly enhanced and customized for French. The
annotator service processes raw textual descriptions, tags them with relevant
biomedical ontology concepts and returns the annotations to the users in several
formats such as JSON-LD, RDF or BRAT. A preliminary evaluation [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] showed
that the web service matches the results of previously reported work in French,
while being public, functional and turned toward semantic web standards. However,
this evaluation precedes the CLEF eHealth French task series and was not
satisfactory. We had the motivation of evaluating the annotation service on the
Quaero and CepiDC corpora used in the CLEF eHealth 2015-2017 French text
data annotation tasks.
      </p>
      <p>The SIFR BioPortal also o ers the possibility of querying the original NCBO
Annotator for English text via a dedicated proxy that extends the original</p>
    </sec>
    <sec id="sec-2">
      <title>1 http://bioportal.lirmm.fr/annotator</title>
      <p>
        functionality. Thus, in this case, the +600 ontologies of the NCBO BioPortal2
may be used. This service, called the NCBO Annotator+3, is querying the original
NCBO Annotator while o ering new functionalities by pre processing of the
input text and/or post processing of the original results. For instance, we have
implemented the scoring of the results for both the SIFR and NCBO Annotator+
thanks to that proxy architecture [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Despite its wide and various uses and
multiple evaluations, the NCBO Annotator has never been evaluated in the
context of the CLEF eHealth tasks, and we believed it would be appropriate and
relevant for the community to o er such an evaluation.
      </p>
      <p>
        In this paper, we present our participation to the task 1 [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] of the CLEF
eHealth 2017 challenge[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which tackles the problem of information extraction
(diagnostic coding) in written clinical texts (death certi cates). The objective
of the task is to annotate each line of several death certi cates, provided by the
French Centre d'epidemiologie sur les causes medicales deces,(CepiDC) with an
International Classi cation of Diseases, 10th revision (ICD-10) diagnostic code
(French aligned task) or to annotate each document with the set of relevant ICD-10
diagnostic codes (French raw and English raw tasks). Considering that ICD-10 was
never conceived to be used by automatic lexical tools, annotating the CepiDC data
using only ICD-10 as source dictionary would have o ered poor results, therefore,
we have built custom SKOS vocabularies from the CepiDC/CDC dictionaries
as well as the development and training corpora provided in the CLEF eHealth
2017 task 1 datasets. In the following, we will describe the construction of these
custom vocabularies and present the results obtained both by the SIFR and
NCBO Annotators used without any speci c customization for the CLEF eHealth
2017 task 1. We obtained, for our best runs on each corpus, the following results:
French Aligned (50.63% P, 52.97% R, 51.77% F1); French Raw (54.11% P, 48.00%
R, 50,87% F1); English raw (69.08% P, 51.37% R, 58,92% F1). We will discuss
the advantages and limitations of the annotators and possible perspectives for
enhancing the performance of the speci c task of coding death certi cates or
clinical notes. To us, in addition to technical performance (i.e., precision and
recall) there are other aspects of the services that we think are crucial if we want
to make the use of ontologies for annotation of clinical data mainstream. For
instance, interoperability, ease of use as a service, openness and the adoption of
the semantic web. A good annotation service can be used in research or clinical
environments, without any explicit knowledge of the technologies, the ontologies
or the natural languages processing techniques involved.
2
2.1
      </p>
      <sec id="sec-2-1">
        <title>Materials and methods</title>
        <sec id="sec-2-1-1">
          <title>SIFR BioPortal and SIFR Annotator</title>
          <p>The SIFR Annotator work ow is composed of several steps: dictionary creation
from ontologies, text pre-processing, concept recognition, semantic expansion</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2 http://bioportal.bioontology.org</title>
    </sec>
    <sec id="sec-4">
      <title>3 http://bioportal.lirmm.fr/ncbo_annotatorplus</title>
      <p>
        (with mappings and hierarchy), annotation post-processing. For instance, in the
nal step, annotations are scored with relation to the context from which they
have been generated, which is a requirement when they are used to index the
original data. The SIFR Annotator can also recognize negation, experiencer and
temporality based on a customized French implementation of the NegEx/Context
technique [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Only the concept recognition step is evaluated in the context of
the CLEF eHealth 2017 task 1, therefore, we will not describe into more detail
the rest of the SIFR Annotator work ow here. For a presentation of the original
NCBO Annotator service we point the readers to [
        <xref ref-type="bibr" rid="ref13 ref8">8,13</xref>
        ]. SIFR Annotator (Figure
1) allows users to input free text and to annotate the text with ontology concepts.
SIFR Annotator, uses a dictionary composed of a at list of terms build the
concept labels and synonym labels from all the resources uploaded in SIFR
Bioportal (ontologies, terminologies, vocabularies, dictionaries). SIFR BioPortal
currently contains about 255K concepts and around twice that number of terms.
Enabling the service to use additional ontologies is as simple as uploading them
to the portal (the indexing and dictionary generation are automatic).
      </p>
      <p>
        Depending on the type of biomedical text, the annotator allows users to
annotate with only a subset of the ontologies available. The annotator is based
on the Mgrep [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] concept recognizer. Mgrep and/or the NCBO Annotator have
been evaluated [
        <xref ref-type="bibr" rid="ref13 ref15 ref16 ref17 ref18">15,16,13,17,18</xref>
        ] on di erent English-language datasets and usually
perform very well in terms of precision e.g., 95% in recognizing disease names [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
A comparative evaluation of MetaMap [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] and Mgrep within NCBO Annotator
also exists [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. However, there are no evaluations on French text. Mgrep uses
no natural language processing techniques for the recognition, but o ers a fast
and reliable (precision) matching that enables its use in real-time high load
web-services. One therefore relies on the creation of the dictionary to augment
the recall by adding alternate syntactic forms.
      </p>
      <p>An important aspect for the SIFR Annotator is to be available as a web service.
The service results may be described in multiple syntaxes (XML or JSON-LD)
and format (e.g., RDF/XML described with the Annotation Ontology or BRAT).
A speci c CLEF eHealth output format has also been implemented to evaluate
the service against previous campaigns (Quaero corpus). Akthough there is a web
interface (Figure 1), the service is meant to be used through the REST application
programming interface (API). We have created a Docker (www.docker.com)
packaging that allows for an easy local installation to allow for the processing
of sensitive in-house data, a common requirement when manipulating clinical
data. All the code is open source and available on GitHub (https://github.
com/sifrproject).</p>
      <p>
        Within the SIFR project, we also developed an enhanced version of the
NCBO Annotator to annotate English biomedical text data, without having to
serve English ontologies locally. The NCBO Annotator+ uses a proxy service
architecture that enhances the capabilities of the original annotation service by
encapsulating around the original application programming interface. All the
extension implemented for the French annotator are thus automatically also
available for English e.g., the scoring of annotations [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] or more recently the
detection of negation.
      </p>
      <p>For CLEF eHealth 2017 task 1, we have used the SIFR and NCBO Annotators
(the software implementations) \as it is" without any speci c customization for
the task, alough we used speci cally tailored dictionaries. For all the runs, the
longest match only parameter was enabled, and we used no semantic expansion
of the annotations, scoring or contextualization.
2.2</p>
      <sec id="sec-4-1">
        <title>Task and corpus</title>
        <p>The objective of CLEF eHealth 2017 task 1 is to annotate death certi cates with
ICD-10 codes both in French and in American English. For English, a corpus
of death certi cates from the CDC was provided, split in a training and a test
corpus. The training corpus contains 13,329 death certi cates, for a total of 32,714
lines. The test corpus contains 6,665 certi cates containing a total of 14,834 lines.
For French, a corpus of death certi cates from CepiDC was provided: a training
corpus of 65,844 documents and 195,204 lines, a development corpus of 27,851
document and 80,900 lines and a test corpus of 31,683 documents and 91,954
lines. The corpora are digitized versions of actual death certi cates lled in by
clinicians. Although the punctuation is not always correct or present, the corpus
is already segmented in lines (as per the standard international death certi cate
model) which for the most part only contain single sentences.</p>
        <p>The French corpus was provided in both an aligned and a raw format, while
the English corpus was only provided in the raw format. The raw format provides
two les, a CausesBrutes le and an Ident le. The former contains semicolon
separated values for the Document identi er (DocID), the year the certi cated
was coded (YearCoded), the line identi er (LineID), the raw text as it appears in
the certi cate (RawText), an interval type during which the condition occurred
(IntType - seconds, minutes, hours, weeks, years) and an interval value (IntValue).
The Ident le contains a document identi er, the year the certi cate was coded,
the gender of the person, the code for the primary cause of death, the age
and the location of death. The aligned format is a reconciliation between the
CausesBrutes and Ident les, where the elds have been aligned at the document
and line number level. Thus, the aligned le contains the same unique elds that
the original ones to which an extra eld is added in the gold standard dataset
providing a standardized text that represents the manually annotated code.</p>
        <p>For CLEF eHealth 2017 task 1, we have used only the \RawText"
information of both the aligned and raw datasets. We did not use any other
information/features such as age or gender contained in the les.
2.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>Dictionaries construction</title>
        <p>
          SIFR BioPortal already contained the French ICD-104 (CIM-10) reference
terminology. This OWL version was originally produced by the CISMeF team from
an automatic export from the HeTOP ontology/terminology server [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
Respectively, the NCBO BioPortal already contained the English ICD-10.5 This RDF
version was automatically exported from the Uni ed Medical Language System
(UMLS) with the umls2rdf tool. 6 However, the purpose of ICD-10 is to serve
as a general purpose reference to code medical acts, and not to be directly used
for text annotation and, especially not in a particular clinical task such as death
certi cate coding. Indeed, from our experiments, using ICD-10 classi cation alone
for annotation leads to a F1 score below 10%.
        </p>
        <p>For the French tasks, a set of dictionaries was provided by CepiDC that give
a standardized description text of each of the codes that appear in the corpora.
Additionally, the data from the aligned corpus (French only) could also be used
to enrich the lexical terms of ICD-10. A similar dictionary was provided for the
English task. In order to use these dictionaries within the SIFR and NCBO
Annotator, we had to encode them using a format accepted as input within
the portal, which includes RDFS, OWL, SKOS, OBO or RRF (UMLS format).
In this case, the ideal choice in terms of standardization, potential reusability
and simplicity was to use SKOS (Simple Knowledge Organization System) a
W3C Recommendation specialized for vocabularies and thesaurus. For CLEF
eHealth 2017 task 1, we produced two groups of SKOS dictionaries: CIM-10DC*
for French, based on the French dictionaries and aligned corpus; ICD-10-CDC*
for English based on the CDC corpus dictionary alone.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4 http://bioportal.lirmm.fr/ontologies/CIM-10</title>
    </sec>
    <sec id="sec-6">
      <title>5 https://bioportal.bioontology.org/ontologies/ICD-10</title>
    </sec>
    <sec id="sec-7">
      <title>6 https://github.com/ncbo/umls2rdf</title>
      <p>
        We set out in this construction process by rst de ning the appropriate
schema to represent the SKOS dictionaries. We chose to use the same URIs as
concepts identi ers for skos:Concept that the owl:Class in the available
CIM10/ICD-10 which allows our dictionaries to be fully aligned ontologically speaking
with the original terminologies they enrich. Each of the codes was represented by
a skos:Concept. The URIs are composed of a base URI and a class identi er
that represents the ICD codes, in the following format: "[A-Z][
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">0-9</xref>
        ][
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">0-9</xref>
        ] ?_[
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">0-9</xref>
        ]?"
(e.g. G12.1 or A10). The identi er is slightly di erent from the codes from the
task dictionaries: there is a dot before the last digit and if the last digit is zero,
then the dot and the last zero are omitted. Thus, G12.1 in CIM-10 corresponds to
G121 in the corpus while A10 in CIM-10 corresponds to A100 in the corpus. The
corresponding URI in CIM-10 and this in CIM-10DC are: http://chu-rouen.fr/
cismef/CIM-10#G12.1, where http://chu-rouen.fr/cismef/CIM-10# is the
base URI and G12.1 the code identi er. In ICD-10 and ICD-10CDC the URIs are
like so: http://purl.bioontology.org/ontology/ICD-10/P08.0, where the
base URI is http://purl.bioontology.org/ontology/ICD-10/ and the code
identi er is P080. When building the SKOS dictionaries, we used the same chapter
hierarchy as ICD-10 for the sake of convenient browsing and visualization of the
dictionaries in the NCBO and SIFR BioPortals.
      </p>
      <p>Construction algorithm We built the French SKOS dictionary from the
aligned corpus and all the CepiDC dictionaries. We built the English one only
from the raw corpus and the CDC dictionary. We rst built a code index, that
to each code associated the list of labels retrieved from: the DiagnosisText eld
in the dictionary, associated to codes through the ICD1 and ICD2 elds 7; the
RawText and StandardText (only for French) elds from the corpus associated
to codes through the ICD-10 eld in the corpus 8.</p>
      <p>For each code concept the CepiDC and CDC dictionaries contained multiple
labels. In order to follow SKOS speci cation, we had to select a preferred name
automatically (skos:prefLabel) and assign the other labels as alternative labels
(skos:altLabel). Note that this selection would note in uence the annotation
process as both preferred name and synonyms are included in the concept
recognizer dictionaries. The selection heuristic took the shortest label that does not
contain three or more consecutive capital letter (likely and acronym).</p>
      <sec id="sec-7-1">
        <title>Ambiguous label selection heuristics An important issue when building the</title>
        <p>SKOS dictionaries was to assign ambiguous labels (i.e., identical labels which
correspond to di erent codes). Indeed, those labels create ambiguity in the
annotations and leads to better recall at the price of a low precision. For example,
the label "choc septique" was present as preferred label or synonyms for 58
di erent codes. Therefore, we had to implement a selection heuristic to determine
the most suitable code to which the label should be bound.
7 For French, we used a concatenation of all the dictionaries and for English we used
the one dictionary le provided.
8 AlignedCauses 2013Full for French, CausesCalculees EN training for English.</p>
        <p>When using both the standard text and the raw text elds from the corpus,
if standard text label is ambiguous, a simple heuristic is to not add it to any
code but just use the raw text instead. Given that the raw text is unique, the
ambiguity related to the inclusion of the test corpus is removed. We called this
rst strategy "Adaptive dictionary generation" and created a CIM-10DCA French
SKOS dictionary to evaluate it. Given that this strategy relied on the availability
of a standardized text, it was con ned to the French aligned task, as the raw
English corpus contained no standard labels.</p>
        <p>The drawback with the previous heuristic is that we lose some labels that
would otherwise have potentially increased recall. Thus, we searched a way of
assigning ambiguous labels to one code only. Taking inspiration from the idea of
the most frequent sense baseline often used in Word Sense Disambiguation tasks,
we adopted a heuristic that assigns ambiguous labels to the most frequent code
only. We use the training corpus to estimate the frequencies of use of the codes
(gold standard annotations) so that when a label can belong to several codes, we
can sort the codes by frequency and chose either the most frequent code (MFC)
or the top k most frequent codes (kMFC).</p>
        <p>In practice the "Adaptive" strategy led to a much lower recall without
particularly improving precision. Final F1 scores were worse than with the MFC strategy
which led to a precision and recall that were balanced. This is the strategy we
have nally used in the reported results.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Availability of the SKOS dictionaries The nal SKOS dictionaries built</title>
        <p>with the best ambiguous label strategy have been uploaded respectively on the
SIFR and NCBO BioPortal, and are accessible in private mode, only for the
replication track of the task. CIM-10DC-ALL contains 6817 concepts for a total
of 295,385 labels. CIM-10DC-ALLMFC contains the same number of concepts
but only 249,524 labels.ICD-10CDC contains 3738 concepts with 166,500 labels.</p>
        <p>We have also created a resource solely from the CepiDC dictionaries (without
using the corpus) called CIM-10DCD, that does not contain any sentences
originally present in death certi cates. We are currently discussing with the CepiDC,
for a potential public release of the SKOS dictionaries as well as the Work ow to
update them on a regular basis. Indeed, we believe it is also part of the SIFR
project mission to facilitate open access to resources (and adopting standard
ways of describing them e.g., semantic web standards), when licensing permits.</p>
      </sec>
      <sec id="sec-7-3">
        <title>2.4 ICD-10/CIM-10 Coding with ontology concepts</title>
        <p>Given that we used the SIFR and NCBO Annotators, besides manually curating
the created SKOS dictionaries, the nal step to obtaining a working system for
the task was to write a complete work ow to:9
1. Read the corpus in the raw or aligned formats;
9 For this purpose, we used the Java language.
2. Send the text to the Annotators REST API with the right ontologies and
annotation parameters and retrieve the annotations produced by the
Annotators;
3. Optionally prune some annotations (post-annotation heuristic);
4. Produce the output in the right raw or aligned format.</p>
        <p>We implemented two post-annotation heuristics. Most Frequent Code, where
if a particular line was annotated with several codes, only keep the most frequent
code based on the code distribution of the training corpus. Code Frequency Cuto ,
we calculate a normalized probability distribution of the codes that annotate a
particular line and only keep the codes below a cumulative probability threshold.</p>
        <p>The parameters of the entire system are the combination of the
parameters of NCBO or SIFR annotator (list of ontologies, longest match (T/F),
expand mappings (T/F)) with any post-annotation heuristic parameter.
2.5</p>
      </sec>
      <sec id="sec-7-4">
        <title>Reproducibility</title>
        <p>Using the SIFR annotator is as simple as sending a request through the HTTP
REST API, for example to annotate the sentence: \Absence de tumeur maligne"
with the French version of WHO-ART and MedDRA, it can be done with:10
http://services.bioportal.lirmm.fr/annotator/?text=Absence%20de%
20tumeur%20maligne&amp;negation=true&amp;ontologies=WHO-ARTFRE,MDRFRE</p>
        <p>Because of the sensitive nature of the CepiDC data, we have used a local
version of the NCBO and NCBO Annotators in order to avoid sending the data
out on the network. However, to reproduce our results one would require to
have access to the private SKOS dictionaries on the NCBO and SIFR BioPortal
and in that case, use the user interface or the REST web service API. The
reproduction instructions are available here: https://twktheainur.github.io/
bpannotatoreval/LIRMMCLEF2017Task1Instructions.html
3</p>
        <sec id="sec-7-4-1">
          <title>Results</title>
          <p>Using the previously described work ow, we have performed six runs:
French Run 1 (Aligned and Raw): Annotation with longest only parameter
on and running on a local instance of the SIFR Annotator with CIM-10 and
CIM-10DC-ALLMFC as target resources.</p>
          <p>French Run 2 (Aligned and Raw): A fallback strategy starting from the
result le of Run 1, and, for each line without any annotations, takes the
annotations from a second run, which used CIM-10 and CIM-10DC-ALL
as target resources. This is, in essence a late fusion technique, that aims at
increasing the recall, without sacri cing precision.
10 The REST API requires an APIkey to be used (obtained by creating an account on the
portal). In that example, one can use the demo API key and add
&amp;apikey=c34b76530639-4946-81af-8ac76fe809dd at the end of the call.</p>
          <p>English Run 1 (Raw): Annotation with longest only parameter on and
running on a local instance of NCBO Annotator with ICD-10 and ICD-10CDC
as target resources.</p>
          <p>English Run 2 (Raw): Same as Run 1 but with ICD-10CM as additional
target (also available in the NCBO BioPortal): the Clinical Modi cation of
ICD-10 made in the USA for the classi cation of morbidity causes.
3.1</p>
        </sec>
      </sec>
      <sec id="sec-7-5">
        <title>English raw results</title>
        <p>11 teams participated for English raw. Table 1 presents the results obtained by
our two runs against the average and median results of the runs submitted to
this task. The NCBO Annotator obtained results that are exactly the median
value of all the results submitted (all causes). We can measure a slight decrease
in precision and increase in recall with the introduction of ICD-10-CM in Run 2.
Regarding the external causes, the NCBO Annotator obtains a better precision
and f-measure than the average and median results submitted to the challenge.
13 runs have been submitted by 9 teams to the French raw evaluation. 7 runs
have been submitted by 5 teams to the French aligned evaluation. Tables 2 and
3 present the results obtained by our two runs against the average and median
results of the runs submitted to this task. As expected, the SIFR Annotator
did perform similarly on the raw and aligned datasets (as they were processed
exactly with the same work ow). The results are exactly the median value of
all the results with the raw dataset, but slightly under the median value for the
aligned datasets (all causes). Indeed, teams that have used other information
from the aligned dataset probably performed better than the SIFR Annotator
here. Regarding the external causes, we obtain better precision and F1 than the
average and median results submitted to the challenge.</p>
        <p>
          Run1
Run2
The results obtained are in line with what could be expected from our approach,
which really is an evaluation of how the SIFR and NCBO Annotators concept
recognition component works. The simple string matching approach adopted by
Mgrep generally o ers a good precision (around 80%) if the ontology is properly
lexicalized and concretely captures the terms used in the text to annotate. For
CLEF eHealth 2017 task 1, the loss in precision is explained by the nature of
the dictionaries used as a source to produce our terminology: a same label can
correspond to several di erent classes (here ICD10 codes). This practice is usually
strongly avoided when designing ontologies as it inevitably creates ambiguities.
Concerning recall, our performance is also limited since the concept recognizer
does not include any natural language processing techniques that would increase
the amount of matches, handle morphological variants (as simple as plural forms)
or any other alternative concepts. All phenomena that are common in reality
but not captured as synonyms by the source ontologies will not be recognized
properly. Previous evaluation of the NCBO Annotator [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] already identi ed such
limitations. Unsurprisingly, we found the same limitations apply to the French
version. These issues are particularly important when using resources such as
ICD10 that are not designed to be automatically used for annotation (in spite of
it's original mission that is indeed coding medical acts).
        </p>
        <p>NCBO and SIFR Bioportal already include the standard versions if ICD10
and CIM-10, but as mentioned in Section 2.3, they are not meant to be used for
clinical annotation but to serve as a reference for clinicians independently from
actual clinical text, just using these for the tasks yields unsatisfactory results (FR
Raw: P=27.8, R=03.8, F=06.7; FR Aligned. P=27.2, R=04.3, F=07.4; EN Raw
P=31.1, R=08.7, F=13.6), which explains the necessity of using the dictionaries
and corpora provided with the task.</p>
        <p>The purpose of the SIFR Annotator, and originally of NCBO Annotator, was
not to beat task-speci c state-of-the-art systems. The concrete advantages of
the services, both connected to their respective portals come from: (i) the size
and variety of their dictionaries coming from ontologies, (ii) their availability
as a web service that can be easily included in any semantic indexing work ow,
and nally (iii) their adoption of a semantic web vision that strongly encourages
using dereferenceable URIs that can then reused to facilitate data integration and
semantic interoperability. One should also note that the semantic expansion step
(which uses the mappings between ontologies and the is a hierarchies to generate
additional annotations) as well as the post-processing of the annotations (which
scores and contextualizes the annotations) are interesting exclusive features that
not evaluated within CLEF eHealth 2017 task 1.</p>
        <p>Despite of their limitations, the NCBO and SIFR Annotators obtained median
results when compared to the performance of all the participating systems.
Therefore, considering the other discussed advantages, we believe they are two
services that can help in a wide class of text mining or annotation problems, but
of course not for all. It is important to note the systems were not tailored for this
task and their performance will highly vary depending of the data to annotate
and the ontologies targeted.</p>
        <p>Participation in CLEF eHealth 2017 task 1 is a good way of improving our
SIFR Annotator and potentially the NCBO Annotator also. Such improvements
shall be either generic (changes to the overall work ow, independent of task)
or tailored for improving the results to the CLEF eHealth series (Quaero or
CepiDC corpora). In order to better understand the shortcoming of the system,
we sampled 100 false positives and false negatives from the best runs of the SIFR
Annotator on the French aligned development dataset and proceeded to manually
determine the case of the error. Some class of errors are as follows:
{ Errors because of missing synonyms. A few good illustrative examples include:
The code R09.2 \arr^et respiratoire" was not identi ed within the text
\arr^et cardio respiratoire" or \detresse cardiorespiratoire."
The code J96.0 \insu sance respiratoire aigue" was not identi ed within
the text \detresse respiratoire."
Although some of these false negative could be avoided with a richer dictionary
(1st case) or simple synonym generation (2nd case), we found some that
strongly relies on some medical expertise that can hardly be captured by
a dictionary based approach (maybe by a machine learning one, assuming
there is enough data to train the tool).
{ Single match returned whereas a multiple match was expected. Indeed, the
MFC strategy resulted in assigning the synonyms to only one code in
CIM10-DC therefore, when for the same text, several annotations where expected,
we found only one. For instance, the code G40.9 "epilepsie, sans precision"
was found with the text "epilepsie avec etat de mal" but not the code G41.9
"etat de mal epileptique, sans precision." Note that those errors are only
present with the MFC strategy.
{ Morphosyntactic or lexical variation (e.g., accent, dash, comma, spelling).</p>
        <p>For instance, the code J18.9 \emphyseme, sans precision" was not identi ed
within the text \emphyseme pumonaire secodaire tabagisme actif" because
of the spelling of "pulmonaire."
{ Annotations were made with a more general (i.e., parent in ICD-10 hierarchy),
often because of a partial match within an expression.
{ Errors cased by implicit semantic information that requires medical knowledge
to identify. E.g. I10 \hypertension essentielle (primitive)" was not found from
the text \TC suite a une chute avec epilepsie sequellaire et tr cognitifs" as
it was annotated within the corpus. Or the code R68.8 \autres sympt^omes
et signes generaux precises" was not identi ed within the text \atteinte
polyviscerale di use."</p>
        <p>From this review of the pitfalls of the SIFR Annotator on the CepiDC corpus,
and from other in-house experiments, we clearly identi ed the need to improve
the dictionary generation process when extracting the labels from the source
ontologies. The terms can be enriched by adding other alternative synonyms
or morphosyntactic or lexical variations, such that we increase recall without
decreasing precision.</p>
        <p>Another source of improvement for the SIFR Annotator comes from generating
and curating alignments between ontologies. For instance, on the CepiDC corpus,
when used with mappings to other ontologies, the SIFR Annotator was able to
identify unambiguous concepts. The system uses the mappings between ontologies
to expand the original direct annotations made from the text, bu only if the
exist and are uploaded to SIFR BioPortal. In the case of ICD10, there exists
multiple sources of published mappings that we plan to upload in the future so
as to improve recall.</p>
        <p>
          The limitations on one particular application do not generalize to others. We
are currently evaluating the performance of the SIFR Annotator on the Quaero
corpus [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] used during previous editions of CLEF eHealth and we identi ed
other problems such as ambiguity (as we use several ontologies whereas Quaero
is annotated with unique UMLS Metathesaurus concepts) or missing translated
terms (as the Quaero corpus used English UMLS concept directly through a
translation approach that disadantages French-only systems). Overall, despite
of these limitations, our results on the Quaero corpus vary between 63 and
70% F1 on plain entity recognition (UMLS semantic groups) and 36-37% F1 on
normalized entity recognition (UMLS concept unique identi ers) on the EMEA
dataset. And respectively 58-69% F1 and 32-33% the Medline dataset. These
results put the SIFR Annotator among the top systems for plain NER and in
the leading top half for normalized NER. Better Quaero results analysis and
reporting shall be the subject of another speci c future communication.
        </p>
        <p>
          We would also like to point to some recent improvements that we have made
to the SIFR Annotator that are currently under evaluation. When annotating
clinical notes with medical conditions, it is important to lter out negated conditions
and distinguish present conditions from antecedents or conditions experienced
by someone other than the patient. For this reason, several methods have been
proposed to detect the context of already identi ed clinical conditions, especially
for English medical text. The English-language system, NegEx/ConText, is one
of the best and fastest algorithms for the determination of the context of medical
conditions [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. ConText is based on lexical cues (trigger terms) that modify
the status of medical conditions appearing in the scope of the cues. We have
adapted this system to the French language. Our approach consisted in compiling
an extensive list of French lexical cues by a process of automatic and manual
translation and enrichment. Then, we interconnected the NegEx/ConText
program with the NCBO and SIFR Annotators thanks to the proxy architecture
previously mentioned. This feature can already be used on the SIFR BioPortal
for the French and English annotation services. Our rst evaluation con rms the
ability to detect negation with a very high F1 score, slightly improving previous
published work done in the past to adapt NegEx for French. This study shall
also be part of another speci c future communication on the SIFR Annotator.
Due to time limitation, we have not use NegEx/ConText on the CepiDC corpus
although we are not sure about the impact of this feature on the results.
5
        </p>
        <sec id="sec-7-5-1">
          <title>Conclusions</title>
          <p>In this paper, we presented our participation to the task 1 of the CLEF eHealth
2017 challenge using the NCBO and SIFR Annotators. Our results are encouraging:
around 50-60% of F1 score means that more than half of the task of coding death
certi cates with ICD-10 codes can be automatized. Especially considering that
we have not implemented anything speci c to process these data. But of course,
we will have to improve these results to be among the best performing systems.
Some improvements perspectives have been discussed.</p>
          <p>We have also argued in this paper that according to us the technical
performance (F1) shall not be the only argument in evaluating a semantic annotation
tool. The SIFR project has invested a signi cant amount of e ort in order to o er
a generic, open and quite robust platform that could easily be used \at the click
of the mouse" to annotate biomedical text data and access French biomedical
ontologies and terminologies. Someone can upload a new resource to the SIFR
BioPortal and get a dedicated annotation service, interconnected to other existing
ontologies, in a couple of hours.</p>
          <p>Indeed, the SIFR Annotator is di erent from other related tool in French
biomedical text mining as: (i) it is a dynamic web service with JSON-LD outputs
which can be integrated in current programmatic work ows; (ii) it uses public
ontologies both to create annotations and to expand them; (iii) it has access
to one of the largest available sets of publicly available biomedical ontologies
in French. We believe the SIFR Annotator can therefore be used in a large
span of biomedical applications including annotating clinical text data. We are
currently using the service in the context of the French PractiKPharma project
(http://practikpharma.loria.fr) which aims to validate pharmacogenomics
state-of-the-art knowledge on the basis of practice-based evidences, i.e., knowledge
extracted from electronic health records.</p>
        </sec>
        <sec id="sec-7-5-2">
          <title>Acknowledgements</title>
          <p>This work is supported by the French National Research Agency within the
PractiKPharma (grant ANR-15-CE23-0028) and SIFR (grant
ANR-12-JS0201001) projects as well as by the European H2020 Marie Curie actions (grant
701771), the University of Montpellier and the CNRS. We also thanks the US
National Center for Biomedical Ontology for their assistance with the NCBO
Annotator and the CLEF eHealth 2017 organizers for their help.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Blake</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Bio-ontologies|fast and furious</article-title>
          .
          <source>Nature Biotechnology</source>
          <volume>22</volume>
          (
          <year>June 2004</year>
          )
          <volume>773</volume>
          {
          <fpage>774</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N.F.</given-names>
          </string-name>
          :
          <article-title>Biomedical ontologies: a functional perspective</article-title>
          .
          <source>Brie ngs in Bioinformatics</source>
          <volume>9</volume>
          (
          <issue>1</issue>
          ) (
          <year>2008</year>
          )
          <volume>75</volume>
          {
          <fpage>90</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Uren</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cimiano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iria</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Handschuh</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vargas-Vera</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motta</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ciravegna</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Semantic annotation for knowledge management: Requirements and a survey of the state of the art</article-title>
          .
          <source>Web Semantics: Science, Services and Agents on the World Wide Web</source>
          <volume>4</volume>
          (
          <issue>1</issue>
          ) (
          <year>January 2006</year>
          )
          <volume>14</volume>
          {
          <fpage>28</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Language Resources for French in the Biomedical Domain</article-title>
          . In Calzolari, N.,
          <string-name>
            <surname>Choukri</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Declerck</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loftsson</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maegaard</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mariani</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moreno</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Odijk</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piperidis</surname>
          </string-name>
          , S., eds.
          <source>: 9th International Conference on Language Resources and Evaluation</source>
          , LREC'
          <fpage>14</fpage>
          , Reykjavik, Iceland, European Language Resources Association (May
          <year>2014</year>
          )
          <volume>2146</volume>
          {
          <fpage>2151</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Annane</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bouarech</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Emonet</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Melzi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : SIFR BioPortal :
          <article-title>Un portail ouvert et generique d'ontologies et de terminologies biomedicales francaises au service de l'annotation semantique</article-title>
          .
          <source>In: 16th Journees Francophones d'Informatique Medicale, JFIM'16</source>
          ,
          <string-name>
            <surname>Geneve</surname>
          </string-name>
          ,
          <source>Suisse (July</source>
          <year>2016</year>
          )
          <fpage>16</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Noy</surname>
            ,
            <given-names>N.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Whetzel</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dorf</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gri</surname>
            <given-names>th</given-names>
          </string-name>
          , N.B.,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Storey</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chute</surname>
            ,
            <given-names>C.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>BioPortal: ontologies and integrated data resources at the click of a mouse</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>37</volume>
          (
          <article-title>(web server))</article-title>
          (May
          <year>2009</year>
          )
          <volume>170</volume>
          {
          <fpage>173</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Whetzel</surname>
            ,
            <given-names>P.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Team</surname>
          </string-name>
          , N.: NCBO Technology:
          <article-title>Powering semantically aware applications</article-title>
          .
          <source>Biomedical Semantics 4S1(S8) (April</source>
          <year>2013</year>
          )
          <fpage>49</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>The Open Biomedical Annotator</article-title>
          . In: American Medical Informatics Association Symposium on Translational BioInformatics, AMIA-TBI'
          <fpage>09</fpage>
          , San Francisco, CA, USA (March
          <year>2009</year>
          )
          <volume>56</volume>
          {
          <fpage>60</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Melzi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Scoring semantic annotations returned by the NCBO Annotator</article-title>
          . In Paschke,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Burger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Romano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Marshall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Splendiani</surname>
          </string-name>
          , A., eds.:
          <article-title>7th International Semantic Web Applications and Tools for Life Sciences</article-title>
          ,
          <volume>SWAT4LS</volume>
          '
          <fpage>14</fpage>
          . Volume 1320 of CEUR Workshop Proceedings., Berlin, Germany, CEUR-WS.
          <source>org (December</source>
          <year>2014</year>
          )
          <fpage>15</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert N. Anderson</surname>
            ,
            <given-names>C.K.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lavergne</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rey</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rondet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Clef ehealth 2017 multilingual information extraction task overview: Icd10 coding of death certi cates in english and french</article-title>
          . In:
          <article-title>CLEF 2017 Evaluation Labs</article-title>
          and Workshop: Online working Notes, CEUR-WS,
          <year>September</year>
          ,
          <year>2017</year>
          . (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robert</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kanoulas</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spijker</surname>
          </string-name>
          , R., ao Palotti, J.,
          <string-name>
            <surname>Zuccon</surname>
          </string-name>
          , G.:
          <article-title>Clef 2017 ehealth evaluation lab overview</article-title>
          .
          <source>In: CLEF 2017 - 8th Conference and Labs of the Evaluation Forum, Lecture Notes in Computer Science (LNCS)</source>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Harkema</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dowling</surname>
            ,
            <given-names>J.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thornblade</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W.W.</given-names>
          </string-name>
          :
          <article-title>Context: an algorithm for determining negation, experiencer, and temporal status from clinical reports</article-title>
          .
          <source>Journal of biomedical informatics 42(5)</source>
          (
          <year>2009</year>
          )
          <volume>839</volume>
          {
          <fpage>851</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhatia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jonquet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rubin</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiang</surname>
            ,
            <given-names>A.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Comparison of concept recognizers for building the Open Biomedical Annotator</article-title>
          .
          <source>BMC Bioinformatics</source>
          <volume>10</volume>
          (
          <issue>9</issue>
          :S14) (
          <year>September 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>N.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xuan</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Musen</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watson</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Athey</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meng</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>An E cient Solution for Mapping Free Text to Ontology Terms</article-title>
          .
          <source>In: AMIA Symposium on Translational BioInformatics</source>
          , AMIA-TBI'
          <fpage>08</fpage>
          , San Francisco, CA, USA (March
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Simon</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Twigger</surname>
            , Joey Geiger,
            <given-names>J.S.:</given-names>
          </string-name>
          <article-title>Using the NCBO Web Services for Concept Recognition and Ontology Annotation of Expression Datasets</article-title>
          . In Marshall, M.S.,
          <string-name>
            <surname>Burger</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Romano</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paschke</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Splendiani</surname>
          </string-name>
          , A., eds.: Workshop on Semantic Web Applications and
          <article-title>Tools for Life Sciences</article-title>
          ,
          <volume>SWAT4LS</volume>
          '
          <fpage>09</fpage>
          . Volume 559 of CEUR Workshop Proceedings., Amsterdam, The Netherlands,
          <article-title>CEUR-WS</article-title>
          .
          <source>org (November</source>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Sarkar</surname>
            ,
            <given-names>I.N.</given-names>
          </string-name>
          :
          <article-title>Leveraging Biomedical Ontologies and Annotation Services to Organize Microbiome Data from Mammalian Hosts</article-title>
          . In: American Medical Informatics Association Annual Symposium, AMIA'10, Washington DC., USA (November
          <year>2010</year>
          )
          <volume>717</volume>
          {
          <fpage>721</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Groza</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oellrich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Collier</surname>
          </string-name>
          , N.:
          <article-title>Using silver and semi-gold standard corpora to compare open named entity recognisers</article-title>
          .
          <source>In: Bioinformatics and Biomedicine (BIBM)</source>
          ,
          <source>2013 IEEE International Conference on, IEEE</source>
          (
          <year>2013</year>
          )
          <volume>481</volume>
          {
          <fpage>485</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Funk</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baumgartner</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roeder</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bada</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hunter</surname>
            ,
            <given-names>L.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verspoor</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Large-scale biomedical concept recognition: an evaluation of current automatic annotators and their parameters</article-title>
          .
          <source>BMC bioinformatics 15(1)</source>
          (
          <year>2014</year>
          )
          <fpage>59</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Xuan</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mirel</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Athey</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Watson</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meng</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Interactive Medline Search Engine Utilizing Biomedical Concepts and Data Integration</article-title>
          . In: BioLINK: Linking Literature,
          <article-title>Information and Knowledge for Biology, SIG</article-title>
          , ISMB'
          <fpage>08</fpage>
          , Vienna, Austria (
          <year>July 2007</year>
          )
          <volume>55</volume>
          {
          <fpage>58</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.:</given-names>
          </string-name>
          <article-title>E ective mapping of biomedical text to the UMLS Metathesaurus: the MetaMap program</article-title>
          . In: American Medical Informatics Association Annual Symposium, AMIA'01, Washington, DC, USA (November
          <year>2001</year>
          )
          <volume>17</volume>
          {
          <fpage>21</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merabti</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , Gri on, N.,
          <string-name>
            <surname>Dahamna</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Multiterminology cross-lingual model to create the European Health Terminology/Ontology Portal</article-title>
          .
          <source>In: 9th International Conference on Terminology and Arti cial Intelligence</source>
          ,
          <source>TIA'11</source>
          , Paris, France (
          <year>November 2011</year>
          )
          <volume>119</volume>
          {
          <fpage>122</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leixa</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosset</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>The QUAERO French medical corpus: A ressource for medical entity recognition and normalization</article-title>
          .
          <source>In: Proc of BioTextMining Work</source>
          . (
          <year>2014</year>
          )
          <volume>24</volume>
          {
          <fpage>30</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>