<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SIBM at CLEF eHealth Evaluation Lab 2016: Extracting Concepts in French Medical Texts with ECMT and CIMIND</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chloe Cabot</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lina F. Soualmia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Badisse Dahamna</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan J. Darmoni</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>French National Institute for Health</institution>
          ,
          <addr-line>INSERM, LIMICS UMR-1142</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Normandie Univ., SIBM, TIBS - LITIS EA 4108, Rouen University and Hospital</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents SIBM's participation in the Multilingual Information Extraction task 2 of the CLEF eHealth 2016 evaluation initiative which focuses on named entity recognition in French written text. We report on the indexing of the provided QUAERO dataset with multiple knowledge organization systems (KOS) partially or totally translated in French. The extraction method is available online in the form a webbased service that requests the KOS to extract clinical concepts from Electronic Health Records. It is also available via a user-friendly interface developed for clinicians. We addressed the identi cation of relevant clinical entities within the International Classi cation of Diseases version 10 in the CepiDC dataset with a system based on natural language processing and approximate string matching methods. The results obtained this year were rather satisfactory and attested signi cant progress, particularly in exact match recognition, since our last year's participation.</p>
      </abstract>
      <kwd-group>
        <kwd>Information extraction</kwd>
        <kwd>Bagging</kwd>
        <kwd>Lexical semantics</kwd>
        <kwd>Natural Language Processing</kwd>
        <kwd>Information storage and retrieval</kwd>
        <kwd>Vocabulary controlled</kwd>
        <kwd>Systematized Nomenclature of Medicine</kwd>
        <kwd>Medical Subject Headings</kwd>
        <kwd>International Classi cation of Diseases</kwd>
        <kwd>Uni ed Medical Language System</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Since the amount of digital medical documents has widely expanded in the last
twenty years, the information retrieval from such heterogeneous documents has
become a signi cant challenge to address a large variety of tasks in clinical
and biomedical research as well as personalized medicine. Since 1995, the
department of BioMedical Informatics of the Rouen University Hospital (SIBM,
URL: www.cismef.org) has been working on developing tools to access health
knowledge (information retrieval and automatic indexing) in French [1{6]. More
recently, our team has worked on the evaluation of health information systems
and information retrieval and indexing in Electronic Health Records (EHRs) [
        <xref ref-type="bibr" rid="ref7 ref8">7,
8</xref>
        ]. In this context, a user-friendly tool and a web-based service ECMT
(Extracting Concepts with Multiple Terminologies) is developed. It has been included
in several projects subsidized by the French national research agency [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ]. To
evaluate the precision of ECMT, our team participated in 2015 for the rst time
to the CLEF eHealth evaluation initiative [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], precisely to the clinical named
entity recognition task 1b [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ]. The results obtained during this previous
edition were not satisfactory, partially due to our late-joining participation without
training. This year, based on 2015 results, we participated in the multilingual
information extraction task 2 (phases 1 et 2) [
        <xref ref-type="bibr" rid="ref14">14, 15</xref>
        ]. It aims to fully automatically
identify clinically relevant entities in medical texts in French with several types
of documents: abstracts titles, documents about marketed drugs and death
certi cates. The main motivation in participating is to improve the functionalities
of the tool and to determine the progress achieved since our last year's
participation and our ability to address the issues detected then. ECMT uses natural
language processing (NLP), patterns and exploit several biomedical knowledge
organization systems (KOS).
      </p>
      <p>The rest of the paper is organized as follows. In Section 3 we introduce our
extraction approach and tools used in QUAERO and CepiDC tasks and we
describe our experimental setup. Section 4 reports on our results and on error
analysis and re ections. Finally, Section 5 wraps up concluding remarks and
outlines future work.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Material and methods</title>
      <sec id="sec-2-1">
        <title>Extracting Concepts with Multiple Terminologies: ECMT</title>
        <p>
          ECMT is developed to extract as accurately as possible from texts as input,
a list of candidate health concepts from the 55 KOS included in the Health
Terminology / Ontology Portal (HeTOP) [16]. The extraction is performed at
the phrase level of the text. A SOAP and REST Web services allow to provide a
response in XML for each concept and contains: the o set of the rst and the nal
word contained in the health concept, and which led to a medical concept in the
nal list, the identi er and its semantic type if the health concept is included
in the UMLS Metathesaurus, and the medical specialty of the concept. The
latter is based on manual semantic links between general medical specialties (e.g.
dermatology, oncology, etc.) [17] and the KOS included in HeTOP. ECMT relies
on bag-of-words and also pattern-matching designed for discharge summaries,
procedure reports or laboratory results which contain symbolic data (presence or
absence), numerical data and units of measurement. The method of bag-of-words
[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] was developed initially for information retrieval and it has been adapted for
indexing i.e. only the largest set of words that maps a concept label is extracted,
even if is subsets map other concepts. The method is considered as being more
precise and avoiding noise. The text in the input is normalized and each phrase
is processed separately to extract the concepts. ECMT has also a user-friendly
interface (Figure 1) accessible after authentication (http://ecmt.chu-rouen.fr/).
Several options are available to index the text and described in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>A new option named prioritization was added since 2015. It addresses the
speci c issue related to the noise generated by multiple-terminology indexation.
If this option is active, ECMT returns only the concept from the most reliable
terminology, according to its semantic type (default value: false). When n
identical terms from several terminologies are retrieved, semantic types related to
these terms are computed and the most relevant is determined using set-theoretic
operations. Then, the most pertinent term is retained based on a classi cation
of the HeTOP resources devised manually for each semantic type available in
UMLS. For example, indexing the term \asthme" (asthma) with ECMT results
with 7 concepts retrieved within 7 di erent resources: SNOMED-int, NCIT,
MeSH, Medline Plus, HPO, ICD-10 and ICDC. With the prioritization option
activated, only one concept is retrieved according to the semantic type
corresponding with \asthme" (T47-disease in this case) which is a MeSH concept. If
no MeSH concept could be retrieved for a T47-disease concept, then an NCIT
concept should be prioritized and retrieved, and so on. At this time, only 29 over
128 existing semantic types can be processed with this option.</p>
        <p>Figure 2 gives an example of processing the phrase Cholestases intrahepatiques
brogenes familiales et anomalies hereditaires du metabolisme hepatocytaire des
acides biliaires with all ECMT default options but the activated prioritization
option. ECMT extracts the MeSH terms acides et sels biliaires (CUI C0005391),
cholestase intrahepatique (CUI C0008372), the ICD-10 term E70-E90 anomalies
du metabolisme and the NCI term hereditaire (CUI C0439660). The user can
also visualize the alternative terms and categories.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Extracting Concepts from Death Certi cates with ICD-10:</title>
      </sec>
      <sec id="sec-2-3">
        <title>CIMIND</title>
        <p>The CepiDC track aims at identifying only ICD-10 terms with several versions of
this resource manually curated by CepiDC (see section 2.4). This dataset made of
death certi cates revealed that many of the raw texts provided included spelling
mistakes (french accents, inversions etc.). As ECMT is designed to perform only
exact match using multiple terminologies, poor results have been obtained while
analyzing the CepiDC corpus during the training phase. In this way, we choose
to build CIMIND especially for the CepiDC track to focus on these particular
issues.</p>
        <p>CIMIND is designed to match ICD-10 concepts from the texts as input to
ICD-10 terms in the relevant version of the ICD-10. The extraction is performed
at the phrase level of the text using natural language processing techniques. The
system is built using Python and Python/C extensions and provides a response
in CSV format for each identi ed concept with: (i) the entry text, (ii) the o set
of the rst and the nal word contained in the health concept, (iii) the
ICD10 identi er and (iv) the ICD-10 term. CIMIND performs three main steps to
identify ICD-10 concepts:
Tokenization The input text is sliced into phrases, then words. Afterwards, stop
words are ltered. Finally, spell checking is performed using the Enchant library.
The Enchant library is a generic spell checking library with a C API providing
dictionaries and corrections for a misspelled word.</p>
        <p>Candidate selection To select ICD-10 term candidates eventually matching the
input phrase, a method based on the phonetic encoding algorithm Double
Metaphone (DM) [18] is used to operate a rst approximate term search. In this
way, our system relies on a database storing pre-computed DM codes for each
word available in each ICD-10 version dictionary. First, CIMIND computes DM
codes for each word included in the analyzed phrase. Then, ICD-10 candidates
with corresponding DM codes are retrieved from this database. This step
provides quickly a list of relevant ICD-10 term candidates and allows us to perform
time-consuming analyses on a reduced set of terms in the nal step.
Candidate ranking Finally, a combination of the longest common substring and
Levenshtein distance algorithms provides the candidate ranking. The most likely
term having the highest score is retained as the matching ICD-10 term.</p>
        <p>Figure 3 gives an example of processing the phrase HEMATOME
INTRACEREBRAL AVEC OEDEME ET ENGAGEMENT SOUSFALCIFORME with
CIMIND. CIMIND extracts the ICD-10 concepts engagement sous-falciforme
(G935), hematome intracerebral (I619), and oedeme (R609).</p>
        <p>82944;2013;1;85;2;1;HEMATOME INTRACEREBRAL AVEC OEDEME ET ENGAGEMENT
SOUSFALCIFORME;NULL;NULL;;engagement sous-falciforme;G935
82944;2013;1;85;2;1;HEMATOME INTRACEREBRAL AVEC OEDEME ET ENGAGEMENT
SOUSFALCIFORME;NULL;NULL;;hematome intracerebral;I619
82944;2013;1;85;2;1;HEMATOME INTRACEREBRAL AVEC OEDEME ET ENGAGEMENT
SOUSFALCIFORME;NULL;NULL;;oedeme;R609</p>
        <p>Regarding execution time, CIMIND is able to process a death certi cate as
provided in the CepiDC corpus in about 80ms.
2.3</p>
      </sec>
      <sec id="sec-2-4">
        <title>Biomedical Knowledge Organisation Systems</title>
        <p>The information retrieval system of HeTOP, and thus of ECMT, operates on
more than 55 terminologies in both French and English partially or totally
translated into French, aligned with semantic relations. At the date of the challenge
of the CLEF-eHealth 2016 task 2, thirteen KOS were migrated to In nispan,
a distributed in-memory key/value data store with optional schema, and were
available for ECMT: the Medical Subject Headings (MeSH), the Anatomical
Therapeutic Chemical classi cation (ATC), the Classi cation Commune des
Actes Medicaux (CCAM), the International Classi cation of Diseases version
10 (ICD-10), MedlinePlus, the Systematized Nomenclature of MEDicine
International (SNOMED-Int), Pharmacology, the International Classi cation of
Primary Care (ICPC), the Foundational Model of Anatomy Ontology (FMA), the
Human Phenotype Ontology (HPO), the NCI Thesaurus (NCIT), the Online
Mendelian Inheritance in Man compendium (OMIM) and the Human Rare
Diseases Ontology (HRDO). Table 1 contains their metrics. Each concept of these
KOS, when it is available in the UMLS, has a Concept Unique Identi er. It is
the case for example for the ICD-10 and not for the CCAM.
The QUAERO dataset The QUAERO French Medical Corpus dataset has
been developed as a resource for named entity recognition and normalization in
2013 [19]. The data set has been created by Neveol et al. in the wake of the 2013
CLEF-ER challenge, with the purpose of creating a gold standard set of
normalized entities for French biomedical text. A selection of the MEDLINE titles
and EMEA documents used in the 2013 CLEF-ER challenge were selected for
human annotation and are used in this challenge. Annotations are provided in
the BRAT3 stando format and the annotation process was guided by concepts
in the UMLS. Ten types of clinical entities which are UMLS Semantic Groups
were annotated: Anatomy, Chemical and Drugs, Devices, Disorders, Geographic
Areas, Living Beings, Objects, Phenomena, Physiology, Procedures. The
annotations were made in a comprehensive fashion, so that nested entities were marked,
and entities could be mapped to more than one UMLS concept. In particular:
(i) If a mention can refer to more than one Semantic Group, all the relevant
Semantic Groups should be annotated. For instance, the mention \recidive"
(recurrence) in the phrase \prevention des recidives" (recurrence prevention) should
be annotated with the category \DISORDER" (CUI C2825055) and the
category \PHENOMENON" (CUI C0034897); (ii) If a mention can refer to more
than one UMLS concept within the same Semantic Group, all the relevant
concepts should be annotated. For instance, the mention \maniaques" (obsessive) in
the phrase \patients maniaques" (obsessive patients) should be annotated with
CUIs C0564408 and C0338831 (category \DISORDER"); (iii) Entities which
span overlaps with that of another entity should still be annotated. For instance,
in the phrase \infarctus du myocarde" (myocardial infarction), the mention
\myocarde" (myocardium) should be annotated with category \ANATOMY" (CUI
C0027061) and the mention \infarctus du myocarde" should be annotated with
category \DISORDER" (CUI C0027051).</p>
        <p>The CepiDC dataset Since 1968, the CepiDC, a French National Institute
for Health and Medical Research (Inserm) laboratory, is dedicated to elaborate
annually the national medical causes of death statistics in association with the
French National Institute for Statistics and Economic Studies (Insee), the
dissemination of the data and the studies and researches on the medical causes of
death. These statistics are built from information from the certi cate of death.
The CepiDC team handles a database containing more than 18,000,000 death
records [20]. The CepiDC task consists of extracting ICD-10 codes from the raw
lines of death certi cate text. The task is an information extraction task that
relies on the text supplied to extract ICD-10 codes from the certi cates, line by
line. The dataset includes 65,843 death certi cates processed by CepiDC over
the period 2006-2012. The corpus is supplied in CSV format and each row
contains twelve information elds associated with a raw line of text from an original
death certi cate. The output comprises the 9 input elds plus two text elds
used to report evidence text supporting the ICD-10 code supplied in the twelfth,
nal eld. The tenth eld should contain the excerpt of the original text that
supports the ICD code prediction. The dataset also includes four versions of a
manually curated ICD-10 dictionary developed at CepiDC.
3 http://brat.nlplab.org/stando .html</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and discussion</title>
      <sec id="sec-3-1">
        <title>QUAERO track</title>
        <p>For each track, the MEDLINE abstract titles and EMEA documents, the
webbased service of ECMT is used. Before submitting our runs, we have tested
ECMT with the following options actives: refined, categorizing, semantic
network, prioritization and with the 7 (run2) or 13 (run1) available KOS for
extracting entities and normalized entities. Run 1 uses the following resources:
ATC, CCAM, ICDC, FMA, HPO, IDC-10, Medline Plus, MeSH, NCIT, OMIM,
HPO, Pharma, SNOMED-Int. Run 2 uses the following resources: ATC, CCAM,
ICD-10, Medline Plus, MeSH, Pharma, SNOMED-Int. For the concerns of the
task and the evaluation, the ECMT output is converted into the BRAT format.
Figure 4 is an example of the annotation le obtained with the following
sentence: L' hyperplasie medullosurrenalienne: une etiologie rare de l' hypertension
arterielle { rapport d' un cas.</p>
        <p>T1 DISO 3 35 hyperplasie medullosurrenalienne
#1 AnnotatorNotes T1 C0020507
T2 DISO 63 86 hypertension arterielle
#2 AnnotatorNotes T2 C0020538
T3 ANAT 76 86 arterielle
#3 AnnotatorNotes T3 C0003842</p>
        <p>The results obtained for the challenge are presented in tables 2, 3, 4, 5 (phase
1 entities and normalized entities) and tables 6 and 7 (phase 2 normalization).</p>
        <p>
          To support our discussion, the results obtained in CLEF eHealth 2015 are
presented in table 8 [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>Phase 1: entities and normalized entities The results obtained for the phase
1 challenge are rather satisfactory, especially with the entity recognition with
the following results: in exact match processing, we obtain a precision of 0.5381
and a recall of 0.3784 (run1) with the EMEA corpus and a precision of 0.6407
and a recall of 0.4375 (run2) with the MEDLINE corpus. In inexact match
processing, we obtain a precision of 0.649 and a recall of 0.4869 (run1) with the
EMEA corpus and a precision of 0.7668 and a recall of 0.5865 (run2) with the
MEDLINE corpus.</p>
        <p>For normalized entities, in exact match processing, we obtain a precision
of 0.38 and a recall of 0.2687 (run1) with the EMEA corpus and a precision
of 0.4776 and a recall of 0.3271 (run2) with the MEDLINE corpus. In inexact
match processing, we obtain a precision of 0.4005 and a recall of 0.2842 (run1)
with the EMEA corpus and a precision of 0.4974 and a recall of 0.3412 (run2)
with the MEDLINE corpus.</p>
        <p>As of last year, our results have been improved, especially in exact match
entity recognition. For the MEDLINE track, we improved the precision in exact
match entity recognition by 280% and the recall is improved by more than
three times. Since we corrected the processing of special characters in documents
and the computed o sets, we have been able to actually process the EMEA
documents in exact match and improve our results in inexact match as 2015 F1
is 0.35390 and 2016 F1s are 0.5564 (run1) and 0.5233 (run2).</p>
        <p>Indexing with multiple terminologies leads to having duplicate terms in the
results that decrease the precision. This fact explains the di erences that can be
observed between run1 (13 terminologies) and run2 (7 terminologies). Compared
to last year, this issue has been considered and a new option has been added in
ECMT. This option prioritization allows to retain only the most pertinent
terms when several terminologies add up a same term in the output, and
therefore reduce the noise. This ranking is operated according to the term semantic
types. For each semantic type, a list of the most pertinent terminologies to be
uppermost retained has been devised manually. However, as of today, only 29
semantic types over 128 are processed. The noise introduced by using multiple
terminologies could then be even more reduced in the future.</p>
        <p>Also, some errors in exact match results (compared to inexact match results)
could be explained by slight di erences in terms used. The gold standard uses
UMLS labels while ECMT outputs preferred labels in the original KOS. This
leads to minor di erences between CLEF and ECMT outputs, such as \douleur"
in CLEF output vs. \douleurs" in ECMT output. Finally, as no speci c
processing was done to extract overlapping entities as described in the task, several
nested entities are missed. Other entities are extracted with ECMT but are not
in the gold standard. As they are more precise, these concepts should not be
considered as noise.</p>
        <p>Phase 2: normalization The results obtained for the phase 2 challenge which we
participated for the rst time are also rather satisfactory. We obtain the following
results: in exact match processing, we obtain a precision of 0.6044 and a recall
of 0.4626 (run2) with the EMEA corpus and a precision of 0.5936 and a recall of
0.515 (run1) with the MEDLINE corpus. In inexact match processing, we obtain
a precision of 0.605 and a recall of 0.463 (run2) with the EMEA corpus and a
precision of 0.5938 and a recall of 0.5153 (run2) with the MEDLINE corpus.</p>
        <p>In this phase as in the normalized entities track in phase 1, most errors in
CUIs retrieved are due to di erences between our data and the gold standard's.
As we used up to 13 terminologies from various sources, and HeTOP does not
track versions of these resources yet, most of these errors are related to the data
sources and can also be related to alignments between these sources (and their
di erent versions) and the UMLS.</p>
        <p>Track Precision Recall F1
entities, exact match 0.22840 0.13350 0.16850
entities, inexact match 0.70910 0.63660 0.67090
normalized entities, exact match 0.29530 0.18610 0.22830
normalized entities, inexact match 0.50030 0.36380 0.42130
entities, exact match 0.00400 0.00220 0.00280
entities, inexact match 0.43450 0.29860 0.35390
normalized entities, exact match 0.00440 0.00240 0.00310
normalized entities, inexact match 0.23050 0.14400 0.17730
CIMIND is used to analyze the CepiDC dataset and outputs the results in CSV
format. The results obtained from this CepiDC track is presented in table 9. In
this track, we obtained a precision of 0.6964 and a recall of 0.6634. The number
of terms retrieved are rather decent, but comparing to results of other teams
participating in this track, our error rate is not satisfactory. As the CIMIND
system has been built expressly for the CLEF eHealth 2016 challenge, we lacked
time to improve the nal step performed by our system by testing more edit
distances and combinations of these methods and then upgrade performances.
In this way, it would be quite interesting to participate again in such a task in
the future.</p>
        <p>SIBM-run1
Average
Median</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and perspectives</title>
      <p>For the second year, the multilingual information extraction task 2 of the CLEF
eHealth 2016 evaluation initiative allowed us to evaluate ECMT in a very speci c
context (indexing MEDLINE titles and EMEA documents in French). ECMT is
developed to index EHRs via a web-based service and also via a user-friendly
interface. The actual version of ECMT (v3) is optimized to process around 70,000
EHR per day. Then, ECMT is not quite designed for the kind of datasets,
abstract titles and EMEA documents, proposed in this challenge. Nevertheless,
since our rst participation in 2015, we have been able to improve ECMT
performances thanks to the rst evaluation which then revealed several issues related
mainly to special characters and o sets computed by ECMT.</p>
      <p>The main conclusion of this work and the obtained results is that
improvements are still to be performed to reduce the noise related to multiple
terminologyindexing as our di erent runs have revealed. Also, the recognition itself could
still be enhanced. Moreover, this year's edition has revealed that version tracking
of the resources available in HeTOP could be a major improvement for ECMT
in the future. Regarding the CepiDC track, progress could have been achieved
with more time and prior knowledge of the documents provided in the challenge.
We plan on deepen these two approaches and to participate to other challenges
in the future to keep track of our developments.
15. Neveol, A., Goeuriot, L., Kelly, L., Cohen, K., Grouin, C., Hamon, T., Lavergne, T.,
Rey, G., Robert, A., Tannier, X., Zweigenbaum, P.: Clinical information extraction
at the CLEF eHealth Evaluation lab 2016. In: CLEF 2016 Evaluation Labs and
Workshop Online Working Notes, CEUR-WS, September, 2016.
16. Grosjean, J., Merabti, T., Dahamna, B., Kergourlay, I., Thirion, B., Soualmia,
L.F., Darmoni, S.J.: Health multi-terminology portal: a semantic added-value for
patient safety. Stud Health Technol Inform 166(66) (2011) 129{138
17. Darmoni, S.J., Neveol, A., Renard, J.M., Gehanno, J.F., Soualmia, L.F., Dahamna,
B., Thirion, B.: A medline categorization algorithm. BMC medical informatics and
decision making 6(1) (2006) 7
18. Philips, L.: The double metaphone search algorithm. C/C++ users journal 18(6)
(2000) 38{43
19. Neveol, A., Grouin, C., Leixa, J., Rosset, S., Zweigenbaum, P.: The Quaero french
medical corpus: A resource for medical entity recognition and normalization. In: Proc
BioTextM, Reykjavik, Citeseer (2014)
20. Pavillon, G., Laurent, F.: Certi cation et codi cation des causes medicales de
deces. Bulletin epidemiologique hebdomadaire 30(31) (2003) 134{138</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thirion</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leroy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douyere</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lacoste</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Godard</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rigolle</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brisou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Videau</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goupy</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Piot</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quere</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ouazir</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abdulrab</surname>
          </string-name>
          , H.:
          <article-title>A search tool based on 'encapsulated' mesh thesaurus to retrieve quality health resources on the internet</article-title>
          .
          <source>Medical Informatics and the internet in medicine 26(3)</source>
          (
          <year>2001</year>
          )
          <volume>165</volume>
          {
          <fpage>178</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>Combining di erent standards and di erent approaches for health information retrieval in a quality-controlled gateway</article-title>
          .
          <source>International Journal of Medical Informatics</source>
          <volume>74</volume>
          (
          <issue>2</issue>
          ) (
          <year>2005</year>
          )
          <volume>141</volume>
          {
          <fpage>150</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rogozan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Automatic indexing of online health resources for a french quality controlled gateway</article-title>
          .
          <source>Information processing &amp; management 42(3)</source>
          (
          <year>2006</year>
          )
          <volume>695</volume>
          {
          <fpage>709</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakji</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Letord</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rollin</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          :
          <article-title>Improving information retrieval with multiple health terminologies in a quality-controlled gateway</article-title>
          .
          <source>BMC Health Information Science and Systems</source>
          <volume>1</volume>
          (
          <issue>1</issue>
          ) (
          <year>2013</year>
          ) 1{
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Gri on, N.,
          <string-name>
            <surname>Schuers</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kerdelhue</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kergourlay</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dahamna</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>A search engine to access pubmed monolingual subsets: Proof of concept and evaluation in french</article-title>
          .
          <source>Journal of medical Internet research</source>
          <volume>16</volume>
          (
          <issue>12</issue>
          ) (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Chebil</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Omri</surname>
            ,
            <given-names>M.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>Indexing biomedical documents with a possibilistic network</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>67</volume>
          (
          <issue>4</issue>
          ) (
          <year>2016</year>
          )
          <volume>928</volume>
          {
          <fpage>941</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cabot</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lelong</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lefebvre</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lecroq</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.</given-names>
          </string-name>
          :
          <article-title>Omic data modelling for information retrieval</article-title>
          . In: IWBBIO,
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          (
          <year>2014</year>
          )
          <volume>415</volume>
          {
          <fpage>424</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Lelong</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merabti</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joulakian</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Gri on, N.,
          <string-name>
            <surname>Dahamna</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cuggia</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grabar</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thiessard</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , et al.:
          <article-title>Moteur de recherche semantique au sein du dossier du patient informatise: langage de requ^etes speci que. 15es Journees francophones d'informatique medicale (</article-title>
          <year>2014</year>
          )
          <volume>139</volume>
          {
          <fpage>151</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Dupuch</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Segond</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bittar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gicquel</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Metzger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Separate the grain from the cha : make the best use of language and knowledge technologies to model textual medical data extracted from electronic health records</article-title>
          .
          <source>In: proceedings of the 6th Language and Technology Conference</source>
          . (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Thiessard</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mougin</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Diallo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jouhet</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cossin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcelon</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campillo-Gimenez</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jouini</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grosjean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massari</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , et al.:
          <article-title>Ravel: retrieval and visualization in electronic health records</article-title>
          .
          <source>In: MIE</source>
          . (
          <year>2012</year>
          )
          <volume>194</volume>
          {
          <fpage>198</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanlen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2015</article-title>
          .
          <article-title>In: Experimental IR Meets Multilinguality</article-title>
          , Multimodality, and Interaction. Springer (
          <year>2015</year>
          )
          <volume>429</volume>
          {
          <fpage>443</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grouin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tannier</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>CLEF eHealth evaluation lab 2015 task 1b: clinical named entity recognition</article-title>
          .
          <source>In: Proceedings of CLEF</source>
          . (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Soualmia</surname>
            ,
            <given-names>L.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cabot</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dahamna</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darmoni</surname>
            ,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>SIBM at CLEF eHealth evaluation lab 2015, CLEF (</article-title>
          <year>2015</year>
          ) Working Notes
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kelly</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Suominen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neveol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palotti</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Overview of the CLEF eHealth Evaluation Lab 2016</article-title>
          .
          <source>In: CLEF 2016 - 7th Conference and Labs of the Evaluation Forum. Lecture Notes in Computer Science (LNCS)</source>
          , Springer, September,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>