<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>LIMSI ICD10 coding experiments on CepiDC death certi cate statements</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierre Zweigenbaum</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Lavergne</string-name>
          <email>lavergne@limsi.fr</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LIMSI, CNRS, Univ. Paris-Sud, Universite Paris-Saclay</institution>
          ,
          <addr-line>F-91405 Orsay</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LIMSI, CNRS, Universite Paris-Saclay</institution>
          ,
          <addr-line>F-91405 Orsay</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We describe LIMSI experiments in ICD10 coding of death certi cate statements with the CepiDc dataset of the CLEF eHealth 2016 Track 2. We tested a classi er with humanly-interpretable output, based on IR-style ranking of candidate ICD10 diagnoses. A tf.idf-weighted bagof-feature vector was built for each training set code by merging all the statements found for this code in the training data. Given a new statement, we ranked candidate codes with Cosine similarity. Features included meta-information and n-grams of normalized tokens. We also prepared an ICD chapter classi er with the same method and used it to rerank the top-k codes (k=2) returned by the code classi er. For development we focused on mono-code statements and obtained a P@1 of 0.749 increased to 0.778 by chapter reranking. On the test data we returned one code for each statement, leaving multiple code assignment for future work, and obtained a precision, recall and F-measure of (0.7650, 0.5686, 0.6524).</p>
      </abstract>
      <kwd-group>
        <kwd>ICD10 coding</kwd>
        <kwd>information retrieval</kwd>
        <kwd>CepiDc</kwd>
        <kwd>death certi cates</kwd>
        <kwd>chapter coding</kwd>
        <kwd>reranking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Coding in the International Classi cation of Diseases (ICD) has been the subject
of a number of studies in the past (e.g., [
        <xref ref-type="bibr" rid="ref2 ref9">9,2</xref>
        ]). Until recently, only one shared
task had taken it as its target [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and after ten more years the CLEF eHealth
2016 Task 2 on multilingual information extraction o ers a much larger-scale
and real-life dataset for ICD10 coding.
      </p>
      <p>
        The main methods for coding or more generally for concept normalization
rely on dictionary-based lexical matching and on machine-learning. In the
medical domain, most dictionary-based methods use the UMLS [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or one of the
vocabularies it contains, including the ICD10 classi cation. For the English
language, MetaMap [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is the best known system and combines the wealth of term
variants of the UMLS MetaThesaurus with lexical knowledge including
morphological rules from the UMLS Specialist Lexicon to map phrases to UMLS
concepts. Language-agnostic methods of approximate term look-up have also
been proposed [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        Koopman et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] studied the classi cation of Australian death certi cates as
pertaining to diabetes, in uenza, pneumonia and HIV with SVM classi ers based
on n-grams and SNOMED CT concepts, and with rules. They also addressed
their classi cation into 3-digit ICD-10 codes such as E10. In another study, the
same team [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] trained SVM classi ers to nd ICD-10 diagnostic codes for
cancerrelated death certi cates. Because they focused on cancer, they used a rst-level
classi er to identify the presence of cancer then applied a second-level classi er
to identify the speci c type of cancer, again as 3-digit ICD-10 codes from C00
to C97. An important di erence from the present work is that they targeted
the underlying cause of death, i.e., one diagnosis per death certi cate, whereas
we aim to determine all the diagnoses mentioned in a given death certi cate,
more precisely at the level of each input statement. Besides, the present work
addresses full four-digit ICD-10 codes (e.g., R09.2) instead of three-digit codes
(e.g., R09).
      </p>
      <p>
        The International Classi cation of Diseases is hierarchical, and previous work
has tried to take advantage of this hierarchy to improve classi cation [
        <xref ref-type="bibr" rid="ref4 ref7">7,4</xref>
        ]. We
have performed limited experiments to leverage the hierarchical structure of the
ICD-10 by using an ICD chapter classi er to rerank base results.
      </p>
      <p>The present work addresses the coding of death certi cate statements with
no restriction on the domain, i.e., into the whole of ICD-10. Rather than aiming
at a method that would obtain the best possible results, we were interested
in exploring simple, humanly-interpretable vector space representations which
would highlight the importance of each word in the classi cation. We selected
an Information Retrieval style method based on the vector space model and
Cosine similarity and made experiments with it, reporting a relatively good
precision (0.7650) on the o cial test set. Since we only targeted one code per
statement and reserved multi-label classi cation for further work, we miss the
many additional codes present in statements with multiple codes and obtain a
limited recall of 0.5686 on the o cial test set.</p>
      <p>We describe how we prepared the source data (Section 2), the methods we
tested to nd ICD10 codes for input statements (Section 3), ICD10 chapter
classi cation and its use for reranking (Section 4). Experiments and results on
the training dataset are presented and discussed along the way. We then provide
the results obtained on the test set, perform a short error analysis based on the
training set, and discuss perspectives for future work (Section 5).
2</p>
    </sec>
    <sec id="sec-2">
      <title>Preparation of the material</title>
      <p>CLEF eHealth 2016 participants were provided with the `Aligned Causes' le
(AlignedCauses 2006-2012.csv) which contains 65,843 death certi cates split into
close to 200,000 diagnostic statements. Each diagnostic statement has associated
metadata and zero, one or more ICD-10 codes. In the provided training and test
data, statements with multiple codes were repeated on as many lines as they had
codes. We considered such repeated statements as one statement, and a large
part of our work focused on statements with exactly one code. Besides, we used
the resulting list of diagnostic statements independently of each other, without
taking into account their grouping into death certi cates.
2.1</p>
      <sec id="sec-2-1">
        <title>Basic processing</title>
        <p>We processed each diagnostic statement as follows:
{ tokenize (French style: NLTK v3.2.1, regular expression);
{ remove stop words (French NLTK);
{ remove diacritics;
{ lower-case;
{ stem (Snowball French stemmer in NLTK).</p>
        <p>All programs here and in the remainder of this work were written using Python
v3.5.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Hyphenated words</title>
        <p>A number of medical words are (often neoclassical) morphological compound
words (e.g., cardiovasculaire). These words are often written with hyphens (e.g.,
cardio-vasculaire), and sometimes even split (e.g., cardio vasculaire). To
normalize these words, we adopted the following strategy:
{ Our goal is to replace variants with the split form: we hope this will allow
to take better account of the components of these words (e.g., the
component vasculaire in cardiovasculaire will share counts with the free-standing
vasculaire in arr^et vasculaire).
{ To build a dictionary which normalizes concatenated forms to split forms,
We collect tokens with hyphens in statements.</p>
        <p>We produce concatenated forms from these tokens.</p>
        <p>We associate the concatenated form to the split form in the dictionary.
Once all forms in the training corpus have been processed, dictionary
expansions are examined to nd tokens which are themselves an entry
in the dictionary: these tokens are replaced with their split form.
{ During tokenization, hyphenated forms are split on hyphens. Their
components are checked for further splitting based on the above dictionary.
Nonhyphenated forms which have an entry in the dictionary are replaced with
the associated split form.
{ In a few cases, components are not meaningful in French (e.g., the English
load word pace-maker, or the misspellings inhalat-ion, aort-ique, ac-fa). A
small set of exceptions (the four above-mentioned cases) was compiled
manually based on an examination of the dictionary built with the training
corpus. Instead of being split during tokenization, they are replaced with their
concatenated form (pacemaker, inhalation etc.).</p>
        <p>The addition of this processing improved precision by about 1pt in all
experiments. For the sake of space, we only present below experiments where this
processing is included.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>ICD-10 data</title>
        <p>We obtained the ICD-10 hierarchy from the UMLS MRREL table and used it
to know which code belongs to which ICD-10 chapter. In some experiments we
also used the actual labels of the ICD-10 diagnoses. For this purpose, we used
the French ICD10, German Modi cation, version of 2014, downloaded on 12 Feb
2015.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>IR-style, vector space representation for coding</title>
      <p>We tested information-retrieval type methods as described below.
3.1</p>
      <sec id="sec-3-1">
        <title>Principles</title>
      </sec>
      <sec id="sec-3-2">
        <title>Representation and training</title>
        <p>{ Each statement of the training corpus is tokenized and normalized as
described in Section 2.
{ One `document' is created for each code: all the statements associated with a
given ICD-10 code Ci are grouped into one text collecting its tokens T (Ci) =
(t1; t2; : : : tn).</p>
        <p>Statements with more than one code are excluded from our training split:
the goal is to gather into our training material only statements which are
fully associated with one code (`mono-code' statements).
{ A bag-of-words, tf.idf model is computed from the token document matrix
(using the tf.idf implementation in gensim v12.4).</p>
      </sec>
      <sec id="sec-3-3">
        <title>Coding as search</title>
        <p>{ Each tested document is represented as a bag-of-words with tf.idf weights
according to this model.
{ This representation is compared to the representations of the training
`documents', which represent the ICD10 codes present in the training corpus.</p>
        <p>Cosine similarity is used.</p>
        <p>{ The top-N most similar documents (codes) are returned.</p>
        <p>Only the top code is used in the prediction. Additionally, the following codes
are computed to evaluate success-at-N metrics (this is detailed in Section
Experiments below). Besides, examining the quality of the top-N codes can inform
us about the potential relevance of reranking methods which could be applied
to these top-N codes.
3.2</p>
      </sec>
      <sec id="sec-3-4">
        <title>Experiments</title>
        <p>Base principle: test on last 10,000 statements, train on rest ({10k*) For our
experiments, we split the CepiDc training corpus into a training split and a test
split. The training split consists of the training corpus except its last 10,000
statements (i.e., the rst 185,203 statements). The test split consists of the last
10,000 statements, which were all coded in year 2012.</p>
        <p>A tf.idf model was created from the training corpus as explained above. Only
mono-code statements were used for training.</p>
        <p>Statements in the test split were coded as explained above. For each such
statement, we measured the following information: whether the correct code was
found among the top N codes (Success@N, or S@N, for N 2 [1 : : : 5], as named
for instance in TREC); note that S@1 is equal to P@1 (precision computed by
examining only the top returned code), and that we therefore often use the term
\precision" to comment success results; for some of the tests, the o cial scoring
program was also applied to the system results and computed Precision, Recall
and F-measure based on all statements in the test split (columns P, R, F).</p>
        <p>We were mostly interested in evaluating results on mono-code statements,
since these are those for which our system is relevant. As we shall see below,
they represent 77% of the statements in the test split. However, for information,
we also evaluated results on all statements with at least one code (98% of the
statements in the test split; the remaining 1.7% have no associated code). These
two settings imply a di erent number of statements submitted to the test.</p>
        <p>Table 1 provides this information for the various experiments described
below. The experiment name ({10k*) codes the characteristics of the experiment,
such as 10k (10,000 statements in the test split), 1 (mono-codes) vs. a (all
statements with at least one code), etc. Column Sttmts shows the actual number of
statements submitted to the test, then the rest of the columns provide
evaluation results as described above. The same statements are evaluated by testing
whether the code found has the correct ICD10 chapter (rows Chapter, shown for
a subset of the evaluations). The statistical signi cance of di erences between
experiments was computed using approximate randomization at p &lt; 0:01.</p>
        <p>We now turn to the description of a series of experiments performed by
adding various features to the representation of statements.</p>
        <p>Features = tokens (Table 1, rows {10k1t) In this test, the only features for a
training or test statement are its normalized tokens. We test our classi er on
the subset of 7669 statements with exactly one code among the 10,000 test
statements (mono-code statements): in this setting, one precision point (1%)
represents 77 statements. As shown on row {10k1t Codes, 71.1% of these
statements obtain a correct code at rank 1, and for 84.1% the top code belongs to the
correct ICD10 chapter (row Chapter). Success increases signi cantly between
S@1 and S@2 both for codes (12pt from 0.711 to 0.836) and for chapters (10pt
from 0.841 to 0.941).</p>
        <p>Use of the standardized form of statements (only available for training) ({10k1ts)
Each CepiDc training data line includes a standardized form of the part of the
statement which supported the choice of a target code (StandardText eld).
In the CepiDc coding process, the human coder records in this eld the text
segment which led to the chosen ICD10 code for the current line, keeping only
content words and correcting spelling errors if any. If training is performed on
the tokens of the standardized form of a statement instead of the tokens of
the original form, performance decreases by 2pt (codes, not shown in table). If
training is performed on the tokens of both the original and the standardized
form of a statement (rows {10k1ts), performance decreases less, by 0.5pt to
S@1=0.706 (codes), and this is con rmed when evaluating on all statements
(not shown in table). We therefore do not pursue with this feature.
Meta-information ({10k1tm) Meta-information, provided in the CepiDc training
and test data, was added as follows to training and test bag-of-words:
{ Gender: two features, =g0 and =g1
{ Age: 5 features corresponding to age ranges based on clusters of initial
veyear age ranges : =a0, =a5-20, =a25-35, =a40-65, =a70-; these clusters were
manually selected based on a study of the distribution of ages per ICD10
chapter in the training set
{ LocationOfDeath: 6 features, one for each LocationOfDeath in the source:
=l1, =l2, . . . , =l6
{ Type of interval: 5 features, one for each LocationOfDeath in the source:
=i1, =i2, . . . , =i5
The addition of these four types of meta-information (rows {10k1tm) slightly
improves S@1 (+0.3pt for codes at 0.714, not signi cant, no improvement for
chapters), and S@2 by 2.6pt (codes: 0.862).</p>
        <p>Bigrams ({10k1tmb), pairs ({10k1tmB), trigrams ({10k1tmbt), triples ({10k1tmbT),
tetragrams ({10k1tmbt4) of tokens were added as follows to training and test
bag-of-words:
{ After tokens were computed and normalized and stop-words removed,
sequences of n consecutive tokens were joined into one n-gram (with
space separator);
the list of normalized tokens was sorted (in alphabetical order); sets of
non-necessarily contiguous tokens were joined into one pair (resp. triple)
(with space separator); the order of tokens in pairs or triples is always the
alphabetical order. Note that pairs include sorted bigrams, and triples
include sorted trigrams.</p>
        <p>The addition of bigrams on top of meta-information (rows {10k1tmb) improves
S@1 by 2.0pt (codes: 0.734, chapters: 0.861). If pairs are used instead of bigrams
(rows {10k1tmB), S@1 improves by another 0.9pt (codes: 0.743) but does not
change for chapters (+0.1pt at 0.861).</p>
        <p>The addition of trigrams on top of bigrams (rows {10k1tmbt) obtains an
improvement of 0.4pt on S@1 (codes: 0.738, chapters: 0.860). The addition of
trigrams to pairs (rows {10k1tmBt) does not change S@1 much (codes: 0.744,
chapters: 0.860).</p>
        <p>The addition of pairs and triples (rows {10k1tmBT) instead of bigrams and
trigrams further improves S@1 (codes: 0.749).</p>
        <p>Adding tetragrams to bigrams and trigrams increases S@1 by 0.3pt ({10k1tmbt4,
codes 0.741). Adding 5-grams ({10k1tmbt45) does not obtain a further increase.
ICD10 labels ({10k1tmbt4i) Some ICD10 codes present in the test corpus may
occur rarely or not at all in the training corpus. To make sure that every ICD10
code is known to the trained model, we add ICD10 terms as additional training
material. They are added to the ICD code `documents':
{ The label of each ICD10 code is tokenized and normalized as described in</p>
        <p>Section 2.
{ One `document' is created for each code: if multiple labels are available for
a code, they are pasted into one such document.
{ These `documents' are merged with those obtained from the training corpus.
{ A bag-of-words, tf.idf model is computed from the token document matrix
as explained in Section 3.1.</p>
        <p>Added to one of the best con gurations (1{4-grams, meta-information), this
does not bring a signi cant change S@1 (+0.1pt = 0.742).</p>
        <p>Testing on statements with at least one code ({10ka*) All tests until now were
performed on the mono-code statements of our test split. Our system is designed
to produce one code per statement, whatever the expected number of codes. In
the case of a multi-code statement, the evaluation considers this code correct if
it matches any of the expected codes. Testing on all statements with at least one
code therefore results in a slightly higher precision: for instance, with 1{4-grams
and meta-information ({10kambt4), +1.2pt for codes and +0.8pt for chapters.
Di erences among the top performing features are not signi cant anymore in
this setting.</p>
        <p>O cial evaluation program We also applied the o cial evaluation program to
the same system output, comparing it to the full set of 10,000 statements in our
test split. Its precision is consistent with Success@1. It adds the measurement
of recall, which is necessarily impaired when our system is run only on the 7669
mono-code subset of the test split, and which is harmed anyway because our
system returns one code per statement even for multi-code statements. This results
in an F-measure of 0.640 for {10katmbt4 (1{4-grams and meta-information),
which we retain for further experiments.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Chapter classi cation and reranking</title>
      <p>We have seen that success increases sharply from S@1 to S@2, suggesting that
reranking the top proposed codes could lead to improvements. For instance,
nding the ICD10 chapter for a statement, if precise enough, could help rerank
candidate codes.
4.1</p>
      <sec id="sec-4-1">
        <title>Chapter classi cation as search</title>
        <p>Therefore, we created an ICD10 chapter classi er in the same way as the ICD10
code classi er:
{ Each tested document is represented as a bag-of-words with tf.idf weights
according to this model.
{ This representation is compared to the representations of the training
`documents', which represent the ICD10 chapters present in the training corpus.</p>
        <p>Cosine similarity is used.</p>
        <p>{ The top-N most similar documents (chapters) are returned.</p>
        <p>Only the top code is used in the prediction (the following codes are only
computed to evaluate success at N metrics).</p>
        <p>Experiments Table 2 shows the results obtained for chapter classi cation on
our test split. Using the best features found for codes (1{4-grams and
metainformation), chapter classi cation reaches a S@1 of 0.873 (mono-codes) or 0.882
(statements with at least one code), which looks reasonably good enough to
attempt to use it for reranking.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Reranking codes based on chapter classi cation</title>
        <p>We use the results of ICD10 chapter classi cation to rerank ICD10 code
classication:
{ Only the top-predicted chapter is used, if any.
{ The top N predicted codes are examined:</p>
        <p>The rst code which belongs to that chapter is pushed to top position,
if any;</p>
        <p>Otherwise, code ranking is kept unchanged.
{ We experimented with N 2 f1 : : : 5g in the training set and found out that
a conservative value of N = 2 obtained the best results.</p>
        <p>Experiments Table 3 shows the results obtained for chapter-reranked code
classi cation on our test split. Using the best features found for codes (1{4-grams
and meta-information, or unigrams, pairs and triples, and meta-information),
chapter-reranked code classi cation boosts S@1 by 4.5pt (0.786, mono-codes) or
3.2pt (0.785, statements with at least one code): both di erences are signi cant.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results and discussion</title>
      <sec id="sec-5-1">
        <title>Results</title>
        <p>The system was trained with parameters {10kBatmbt4 (codes reranked with
chapters, with 1{4grams and meta-information) on the full training set then
applied to the test dataset. It was asked to produce one code per input statement,
as on the training set. Table 4 shows the results it obtained. For comparison, we
recall in the same table the results obtained with the same settings by training
on our training split of the training corpus and testing on our test split of the
training corpus.
Without surprise, recall and F-measure remain lower than precision, as on the
development set, mainly because our system does not include provision for
multilabel classi cation.</p>
        <p>We also observe that the precision, recall and F-measure of the system did not
change much (about {2pt for P, {1pt for R and {1pt for F) from the development
set to the test set: this shows that our system did not over t the training corpus.</p>
        <p>We examined the most frequent false positive codes in our test split as follows.
Given the set of statements associated with an ICD code in the gold standard
codes, we counted the number of incorrect codes in system results (false positives)
for these statements. We ranked them in decreasing order of false positives and
show the top 20 in Table 5.</p>
        <p>Many of the top false-positive codes are among the most frequent codes (e.g.,
A419, R092, etc.).</p>
        <p>The code with the largest number of errors is I619 (Hemorragie intracerebrale,
sans precision [Intracerebral haemorrhage, unspeci ed]). This code frequently
occurs (105 times in our test split) and was confused by our system 102 times
(97%) with other codes. Among these it was confused 40 times with I691 (Sequelles
d'hemorragie intracerebrale [Sequelae of intracerebral haemorrhage]). Table 6
shows the vector space representation (the top ve features) of I619 and its
topthree confused codes. Because of these representations, statements such as AVC
hemorragique (30 occurrences in the test split), which are represented as 'avc',
'hemorrag', 'avc hemorrag', are more similar to I691 than to I619, although they
miss the sequel feature which would make them really relevant for I691.</p>
        <p>P524 (Hemorragie intracerebrale (non traumatique) du f tus et du
nouveaune [Intracerebral (nontraumatic) haemorrhage of fetus and newborn]) is the
second most confused code with I619. The discriminant factor for P524 with respect
to I619 is that it is a perinatal diagnosis in Chapter XVI of ICD10. However,
age=0 (less than 5 years old), an important feature for that ICD10 chapter, did
not make it to these top ve features, i.e., it has a lower tf.idf weight than the
features displayed in Table 6. Statements such as hemorragie cerebrale [Cerebral
haemorrhage] or hemorragie intra-cerebrale [Intracerebral haemorrhage] and the
like, whatever their age metadata, have therefore nearly as strong unigram
similarity to I619 as to P524, and have higher bigram similarity to it. The absence
in a statement of the feature age=0, which is present in all examples of P524 in
the training corpus, does not prevent the IR-style classi er from considering the
representation of hemorragie cerebrale (with an age di erent from 0) as more
similar to that of P524 (which includes the age=0 feature) than to that of I619.</p>
        <p>Finally, S062 (Lesion traumatique cerebrale di use [Di use brain injury]) is
the third most confused code with I619. Its top ve features also show its strong
similarity to I619 and give higher weights to 'cerebral', 'hemorrag', 'intra', 'intra
cerebral' than I619. Because of that, statements such as hematome intra-cerebral
[Intracerebral haemorrhage] or hematome cerebral [Cerebral haemorrhage]
obtain a representation whose 'cerebral', 'intra', and 'intra cerebral' features get a
higher similarity score to S062 than to I619. The discriminant factor for S062
is the fact that in the context of Chapter XIX of ICD10 (traumas, poisoning,
etc.), it must be a traumatic lesion. In the CepiDc death certi cates, this factor
is often present not in the statement itself but in its context, i.e., in the other
statements of the same certi cate. Contrast for instance hematome intra-cerebral
in Examples 1 and 2 below, each of which shows an excerpt of a death certi
cate as annotated in the gold standard data. In Example 1 there is no speci c
context, therefore it is coded as a cardiovascular diagnosis (I619); in Example 2
the problem in Statement 2 is caused by a fall (chute in Statement 3), therefore
it is coded as a traumatic injury (S062).</p>
        <p>Example 1.
77938;2012;2;75;2;1;detresse respiratoire;3;5;1-1;detresse respiratoire;J960
77938;2012;2;75;2;2;Hematome intra-cerebral;NULL;NULL;2-1;hematome
intracerebral;I619
Example 2.
65397;2012;1;80;2;1;Hypertension intra-cr^anienne;NULL;NULL;1-1;hypertension
intra-cr^anienne;G932
65397;2012;1;80;2;2;hematome intra-cerebral;NULL;NULL;2-1;hematome
intracerebral;S062
65397;2012;1;80;2;3;chute de sa hauteur;NULL;NULL;3-1;chute de sa hauteur;W18
[...]</p>
        <p>This analysis of the top false positive codes for I619 gives an idea of the types
of shortcomings of our method and outlines directions for future work.
5.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>Future work</title>
        <p>Improving precision. We observed that some discriminant features were not given
su cient weight in the classi cation. While using more state-of-the-art IR scores
and similarities such as BM25 or language models might improve upon tf.idf and
Cosine, we believe that using a more classical machine learning classi er based
on individual training examples (instead of global pseudo-documents as we did in
an Information Retrieval style) might lead to more important changes. Since our
submission, we also experimented with various classi ers, among which Logistic
Regression and Support Vector Machines, which obtained a precision of 0.89{
0.91 on mono-codes (+15pt) with only tokens as features.</p>
        <p>The context of a statement, i.e., the other statements of the same certi cate,
may provide important information in some cases, as we observed in our error
analysis. We plan to incorporate this context as additional features for classi
cation. Another possibility will be to use such features in a reranking stage, as
we already started to do with our chapter classi er.</p>
        <p>Improving recall. The main cause of the low recall of our classi er is that it
only proposes one code for each statement, whereas some statements may lead
to multiple codes. A possible direction to address it is to select the top-ranked
code, then to greedily remove the tokens which support it, and to iterate on the
remaining tokens as long as codes are proposed with su ciently high con dence.</p>
        <p>Adding a more traditional strategy based on simple dictionary matching
would be yet another way to identify sequences of diagnoses in a statement.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aronson</surname>
            ,
            <given-names>A.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lang</surname>
            ,
            <given-names>F.M.:</given-names>
          </string-name>
          <article-title>An overview of MetaMap: historical perspective and recent advances</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          <volume>17</volume>
          (
          <issue>3</issue>
          ),
          <volume>229</volume>
          {
          <fpage>36</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Blanquet</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zweigenbaum</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>A lexical method for assisted extraction and coding of ICD-10 diagnoses from free text patient discharge summaries</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>6</volume>
          (
          <issue>suppl</issue>
          ) (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The Uni ed Medical Language System (UMLS): Integrating biomedical terminology</article-title>
          .
          <source>Nucleic Acids Research</source>
          <volume>32</volume>
          (
          <issue>Database issue</issue>
          ),
          <source>D267{270</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Kamkar</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phung</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Venkatesh</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>Stable feature selection for clinical prediction: exploiting ICD tree structure using tree-lasso</article-title>
          .
          <source>J Biomed Inform</source>
          <volume>53</volume>
          ,
          <issue>277</issue>
          {290 (Feb
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karimi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuire</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muscatello</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kemp</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Truran</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Thackway</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          :
          <article-title>Automatic classi cation of diseases from free-text death certi cates for real-time surveillance</article-title>
          .
          <source>BMC Med Inform Decis Mak</source>
          <volume>15</volume>
          ,
          <issue>53</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Koopman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergheim</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grayson</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <string-name>
            <surname>Automatic</surname>
            <given-names>ICD</given-names>
          </string-name>
          <article-title>-10 classi cation of cancers from free-text death certi cates</article-title>
          .
          <source>Int J Med Inform</source>
          <volume>84</volume>
          (
          <issue>11</issue>
          ),
          <volume>956</volume>
          {965 (Nov
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Perotte</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pivovarov</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natarajan</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiskopf</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wood</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elhadad</surname>
          </string-name>
          , N.:
          <article-title>Diagnosis code assignment: models and evaluation metrics</article-title>
          .
          <source>J Am Med Inform Assoc</source>
          <volume>21</volume>
          (
          <issue>2</issue>
          ),
          <volume>231</volume>
          {237 (Mar-Apr
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pestian</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brew</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matykiewicz</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hovermale</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Johnson, N.,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>K.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duch</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>A shared task involving multi-label classi cation of clinical free text</article-title>
          .
          <source>In: Biological, translational, and clinical language processing</source>
          . pp.
          <volume>97</volume>
          {
          <fpage>104</fpage>
          . Association for Computational Linguistics, Prague, Czech Republic (
          <year>June 2007</year>
          ), http://www. aclweb.org/anthology/W/W07/W07-1013
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Wingert</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rothwell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , Co^te, R.A.:
          <article-title>Automated indexing into SNOMED and ICD</article-title>
          . In: Scherrer,
          <string-name>
            <surname>J.R.</surname>
          </string-name>
          , Co^te,
          <string-name>
            <given-names>R.A.</given-names>
            ,
            <surname>Mandil</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.H</surname>
          </string-name>
          . (eds.)
          <source>Computerised Natural Medical Language Processing for Knowledge Engineering</source>
          , pp.
          <volume>201</volume>
          {
          <fpage>239</fpage>
          . NorthHolland, Amsterdam (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>MaxMatcher: Biological concept extraction using approximate dictionary lookup</article-title>
          .
          <source>In: Proceedings of the 9th Paci c Rim International Conference on Arti cial Intelligence</source>
          . pp.
          <volume>1145</volume>
          {
          <fpage>1149</fpage>
          . PRICAI'
          <volume>06</volume>
          , Springer-Verlag, Berlin, Heidelberg (
          <year>2006</year>
          ), http://dl.acm.org/citation.cfm? id=
          <volume>1757898</volume>
          .
          <fpage>1758065</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>