<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Classi cation of Cancer Noti able Death Certi cates</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luke Butt</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guido Zuccon</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anthony Nguyen</string-name>
          <email>anthony.nguyeng@csiro.au</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anton Bergheim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Narelle Grayson</string-name>
          <email>narelle.graysong@cancerinstitute.org.au</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cancer Institute NSW</institution>
          ,
          <addr-line>Alexandria, New South Wales</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The Australian e-Health Research Centre</institution>
          ,
          <addr-line>Brisbane, Queensland</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <fpage>65</fpage>
      <lpage>76</lpage>
      <abstract>
        <p>The timely noti cation of cancer cases is crucial for cancer monitoring and prevention. However, the abstraction and classi cation of cancer from the free-text of pathology reports and other relevant documents, such as death certi cates, are complex and time-consuming activities. In this paper we investigate approaches for the automatic detection of cases where the cause of death is a noti able cancer from free-text death certi cates supplied to Cancer Registries. A number of machine learning classi ers were investigated. A large set of features were also extracted using natural language techniques and the Medtex toolkit; features include stemmed words, bi-grams, and concepts from the SNOMED CT medical terminology. The investigated approaches were found to be very e ective in identifying death certi cates where the cause of death was a noti able cancer. Best performance was achieved by a Support Vector Machine (SVM) classi er with an overall F-measure of 0.9647 when evaluated on a set of 5,000 free-text death certi cates. This classi er considers as features stemmed token bigrams and information from SNOMED CT concepts ltered by morphological abnormalities and disorders. However, our analysis shows that it is the selection of features that most in uences the performance of the classi ers rather than the type of classi er or the feature weighting schema. Speci cally, we found that stemmed token bigrams with or without SNOMED CT concepts are the most e ective feature. In addition, the combination of token bigrams and SNOMED CT information was found to yield the best overall performance.</p>
      </abstract>
      <kwd-group>
        <kwd>death certi cates</kwd>
        <kwd>Cancer Registry</kwd>
        <kwd>cancer monitoring and reporting</kwd>
        <kwd>machine learning</kwd>
        <kwd>natural language processing</kwd>
        <kwd>SNOMED CT</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Cancer noti cation and reporting is an important and fundamental process for
providing an accurate picture of the impact of cancer, the nature and extent of
cancer, and to direct research e orts for the cure of cancer. Cancer Registries
collect and interpret data from a large number of sources, helping to improve cancer
prevention and control, as well as treatments and survival rates for patients with
cancer.</p>
      <p>The manual coding of documents, such as pathology reports and death
certi cates, with respect to noti able cancers and corresponding synoptic factors
(such as primary site, morphology, etc.) is a laborious and time consuming
process. Cancer Registries strive to provide timely and accurate information on
cancer incidence and mortality in the community. They receive large quantities
of data from a range of sources, including hospitals, pathology laboratories and
Registries of Births, Deaths and Marriages (which issues releases of death
certi cates). It is estimated that incident cases within Cancer Registries that have
death certi cate only noti cations amount to about 1-5% of the total cases;
delays in the processing of this data may cause underestimation of the incidence of
cancer. Computational methods for the automatic abstraction of relevant
information have the possibility to enhance a Cancer Registry's work ow, providing
time and costs savings as well as timely cancer incidence information and
mortality information. This automatic process is however challenging, both for the
complex nature of the language used in the reports, and for the high level of
recall and accuracy required.</p>
      <p>
        Previous works have attempted to provide automatic cancer coding from
free-text pathology reports collected by Cancer Registries. For example, Nguyen
et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] used natural language processing techniques and a rule-based system
to automatically extract relevant synoptic factors from electronic pathology
reports. Similarly, Zuccon et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] showed how these techniques could cope with
character recognition errors generated by scanning free-text pathology reports
stored in paper form. Machine learning approaches have also been considered; for
instance, D'Avolio et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] have tested approaches based on supervised machine
learning (Conditional Random elds and Maximum Entropy) and have shown
its e ectiveness for the classi cation of pathology reports that were consistent
with cancer in the domains of colorectal, prostate, and lung cancer.
      </p>
      <p>Cancer Registries have access to a number of data sources beyond pathology
reports. One such data source is death certi cates. Death certi cates are a rich
source of data that can support cancer surveillance, monitoring and reporting.
These certi cates contain free-text sections that report the cause of the death of
an individual. An example of the free-text content of a death certi cate where
the cause of death is a noti able cancer is given in Figure 1, while Figure 2 is
an example of a non-noti able death certi cate.</p>
      <p>Limited works have focused on computational methods for automatically
classi ng death certi cates with respect to the cause of death. The
SuperMICAR system and its related tools1 provide a semi-automatic coding of the
cause of death in death certi cates. The system identi es keywords and
expressions from the free-text documents that indicate possible causes of death; this
is done through the use of a standard set of expressions encoded in a prede ned
vocabulary. Extracted free-text expressions are then converted to one or more
1 Consult http://www.cdc.gov/nchs/nvss/mmds/super_micar.htm (last visited 19th
November 2012) for further details.</p>
      <p>(I)A) MAXILLARY TUMOR, 2 YEARS B) PULMONARY OEDEMA, 1 WEEK
(II) CEREBROVASCULAR ACCIDENT/DYSPLASIA, 20 YEARS ASTHMA</p>
      <p>
        ICD-10 codes which are then aggregated into a single ICD-10 underlying cause
of death through the use of a rule-base. While doctor reported death certi cates
can be fed directly into the system, Coroner reported ones require additional
pre-processing. A consistent number (between 15 and 20 percent according to a
US study [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]) of death certi cates cannot be coded through SuperMICAR and
related tools, and thus require manual coding. A recent work has successfully
classi ed death certi cates related to pneumonia and in uenza using a natural
language processing pipeline and rule-based system [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However, to the best
of our knowledge, no previous research has been conducted to investigate fully
automatic methods that go beyond keyword spotting of standard cause of death
expressions to classifying death certi cates, in particular focusing on certi cates
where the main cause of death is cancer. Furthermore, while Australian
Cancer Registries can acquire free-text death certi cates on a fortnightly basis from
the Registry of Births Deaths and Marriages, coded causes of death produced
by SuperMICAR (and related products) are released by the Australian Bureau
of Statistics on a yearly basis. Computational methods able to tackle the fast
identi cation of death certi cates where the cause of death is a noti able
cancer would enhance the cancer reporting and monitoring capabilities of Cancer
Registries.
      </p>
      <p>
        In this paper, we focus on the problem of automatically identifying death
certi cates where the main cause of death is cancer. This problem is cast into
a binary classi cation problem, i.e. death certi cates are classi ed as containing
a death cause related to cancer or vice versa as not containing a death cause
related to cancer. Several machine learning classi ers were investigated for this
task. These include support vector machine, Naive Bayes, decision trees, and
boosting algorithms. A state-of-the-art information extraction tool (Medtex [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ])
is used to create di erent set of features that are used to train the classi ers;
different feature weighting schemas were also considered. Features include stemmed
tokens, n-grams, as well as SNOMED CT concept ids and tokens from fully
speci ed names of SNOMED CT concepts, among others. SNOMED CT is a medical
terminology which formally describes in detail the coverage and knowledge of
topics and terminology used in the medical domain.
      </p>
      <p>Our approaches are tested on 5,000 de-identi ed death certi cates acquired
from an Australian Cancer Registry, using 10-fold cross validation for
allowing robust training and testing. Our experimental results demonstrate that the
choice of classi er and weighting schema, although being important, is not
critical for achieving high classi cation e ectiveness. Instead, the choice of features
used to represent content of death certi cates is the determining factor for high
classi cation e ectiveness. Speci cally, stemmed token bigrams are found to be
the single most important features among those extracted. Furthermore, we
found that SNOMED CT features provide consistent increments in classi cation
e ectiveness if used along with stemmed token bigrams; although not providing a
large increment, the combined use of stemmed token bigrams and SNOMED CT
morphology provide the best classi cation e ectiveness in our experiments.</p>
      <p>Next, we detail the approaches adopted in this paper. Then, in Section 3 we
outline our empirical evaluation methodology; classi cation results obtained by
the investigated approaches are reported in Section 4. An analysis of the results
is developed in Section 4.1. The paper concludes in Section 5 summarising our
main contribution and directions for future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Approaches for Automatic Classi cation of Death</title>
    </sec>
    <sec id="sec-3">
      <title>Certi cates</title>
      <p>In this paper we investigate supervised machine learning approaches for the
detection of death certi cates where the cause of death is a noti able cancer.
These approaches are characterised by three main variables: (1) the features
extracted from the documents (Section 2.1), (2) the weighting schemas applied
to the features to represent documents (Section 2.2), and (3) the speci c binary
classi er used to individuate certi cates where the cause of death is a noti able
cancer (Section 2.3).
2.1</p>
      <sec id="sec-3-1">
        <title>Automatic Feature Extraction</title>
        <p>Machine learning algorithms require data to be represented by features, such as
the words that occur in a text document. We used the information extraction
capabilities of the Medtex system2 for obtaining a set of meaningful features
from the free-text of the death certi cates.</p>
        <p>
          The feature sets investigated in this paper are:
stem: a token stem, i.e. the stemmed version of a word contained in a certi cates
stemBigram: the bi-gram formed by two token stems, i.e. a pair of adjacent
stemmed words as found in a certi cates
concept: SNOMED CT concepts as found in the free-text of the certi cates
using the Medtex system
conceptFull: the tokens of the fully speci ed name of the extracted SNOMED CT
concepts
2 Medtex comprises both information extraction capabilities (extracting both low level
information such as word tokens and stems, punctuation, etc., and higher level
semantic information such as UMLS and SNOMED CT concepts [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]) and classi cation
capabilities integrated via its rule-based engine.
concFullMorph: the tokens of the fully speci ed name of extracted SNOMED CT
concepts that are morphologic abnormalities or disorders
concBigram: the bigram formed by two adjacent SNOMED CT concept ids
concFullBigram: the bigram formed by two adjacent tokens in the fully speci ed
name of concepts extracted from SNOMED CT
        </p>
        <p>While features like stem and stemBigram are commonly used for classifying
free-text documents, features based on SNOMED CT concepts and its properties
such as tokens from the fully speci ed name have not been exploited by
previous works that attempted to classify free-text death certi cates. SNOMED CT
provides a standard clinical terminology used to map various descriptions of a
clinical concept to a single standard clinical concept. In this work, the SNOMED
CT ontology was used as an underlying mechanism to classify free-text using
semantically matching SNOMED CT concepts.</p>
        <p>In addition, we also considered pair-wise combinations of features that
provided promising results on preliminary experiments. In this paper we shall
report the results obtained by all features used singularly, and of the combinations
concept + stem, concept + stemBigram, concFullMorph + stemBigram, and
concBigram + stemBigram, which has shown promise in preliminary investigations.</p>
        <p>Next, we consider the example death certi cates given in Figure 1 and
Figure 2 to describe how a feature set is constructed. To build the feature
representations, we examine each death certi cate and for each occurring instance of a
feature in the certi cate we assign a value of 1, while the absence of a feature is
marked by a zero entry value. Note that these values are subsequently modi ed
according to the feature weighting functions, as we shall describe in Section 2.2.
After all certi cates have been processed in this manner, we add a nal feature
cancerNoti able, whose value is obtained from ground truth judgements supplied
with the data. Table 1 shows an extract of the feature data constructed for the
two example death certi cates. The task of the machine learning classi ers is to
predict the value of the cancerNoti able feature, given the learning data supplied.</p>
        <sec id="sec-3-1-1">
          <title>Document</title>
          <p>Figure 1
Figure 2</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Features</title>
          <p>stem stemBigram concept conceptFull ...</p>
          <p>ICCAD LLCAHOO... TRUOMEEKWERAY IISSLPCCAAAYDD I48CCAD... 02ERYA SETRYAAHAM 621550400 ... 306290007 lillfseoaoaaxpNmm... litrrrsccceeeaaaovudbC liltrrrrrsssceeeeoaaobC...</p>
          <p>n i
1 0 0 1 1 1 1 0 0 1 1 1 0 1 1 1 1 0 ... 1
1 1 1 0 0 1 0 1 1 0 0 0 1 1 0 1 1 1 ... 0
e
l
b
a
i
t
o
N
r
e
c
n
a
c</p>
          <p>Note that no further processing is applied to the text, for example, for
removing punctuation, identifying section or list labels, or for removing or correcting
typographical errors present in the free-text. While adequate text pre-processing
may enhance the quality of the text itself and thus of the extracted features,
we left this for future work and instead we focused on investigating weighting
schemas for the selected features and binary classi ers.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Feature Weighting</title>
        <p>A number of weighting schemes for capturing the local importance of a feature
in a report were tested.</p>
        <p>Binary coe cients were used to encode the presence or absence of a feature.
We refer to this schema as binary.</p>
        <p>The weighting schema composed by the feature frequency f (F ) of feature F
was used to capture the number of times a speci c feature appeared within a
document. We shall refer to this weighting schema as frequency.</p>
        <p>Variations of the frequency weighting schema were also experimented with. In
this weighting schema, features frequencies were directly translated into weights,
i.e. weights are linearly derived from frequencies. Variations consider non-linear
functions of the frequency of a feature.</p>
        <p>A rst variation was to scale the appearance of feature F in a free-text
death certi cate by the function 1 + log(f (F )) if f (F ) 1, and 0 if the feature
was absent. This function would capture the fact that little importance is given
to subsequent appearances of a feature F in a document: the logarithm of a
number greater than one plateaus rapidly. In the following, we shall refer to this
weighting schema as LogF, i.e. logarithm of the frequency.</p>
        <p>A second variation was to assign increasing weights to features that appear
with high frequencies within the death certi cate. To this aim, the appearance of
feature F was weighted according to the function ef(F), while a zero value was
assigned to absent features. It is suggested that, given the short length of the
considered death certi cates, the unexpected multiple occurrence of a feature
would provide strong evidence that that feature is important for the document.
Using the exponential function to weight occurrences of a feature would assign
dominating scores to features that occur frequently in a document. We shall refer
to this weighting function as expF.</p>
        <p>Note that only local weighting functions were used to assign scores to
features,that is, weights were computed only by taking into account the frequencies
of appearance of a feature in a text, thus ignoring the distribution of that feature
on a global level, i.e. across the dataset. The incorporation of global occurrence
statistics within the weighting schemas is left to future work.
2.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Automatic Classi cation Methodology</title>
        <p>
          A number of common classi ers were evaluated. These comprised statistical
models (Naive Bayes), support vector machines (SPegasos), decision trees (C4.5), and
boosting algorithms (AdaBoost). We considered the implementations of these
algorithms provided in the Weka toolkit [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>The multinomial Naive Bayes classi er determines the class of a death
certi cate according to the features that occur in the text and their weights. The
SPegasos classi er uses a stochastic gradient descent algorithm and a hinge loss
function to produce the separation hyperplane used by the linear support vector
machine. In the C4.5 classi er, information gain is used for choosing at each
level of the decision tree the most e ective feature able to split the data into
the two binary classes considered here (i.e. death certi cates related to cancers
and those not related to cancer). Adaboost minimises of a convex loss function
built from the prediction of a base weak classi er. A simple binary decision tree
classi er that constructs one-level trees was used as base classi er for Adaboost.</p>
        <p>
          Parameters of all classi ers were set to the default values described in Witten
et al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
3
3.1
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Data</title>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Methodology</title>
      <p>A set of 5,000 free-text death certi cates was acquired from Cancer Institute
NSW, the institutional entity responsible for maintaining the Central Cancer
Registry in New South Wales. Ethics approval was granted by the NSW
Population &amp; Health Services Research Ethics Committee for this study including to
use the de-identi ed data. The free-text documents were short in length,
containing on average 13.08 words; the (unstemmed) vocabulary contained 3,751
unique words (including section headings and labels).</p>
      <p>Cause of death classi cations based on ICD-10 codes accompanied the
reports. This coding set was acquired from the Australian Bureau of Statistics,
who releases coded data yearly. ICD-10 codings were used to determine the
class each death certi cates belonged to. A list of ICD-10 codes that are cancer
noti able was provided by Cancer Institute NSW.</p>
      <p>The 5,000 death certi cates were extracted from Cancer Institute NSW
archives so that documents were uniformly split across the two classes, i.e. 2,500
certi cates were coded with ICD-10 codes that are for noti able cancers
according to the business rules of Cancer Institute NSW, while the remaining 2,500
were not cancer noti able. The causes of death of the 2,500 death certi cates
for noti able cancers span a total of 367 unique ICD-10 codes.
3.2</p>
      <sec id="sec-4-1">
        <title>Evaluation</title>
        <p>A 10-fold cross validation methodology was used to train and test the classi
cation algorithms. In this methodology, the dataset was randomly divided into 10
strati ed3 folds of equal dimensions. A model for each classi er was then learnt
3 Folds were automatically strati ed with respect to the two target classes, not the
ICD-10 codes.
on nine of these folds, leaving one fold out for testing. The process was repeated
by selecting a new fold for testing, while a new model was learnt from the
remaining folds. Classi cation e ectiveness was then averaged across the folds left
out for testing in each iteration.</p>
        <p>F-Measure (F-m) was used as primary metric to evaluate the e cacy of the
implemented classi ers; accuracy, recall (sensitivity, Rec) and precision
(positive predictive value, Prec) were also recorded, along with the number of true
positive (TP), false positve (FP), true negative (TN), and false negative (FN)
classi cations.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>The combination of 10 features, 4 weighting schemas, and 4 classi ers requires
the evaluation of a total of 160 classi er settings (referred to as runs in the
following) on the dataset consisting of 5,000 death certi cates. While we
evaluated all combinations of features, weighting schema and classi ers, given the
large number of combinations, it is not feasible to report the individual results
for each of the runs. Thus, we report only the settings of the 40 most e ective
runs in terms on F-measure, our primary evaluation metric (Table 2), with the
F-measure of each classi er over all experimented settings graphically shown in
Figure 3. Later in the paper we shall consider a summary evaluation of the
variability of results provided by features, weighting schemas, and classi ers. This
analysis will comprise of the results from all runs.</p>
      <p>The results reported in Table 2 suggest that the tested approaches are highly
e ective in discriminating between those death certi cates that contain a cancer
noti able cause of death and those that do not.</p>
      <p>Overall, the best classi er is the support vector machine implementation
provided by SPegasos when used on concFullMorph + stemBigram features, i.e. the
fully speci ed names of concepts associated to morphological abnormalities and
disorders as encoded in SNOMED CT, weighted using raw frequencies. SPegasos
is found to be very e ective also when other combinations of weighting schemas
and features are considered. In addition, this support vector machine classi er
shows the smallest variance across all considered settings (Figure 3).</p>
      <p>Among the best performing classi ers, AdaBoost used in conjunction with
stemmed bigrams features achieved perfect precision (Prec= 1), at the expense
of recall. Although these results are remarkable, high precision may be considered
less important than high recall in such task. In fact, in a Cancer Registry setting,
it is preferable to have high recall and be considering death certi cate that are
incorrectly reported as containing cancer noti able cause of death, than to have
missed cancer cases. This becomes particularly important if the missed cancer
cases refer to rare cancers. AdaBoost also exhibits the highest variance across
experiment settings among the considered classi ers (see Figure 3).
4.1</p>
      <sec id="sec-5-1">
        <title>The Impact of Classi ers, Weighting Schemas, and Features</title>
        <p>To better understand the role of speci c features, weighting schema, and
classiers on the e ectiveness of the tested approaches, an analysis of the empirical
results where each of the three key characteristics were treated as the controlled
variable is performed.</p>
        <p>We start by examining the impact of each classi cation model on the overall
e ectiveness of the approaches. Table 3 reports maximum (Max(F-m)),
minimum (Min(F-m)), di erence ( ), and variance of F-measure over all runs of
each classi er model. SPegasos is found to be the classi er achieving the
highest maximum and minimum F-measure values, thus extending the observations
made on this classi er when examining the results of Table 2. Instead, while the
Naive Bayes classi er was not found to be amongst the most e ective classi
cation models in our experiments, its robustness is second only to that of SPegasos,
with performance ranges between 0.9337 and 0.7428 in F-Measure. While models
such as C4.5 and Adaboost achieve higher values of F-measure than Naive Bayes,
their minimum performances are lower than that recorded for Naive Bayes.</p>
        <sec id="sec-5-1-1">
          <title>Classi er Max(F-m) Min(F-m)</title>
        </sec>
        <sec id="sec-5-1-2">
          <title>Variance</title>
          <p>SPegasos
Naive Bayes</p>
          <p>C4.5
AdaBoostM1
0.9647</p>
          <p>We continue by analysing the in uence of weighting schemas on the classi
cation results of the approaches investigated in this work. Simple raw frequency
weighting, i.e. frequency, is found to be the most e ective weighting schema.
However, no weighting schema appears to be signi cantly better than another: while</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>Weight Max(F-m) Min(F-m) Variance</title>
          <p>frequency achieves the best performance with a F-measure of 0.9647, the highest
F-measure of the worst performing schema is 0.9622 (expF), just 0.003% lower
than frequency. Furthermore, all weighting schema exhibit the same e ectiveness
when considering the worst performing settings. Thus the range of performance
di erences and their variance do not signi cantly di er across weighting schema.
This may be due to the fact that death certi cates are in general short
documents, where features occur uniformly.</p>
          <p>Feature
stemBigram
concept + bigramStem
concFullMorph + stemBigram
concBigram + stemBigram</p>
          <p>concBigram
concFullBigram</p>
          <p>conceptFull
concept + stemBigram
concept
stem</p>
          <p>Feature is the nal variable of our analysis, and the one with the greatest
impact on classi cation results. The use of the concFullMorph + stemBigram feature
provide the highest F-measure (0.9647), while concFullBigram yields the lowest
maximal F-measure (0.7768): a signi cant di erence of 19.48%. The smallest
variance was demonstrated by stemBigram (2:02 10 4), making it the most
robust feature in our experiment; in addition this feature yielded a maximal
Fmeasure of only 0.003% lower than the best value recorded in our experiments.
The minimal F-measure yield by the stemBigram feature was also greater than
the greatest F-measure values obtained when using half of the features
investigated in our study. These results provide strong indication that, of the variables
analysed, the choice of feature provides the greatest contribution to the classi
cation e ectiveness.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5 Conclusions</title>
      <p>Timely processing of cancer noti cations is critical for timely reporting of cancer
incidence and mortality. Death certi cates are a rich source of data on cancer
mortality. Cancer registries acquire free-text death certi cates on a regular (e.g.
fortnightly) basis. However, the cause of death information needs to be classi ed
to facilitate reporting of cancer mortality. Cause of death information classi ed
using ICD-10 codes is only available on an annual basis. In this paper we
investigated the automatic classi cation of death certi cates to individuate cancer
noti able cause of deaths. The investigated approaches achieved overall strong
classi cation e ectiveness, with a support vector machine classi er trained with
token bigram features and information from the SNOMED CT medical
ontology, and weighted by their frequency in the documents yielding an F-measure
of 0.9647. The choice of features, rather than that of classi ers or weighting
schema, was found to be the determining factor for high e ectiveness.</p>
      <p>Future e orts will be directed towards an in depth error analysis, in particular
examining the distance between the prediction produced by a classi er and the
decision threshold. We also plan to extend the investigation to predict the actual
ICD-10 codes associated to cause of death related to cancer, so as to further assist
clinical coders in processing cancer noti cations.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawley</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colquist</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Automatic extraction of cancer characteristics from free-text pathology reports for cancer noti cations</article-title>
          .
          <source>In: Health Informatics Conference</source>
          . (
          <year>2011</year>
          )
          <volume>117</volume>
          {
          <fpage>124</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Zuccon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergheim</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wickman</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grayson</surname>
            ,
            <given-names>N.:</given-names>
          </string-name>
          <article-title>The impact of OCR accuracy on automated cancer classi cation of pathology reports</article-title>
          .
          <source>Studies in health technology and informatics 178</source>
          (
          <year>2012</year>
          )
          <fpage>250</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D</given-names>
            <surname>'Avolio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Farwell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            ,
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Fitzmeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Harris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            ,
            <surname>Fiore</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Evaluation of a generalizable approach to clinical information retrieval using the automated retrieval console (ARC)</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>17</volume>
          (
          <issue>4</issue>
          ) (
          <year>2010</year>
          )
          <volume>375</volume>
          {
          <fpage>382</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Selected data editing procedures in an automated multiple cause of death coding system</article-title>
          .
          <source>In: Proceedings of the Conference of European Statistics</source>
          . (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Staes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duncan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Igo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Facelli</surname>
          </string-name>
          , J.:
          <article-title>Identi cation of pneumonia and in uenza deaths using the death certi cate pipeline</article-title>
          .
          <source>BMC Medical Informatics and Decision Making</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ) (
          <year>2012</year>
          )
          <fpage>37</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Nguyen</surname>
            ,
            <given-names>A.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawley</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>D.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>R.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clarke</surname>
            ,
            <given-names>B.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duhig</surname>
            ,
            <given-names>E.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Colquist</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Symbolic rule-based classi cation of lung cancer stages from freetext pathology reports</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>17</volume>
          (
          <issue>4</issue>
          ) (
          <year>2010</year>
          )
          <volume>440</volume>
          {
          <fpage>445</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Data Mining: Practical Machine Learning Tools and Techniques: Practical Machine Learning Tools and Techniques</article-title>
          . Morgan Kaufmann (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>