<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>David J. Muscatello, Tim Churches, Jill Kaldor, Wei Zheng, Clayton Chiu, Patricia Correll, and Louisa
Jorm. An automated, broad-based, near real-time public health surveillance system using presentations
to hospital emergency departments in New South Wales, Australia. BMC public health</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1471-2458</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Assessing the performance of American chief complaint classi ers on Victorian syndromic surveillance data</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bahadorreza Ofoghi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karin Verspoor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computing and Information Systems</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Health and Biomedical Informatics Centre The University of Melbourne Melbourne</institution>
          ,
          <addr-line>Victoria</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2005</year>
      </pub-date>
      <volume>5</volume>
      <issue>1</issue>
      <abstract>
        <p>Syndromic surveillance systems aim to support early detection of salient disease outbreaks, and to shed timely light on the size and spread of pandemic outbreaks. They can also be used more generally to monitor disease trends and provide reassurance that an outbreak has not occurred. One commonly used technique for syndromic surveillance is concerned with classifying Emergency Department data, such as chief complaints or triage notes, into a set of pre-de ned syndromic groups. This paper reports our ndings on the investigation of the utility and e ectiveness of two existing North American methods for free-text chief complaint classi cation on a large data set of Australian Emergency Department triage notes, collected from two hospitals in the state of Victoria. To our knowledge, these methods have never before been analysed and compared against each other for their applicability and e ectiveness on free text chief complaint classi cation at this scale or in the Australian context.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        existing systems utilize keyword-based, linguistic, statistical, and/or character-level data from CC texts in the
classi cation process. Ivanov et al. [11] conducted a retrospective study to ascertain the potential of free-text
CCs collected in pediatric emergency departments. They used the Bayesian classi er implemented in Complaint
Coder (CoCo) [
        <xref ref-type="bibr" rid="ref13">16</xref>
        ]. On a population of children less than ve years of age, for early detection of respiratory
and gastrointestinal outbreaks, they found that: i) time series of automatically coded free text CCs related to
pediatric patients correlated with hospital admissions, and ii) the same time series preceded hospital admissions
by the mean of 10.3 and 29 days for respiratory and gastrointestinal outbreaks, respectively.
      </p>
      <p>
        In this paper, we present an evaluation of the applicability and e ectiveness of two existing machine
learningbased North American CC classi ers, namely Symptom Coder (SyCo) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Complaint Coder (CoCo). These
tools are both part of a computer-based public health surveillance system named Real-time Outbreak and
Disease Surveillance (RODS) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Our analysis of SyCo and CoCo on free text CC classi cation is based on a syndromic surveillance data set
that includes ED triage notes from two di erent hospitals in the Australian state of Victoria. We focus on
three primary medical conditions of particular interest in public health surveillance: Flu Like Illness, Acute
Respiratory, and Diarrhoea. According to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and to our knowledge, SyCo and CoCo have never been evaluated
against each other on any surveillance data set at the level of complexity and size as the data we used for this
study.
      </p>
      <p>
        The only previous research that compared the two systems is the work in [
        <xref ref-type="bibr" rid="ref14">18</xref>
        ] which only focused on 1,122
CC entries on a single disease, i.e., In uenza Like Illness. Our analysis goes beyond this by considering the
three above-mentioned diseases and a much larger syndromic surveillance data set (with a total of 314,629 CC
records) as will be introduced in section 2.2.
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <sec id="sec-2-1">
        <title>Symptom Coder and Complaint Coder</title>
        <p>
          Symptom Coder (SyCo) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and Complaint Coder (CoCo) [
          <xref ref-type="bibr" rid="ref13">16</xref>
          ] are two chief complaint classi er systems
developed as parts of the Real-time Outbreak and Disease Surveillance (RODS) system [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. Both CoCo and SyCo
implement Nave Bayes text classi cation which assumes conditional independence between features. CoCo is a
direct complaint-into-syndrome classi er which nds the posterior probabilities for each of its eight syndromic
categories (i.e., constitutional, respiratory, gastrointestinal, hemorrhagic, botulism, neurological, respiratory,
and other) given the text of a chief complaint. CoCo can classify CCs into one of the eight syndromic categories
constitutional, respiratory, gastrointestinal, hemorrhagic, botulism, neurological, respiratory and other.
        </p>
        <p>
          The Bayesian probabilities that CoCo calculates are determined using a default probability le (developed
by RODS) that was derived from 28,990 CC strings each manually coded by a physician into a syndrome
category [
          <xref ref-type="bibr" rid="ref13">16</xref>
          ]. Although CoCo has the capability to be retrained by using patient data obtained locally, CoCo,
in our experiments, was used with the default probability le and no further training was performed.
        </p>
        <p>SyCo di ers from CoCo in that it implements a two-layer classi cation procedure from chief complaints into
syndromes. SyCo rst nds the posterior probabilities of a number of symptoms (i.e., 17 in-built symptoms)
given the text of the chief complaint. Consequently, SyCo calculates the posterior probabilities of each syndromic
group given the (posterior) probabilities of the symptoms. The syndrome with the highest posterior probability
is then selected as the syndromic group for the chief complaint.</p>
        <p>
          The textual features that CoCo and SyCo extract from CCs are at the word level only, excluding phrases,
n-grams, or any biomedical terminology. The study conducted by Connor et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] showed that the RODS
CoCo classi er is outperformed by using a Maximum Entropy Model that was able to overcome the word-level
approach of CoCo by considering both sub-word and super-word sequences of characters. The MaxEnt classi er
also improved over CoCos ignorance of conditional dependence between textual features.
        </p>
        <p>
          Silva et al. [
          <xref ref-type="bibr" rid="ref14">18</xref>
          ] compared the performance of syndromic classi cation of CCs achieved using SyCo and CoCo
(as well as with their Geographic Utilization of Arti cial Intelligence in Real-time for Disease Identi cation
and Alert Noti cation System that is not available for us to test); however, this analysis was only performed
on a single disease (i.e., In uenza Like Illness ) with a small data set of 1,122 CC records from a single urban
academic medical centre. They found similar recall (0.3530) for SyCo and CoCo on that data but slightly
di erent precision values (CoCo = 0.9890% and SyCo = 0.9930%).
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Victorian Data Set</title>
        <p>The SynSurv data is data collected from the Emergency Departments of the Royal Melbourne Hospital and
the Alfred Hospital during the period 1998-2010 (predominantly from 2000 to 2009). The data were collected
on behalf of the Victorian Department of Health (Vic Health) for syndromic surveillance during the 2006
Commonwealth Games held in Melbourne. The Vic Health remains the custodian of the data; however, for the
purposes of this project, we were granted permission by Vic Health to work with it.</p>
        <p>Data set
Training
Testing</p>
        <p>Syndromic group
Flu Like Illness
Acute Respiratory
Diarrhoea
Other
Total:
Flu Like Illness
Acute Respiratory
Diarrhoea
Other
Total:</p>
        <p>SynSurv consists of naturally occurring data collected at triage, consisting primarily of free text notes written
by a triage nurse for each patient visit to the EDs of the selected hospitals. These notes have been augmented
with annotations of diagnostic codes using the ICD-10 version of the International Classi cation of Disease.</p>
        <p>The data set contains a total of 918,330 ED visit records. From these records, there are 730,054 visits with
a valid ICD-10 code with the primary diagnosis as entered by a nurse upon triage assessment of the patient. In
SynSurv, there are 456,213 records with any textual comment at all. The total number of records with both
a diagnosis and a textual nurse note is 316,362 entries. From this data set of 316,362 CC records, we used a
subset of 314,629 records that were labelled with one of the three syndromic groups Flu Like Illness, Acute
Respiratory, and Diarrhoea, or as other. Table 1 summarizes the distribution of the di erent syndromic groups
in the data set.
3
3.1</p>
      </sec>
      <sec id="sec-2-3">
        <title>Experimental Set-up</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Supervised Chief Complaint Analysis with SyCo and CoCo</title>
      <p>For the machine learning-based classi cation experiments that follow, the set of records corresponding to a
given syndromic group was used as positive examples for that syndrome and all other records in the relevant
subset of SynSurv (see Table 1) were considered to be negative instances for the syndrome.</p>
      <p>We trained one binary SyCo classi er for each of the syndromic groups with the training portion of the data
set for the given syndromic group and then, tested the e ectiveness of the classi er using the corresponding test
set for the syndromic group.</p>
      <p>To evaluate CoCo, which we did not train with our training data sets, we only used the test sets to measure
the classi cation performance of the tool for each syndromic group. For this, we had to nd a mapping between
our syndromic groups and those eight syndromic categories that CoCo had been trained with. Table 2 shows
the mapping that we considered between the two categories of syndromic groups.</p>
      <p>To understand how the performances of SyCo and CoCo compare with those of trivial baseline systems, we
developed two trivial baseline classi ers:</p>
      <p>All-Positive (All+) Baseline, which assigns a positive class label to every given instance of CCs in the
test set. A positive class label means that the chief complaint is labelled as the given syndrome under
investigation.</p>
      <p>Random-50-50 Baseline, which assigns either a positive or a negative class label to every given instance
of CCs in the test set. The assignment of class labels is based on a binary random generator. To more
accurately account for the randomness of this baseline system, we generated average evaluation results for
10 consecutive independent runs over the same test set related to each syndromic group.</p>
      <p>Syndromic group in Victorian data
Flu Like Illness
Acute Respiratory
Diarrhoea</p>
      <p>Syndromic group in CoCo
Constitutional
Respiratory
Gastrointestinal
We evaluated the classi cation methods using the number of True Positives (TPs), False Positives (FPs),
Precision (i.e., speci city), Recall (i.e., sensitivity), and F1-measure. Table 3 summarizes the results achieved
using SyCo and CoCo as well as the two baseline systems All+ and Random-50-50 on the above-mentioned
data set.
SyCo
All+
Flu Like Illness
Acute Respiratory
Diarrhoea
Flu Like Illness
Acute Respiratory
Diarrhoea</p>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>From the results in Table 3, and considering F1-measure, from the two chief complaint classi cation systems,
SyCo has been demonstrated to perform better on all of the three syndromic groups. In terms of precision and
recall, SyCo outperforms CoCo again in most cases, except for the Acute Respiratory group where CoCo shows
a higher precision (0.3603 vs. 0.2483). This single scenario where CoCo outperforms SyCo corresponds to the
situation in which SyCo returns with a large number of FPs (5,845) versus only 1,843 FPs returned by CoCo
for the same syndromic group. This could not be compensated for by the relatively small di erence between
the numbers of TPs that both methods returned (CoCo: 1,038 vs. SyCo 1,931). One major factor that may
explain the superior performance of SyCo over that of CoCo in our experiments is the fact that CoCo was not
trained with our training data sets whereas each SyCo binary classi er was trained for each of the syndromic
groups with the corresponding training set in our data set.</p>
      <p>
        According to the results in Table 3, SyCo and CoCo result in a large di erence in the recall values for the
Flu Like Illness category (CoCo=0.1561 and SyCo=0.5378). These recall values are both relatively di erent
to what Silva et al. [
        <xref ref-type="bibr" rid="ref14">18</xref>
        ] found on the similar disease category In uenza Like Illness with 1,122 CC entries
(SyCo=CoCo=0.3530). The higher recall value of SyCo here compared with what Silva et al. [
        <xref ref-type="bibr" rid="ref14">18</xref>
        ] demonstrated
may be due to the larger training data set that has been utilised in this work.
      </p>
      <p>The results in Table 3 also demonstrate that: i) neither of the two baseline classi cation systems perform
well on any of the three syndromic groups, and ii) more importantly, both SyCo and CoCo perform relatively
higher than the trivial baseline systems.</p>
      <p>To understand how SyCo and CoCo classify each instance of the CCs in the SynSurv data set, we conducted
a Cohen's Kappa statistical agreement analysis on the output classi cations of the two systems. For this, a
pair of output classi cations for each CC string was created and the entire set of pairs were analysed for their
agreement with each other.</p>
      <p>The relatively weak agreement between SyCo and CoCo on our data suggest that the two classi er systems
model and interpret the CCs in relatively di erent fashions. As a result, there is a dissimilar, possibly
complementary, coverage over true positives and true negatives. A possible next step would be to bring the two
classi ers together using an ensemble approach.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>Syndromic surveillance is a widely-utilised procedure for early detection of salient disease outbreaks. One of the
most commonly used techniques in the area of syndromic surveillance is based on supervised classi cation of
Chief Complaints into a set of pre-de ned syndromic categories. In this paper, we analysed the performance of
two well-known CC classi ers from the Real-time Outbreak and Disease Surveillance (RODS), namely Symptom
Coder (SyCo) and Complaint Coder (CoCo). While CoCo was used as an o -the-shelf component, with no
training in our experiments, SyCo was trained with the training subsets of the data we used in this work. The
results of our analysis on a large data set of Australian ED notes labelled with three syndromic groups Flu Like
Illness, Acute Respiratory, and Diarrhoea suggest that, in most cases, SyCo outperforms CoCo in terms of the
evaluation metrics. Both SyCo and CoCo outperform the two trivial baseline systems that we developed for
performance comparison reasons only. We also found that the two classi ers do not agree on the classi cation
outputs, i.e., the classi ers make di erent mistakes, which may suggest that an e ective approach may be
required to combine the two classi ers.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This work was supported by the Bioterrorism Preparedness task of the Land Personnel Protection Branch, Land
Division of the Australian Defence Science and Technology Organisation (DSTO). We also thank the Victorian
Department of Health and Human Services for their contribution of the SynSurv data set. This research was
conducted under Ethics Approval 21/14 by the Department of Health Human Research Ethics Committee.
[9] International Society for Disease Surveillance. Final recommendation: Core processes and EHR
requirements for public health syndromic surveillance. Technical report, ISDS, 2011. URL
www.syndromic.org/projects/meaningful-use.
[10] C.B. Irvin, P.P. Nouhan, and K. Rice. Syndromic analysis of computerized emergency department
patients chief complaints: An opportunity for bioterrorism and in uenza surveillance. Annals of Emergency
Medicine, 41(4):447{452, 2003.
[11] Oleg Ivanov, Per H. Gesteland, William Hogan, Michael B. Mundor , and Michael M. Wagner. Detection
of pediatric respiratory and gastrointestinal outbreaks from free-text chief complaints. AMIA Annual
Symposium proceedings, pages 318{322, 2003. ISSN 1942-597X.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Abir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mostashari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Atwal</surname>
          </string-name>
          , and
          <string-name>
            <surname>Lurie</surname>
            <given-names>N.</given-names>
          </string-name>
          <article-title>Electronic health records critical in the aftermath of disasters</article-title>
          .
          <source>Prehospital and Disaster Medicine</source>
          ,
          <volume>27</volume>
          (
          <issue>6</issue>
          ):
          <volume>620</volume>
          {
          <fpage>622</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Alexandra</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Bambrick</surname>
          </string-name>
          , Dina B.
          <string-name>
            <surname>Passman</surname>
          </string-name>
          ,
          <string-name>
            <surname>Rachel M. Torman</surname>
          </string-name>
          , Alicia A.
          <string-name>
            <surname>Livinski</surname>
          </string-name>
          , and
          <string-name>
            <surname>Jennifer</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Olsen</surname>
          </string-name>
          .
          <article-title>Optimizing the use of chief complaint &amp; diagnosis for operational decision making: An EMR case study of the 2010 Haiti earthquake</article-title>
          .
          <source>PLoS Currents</source>
          ,
          <volume>6</volume>
          :1{
          <fpage>18</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Kevin</surname>
            <given-names>H.O.</given-names>
          </string-name>
          <string-name>
            <surname>Connor</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kieran M. Moore</surname>
            ,
            <given-names>Bronwen</given-names>
          </string-name>
          <string-name>
            <surname>Edgar</surname>
            , and
            <given-names>Don</given-names>
          </string-name>
          <string-name>
            <surname>Mcguinness</surname>
          </string-name>
          .
          <article-title>Maximum entropy models in chief complaint classi cation</article-title>
          .
          <source>In International Society for Disease Surveillance Conference</source>
          , page p.
          <volume>23</volume>
          ,
          <string-name>
            <surname>Baltimore</surname>
          </string-name>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Mike</given-names>
            <surname>Conway</surname>
          </string-name>
          ,
          <string-name>
            <given-names>John N.</given-names>
            <surname>Dowling</surname>
          </string-name>
          , and
          <string-name>
            <surname>Wendy</surname>
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Chapman</surname>
          </string-name>
          .
          <article-title>Using chief complaints for syndromic surveillance: A review of chief complaint based classi</article-title>
          ers in North America,
          <year>2013</year>
          . ISSN 15320464.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Jeremy</surname>
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Espino</surname>
            , John Dowling, John Levander,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Sutovsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>Michael M. Wagner</surname>
            ,
            <given-names>and Gregory F.</given-names>
          </string-name>
          <string-name>
            <surname>Cooper</surname>
          </string-name>
          .
          <article-title>SyCo: A probabilistic machine learning method for classifying chief complaints into symptom and syndrome categories</article-title>
          .
          <source>Advances in Disease Surveillance</source>
          ,
          <volume>2</volume>
          (
          <issue>5</issue>
          ),
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.U.</given-names>
            <surname>Espino</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.M. Wagner</surname>
            ,
            <given-names>F.C.</given-names>
          </string-name>
          <string-name>
            <surname>Tsui</surname>
            ,
            <given-names>H.D.</given-names>
          </string-name>
          <string-name>
            <surname>Su</surname>
            ,
            <given-names>R.T.</given-names>
          </string-name>
          <string-name>
            <surname>Olszewski</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Lie</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>Z.W.</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
            , and
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Dara</surname>
          </string-name>
          .
          <article-title>The RODS open source project: Removing a barrier to syndromic surveillance</article-title>
          .
          <source>Medinfo</source>
          ,
          <volume>11</volume>
          (Pt 2):
          <volume>1192</volume>
          {
          <fpage>1196</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Aaron</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Fleischauer</surname>
            , Stacy Young, Joshua Mott, and
            <given-names>Raoult</given-names>
          </string-name>
          <string-name>
            <surname>Ratard</surname>
          </string-name>
          .
          <article-title>Disaster surveillance revisited: Passive, active and electronic syndromic surveillance during hurricane katrina, New Orleans</article-title>
          ,
          <source>LA 2005. Advances in Disease Surveillance</source>
          ,
          <volume>2</volume>
          :
          <fpage>153</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Kelly</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Henning</surname>
          </string-name>
          .
          <article-title>Overview of syndromic surveillance: What is syndromic surveillance?</article-title>
          <source>Mortality Weekly Report (MMWR)</source>
          , pages
          <fpage>53</fpage>
          (
          <issue>suppl</issue>
          ):
          <volume>5</volume>
          {
          <fpage>11</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Yacine</surname>
            <given-names>Jernite</given-names>
          </string-name>
          , Yoni Halpern, and
          <string-name>
            <given-names>Steven</given-names>
            <surname>Horng</surname>
          </string-name>
          .
          <article-title>Predicting chief complaints at triage time in the emergency department</article-title>
          .
          <source>In Machine Learning for Clinical Data Analysis and Healthcare NIPS Workshop</source>
          <year>2013</year>
          , pages
          <issue>1{5</issue>
          ,
          <string-name>
            <surname>Lake</surname>
            <given-names>Tahoe</given-names>
          </string-name>
          , Nevada,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.R.</given-names>
            <surname>Landis</surname>
          </string-name>
          and
          <string-name>
            <surname>G.G. Koch.</surname>
          </string-name>
          <article-title>The measurement of observer agreement for categorical data</article-title>
          .
          <source>Biometrics</source>
          ,
          <volume>33</volume>
          :
          <fpage>159</fpage>
          {
          <fpage>174</fpage>
          ,
          <year>1977</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Hsin-Min</surname>
            <given-names>Lu</given-names>
          </string-name>
          , Daniel Zeng, Lea Trujillo, Ken Komatsu, and
          <string-name>
            <given-names>Hsinchun</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Ontology-enhanced automatic chief complaint classi cation for syndromic surveillance</article-title>
          .
          <source>Journal of biomedical informatics</source>
          ,
          <volume>41</volume>
          (
          <issue>2</issue>
          ):
          <volume>340</volume>
          {
          <fpage>56</fpage>
          ,
          <string-name>
            <surname>April</surname>
          </string-name>
          <year>2008</year>
          . ISSN 1532-
          <fpage>0480</fpage>
          . doi:
          <volume>10</volume>
          .1016/j.jbi.
          <year>2007</year>
          .
          <volume>08</volume>
          .009.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Tomi</given-names>
            <surname>Malmstr</surname>
          </string-name>
          <article-title>om, Olli Huuskonen, Paulus Torkki, and Raija Malmstrom. Structured classi cation for ED presenting complaints - from free text eld-based approach to ICPC-2 ED application</article-title>
          .
          <source>Scandinavian journal of trauma, resuscitation and emergency medicine</source>
          ,
          <volume>20</volume>
          (
          <issue>76</issue>
          ),
          <year>January 2012</year>
          . ISSN 1757-
          <fpage>7241</fpage>
          . doi:
          <volume>10</volume>
          .1186/
          <fpage>1757</fpage>
          -7241-20-76.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.A.</given-names>
            <surname>Mikosz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Black</surname>
          </string-name>
          , G. Gibbs,
          <string-name>
            <surname>and I. Cardenas.</surname>
          </string-name>
          <article-title>department-based free-text chief-complaint coding systems</article-title>
          .
          <source>(MMWR)</source>
          , pages
          <fpage>53</fpage>
          (
          <issue>suppl</issue>
          ):
          <volume>101</volume>
          {
          <fpage>105</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Julio</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Silva</surname>
            , Shital C. Shah, Dino P. Rumoro,
            <given-names>Jamil D.</given-names>
          </string-name>
          <string-name>
            <surname>Bayram</surname>
          </string-name>
          ,
          <string-name>
            <surname>Marilyn M. Hallock</surname>
            , Gillian S. Gibbs, and
            <given-names>Michael J.</given-names>
          </string-name>
          <string-name>
            <surname>Waddell</surname>
          </string-name>
          .
          <article-title>Comparing the accuracy of syndrome surveillance systems in detecting in uenza-like illness: GUARDIAN vs</article-title>
          .
          <article-title>RODS vs. electronic medical record reports</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          ,
          <volume>59</volume>
          :
          <fpage>169</fpage>
          {
          <fpage>174</fpage>
          ,
          <year>2013</year>
          . ISSN 09333657. doi:
          <volume>10</volume>
          .1016/j.artmed.
          <year>2013</year>
          .
          <volume>09</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Debbie</surname>
            <given-names>Travers</given-names>
          </string-name>
          , Stephanie W. Haas, Anna E. Waller,
          <string-name>
            <given-names>Todd A.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          , Javed Mostafa, Nakia C. Best,
          <string-name>
            <given-names>and John</given-names>
            <surname>Crouch</surname>
          </string-name>
          .
          <article-title>Implementation of emergency medical text classi er for syndromic surveillance</article-title>
          .
          <source>AMIA Annual Symposium proceedings</source>
          ,
          <year>2013</year>
          :
          <volume>1365</volume>
          {
          <fpage>74</fpage>
          ,
          <year>January 2013</year>
          .
          <article-title>ISSN 1942-597X.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>