<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Crowdsourcing for ICD10 Code to Concept Relationships</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Patrice Seyed</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evan Patton</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Subotin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>3M HIS</institution>
          ,
          <addr-line>Silver Spring, MD</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Computer Science</institution>
          ,
          <addr-line>RPI, Troy, NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this work we leverage crowdsourcing in connection with machine learning techniques to validate candidate ICD10 Code to UMLS concept relationships that we generate. Our immediate use is in natural language understanding and machine learning approaches to automatically code electronic health record documents with ICD codes. Beyond auto-coding, the relationships will aid a wide variety of future medical applications, such as terminology-driven search in support of smart medical assistants.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Initially, EHRs manually annotated with gold standard ICD-10 CM Codes are
also annotated for UMLS concepts, and the code-concept pairs are processed
for pointwise mutual information (PMI) across a corpus [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We consider PMI
scores, as well as concept counts and pair count as thresholds for determining
the set of pairs for submitting for crowdsourcing using Amazon's Mechanical
Turk. The use of similarity (e.g., lexical, semantic) measurements is planned to
assist lter candidates, and as features for training a machine learning model
that will identify the relationship (if it exists) between an arbitrary ICD10 CM
code and concept.
      </p>
      <p>Once we have trained our probabilistic model with crowdsourcing results, we
use the results to predict whether the candidate pairs that have not yet been
crowd-sourced are valid or not (using PMI and other measures as features, and
results of the crowd as labeled data). In the initial phase, we leverage subject
matter experts' knowledge to validate the crowdsourced judgments. Thus the
crowdsourcing data is used for two purposes: training the probabilistic model
and for nal judgments. As similarity measures are developed further within the
project, they can be used as additional features to identify candidates to crowd
source or serve as relationships. Table 1 shows three pairs examples. The overall
work ow is illustrated in Figure 1.</p>
    </sec>
    <sec id="sec-2">
      <title>Designing Questions for the Crowd</title>
      <p>In order to determine how best to formulate natural language questions to ask
as to the direct relationship between a code (i.e., clinical situation) and a
concept, we consider the UMLS semantic type, based primarily on the following:
Disorders, Body Parts, Procedures, and Findings. For example: if the concept is
a disorder, the question options pertain to the taxonomic relation \isa"; if the
concept is a body part, the question pertain to whether it is the nding site of
the disorder; if the concept is a diagnostic procedure, the questions pertain to
whether the procedure is used to diagnose the disorder. Note then, that the
relationships are both taxonomic, compositional, and other direct relations, therefore
supporting structured knowledge source applications. Example questions posed
to Mechanical Turk workers are given in Figure 2.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Evaluation And Future Steps</title>
      <p>
        For evaluating our results, we consider majority for con rming a speci c answer,
and consensus for con rming negation (i.e., none of the above answer) is
accurate. We are in the process of adjusting parameters (pay, questions per task,
quali cation questions, assignment duration, auto-approval). Our next steps
include performing analyses for evaluation techniques for workers, answers, and
question quality. This is useful since disagreement is oftimes signal and not noise
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We aim to increase the volume of results by increasing pay and comparing
results against another crowsourcing platform, CrowdFlower. As for utility, we
will investigate which code-to-concept pairs are directly useful in our rule-based
systems, and which are useful primarily for machine learning approaches for
auto-coding.
5
      </p>
    </sec>
    <sec id="sec-4">
      <title>Vision and Impact</title>
      <p>
        The improved ICD-10 to UMLS relationships generated by our crowd-sourcing
approach will result in less noisy and more robust structured knowledge. It will
enable the use of these relationships are rules as well as evidence for auto-coding
ICD 10 CM. Also, such knowledge will support future medical applications aimed
at aiding practitioners and patients alike. Use of knowledge representation for
building medical expert systems for diagnosis has been well explored (for a
review of techniques, see[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]). The number of structured platforms for
patientoriented smart medical assistants is also growing, for example [
        <xref ref-type="bibr" rid="ref5 ref7">5, 7</xref>
        ]. All of these
applications using structured knowledge approaches stand to bene t from our
crowd-sourced relationships.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Aroyo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Welty</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Measuring crowd truth for medical relation extraction</article-title>
          .
          <source>In: 2013 AAAI Fall Symposium Series</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bodenreider</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>The uni ed medical language system (umls): integrating biomedical terminology</article-title>
          .
          <source>Nucleic acids research 32(suppl 1)</source>
          ,
          <source>D267{D270</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bouma</surname>
          </string-name>
          , G.:
          <article-title>Normalized (pointwise) mutual information in collocation extraction</article-title>
          .
          <source>Proceedings of GSCL</source>
          pp.
          <volume>31</volume>
          {
          <issue>40</issue>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>J.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brear</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scichilone</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giannangelo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carlsen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solbrig</surname>
            ,
            <given-names>H.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fung</surname>
            ,
            <given-names>K.W.</given-names>
          </string-name>
          :
          <article-title>Semantic interoperation and electronic health records: context sensitive mapping from snomed ct to icd-10</article-title>
          . In: MedInfo. pp.
          <volume>603</volume>
          {
          <issue>607</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lewandowski</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arochena</surname>
            ,
            <given-names>H.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naguib</surname>
            ,
            <given-names>R.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chao</surname>
            ,
            <given-names>K.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Perez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Logic-centered architecture for ubiquitous health monitoring</article-title>
          .
          <source>Biomedical and Health Informatics</source>
          ,
          <source>IEEE Journal of 18(5)</source>
          ,
          <volume>1525</volume>
          {
          <fpage>1532</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>S.H.</given-names>
          </string-name>
          :
          <article-title>Expert system methodologies and applications|a decade review from 1995 to 2004</article-title>
          .
          <article-title>Expert systems with applications 28(1</article-title>
          ),
          <volume>93</volume>
          {
          <fpage>103</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Patton</surname>
            ,
            <given-names>E.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGuinness</surname>
            ,
            <given-names>D.L.</given-names>
          </string-name>
          :
          <article-title>Toward next generation integrative semantic health information assistants</article-title>
          .
          <source>In: Proc. AAAI Winter Symp. on Expandings the Boundaries of Health Informatics using Arti cial Intelligence</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>