<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a rule-based support system for the coding of health conditions in the Patient Summary</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elena Cardillo</string-name>
          <email>elena.cardillo@iit.cnr.it</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maria Teresa Chiaravalloti</string-name>
          <email>chiaravalloti@icar.cnr.it</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudio Eccher</string-name>
          <email>cleccher@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Erika Pasceri</string-name>
          <email>erika.pasceri@iit.cnr.it</email>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vincenzo della Mea</string-name>
          <email>vincenzo.dellamea@uniud.it</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lucilla Frattura</string-name>
          <email>lucilla.frattura@regione.fvg.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Guarasci</string-name>
          <email>roberto.guarasci@unical.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bruno Kessler Foundation</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Central Health Directorate, Friuli Venezia Giulia Region</institution>
          ,
          <addr-line>Udine</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Languages and Education Sciences, University of Calabria</institution>
          ,
          <addr-line>Rende (CS)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Department of Mathematics and Computer Science, University of Udine</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Institute for High Performance Computing and Networking</institution>
          ,
          <addr-line>Rende (CS)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Institute of Informatics and Telematics</institution>
          ,
          <addr-line>Rende (CS)</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the frame of federated and interoperable Electronic Health Records (EHRs), specific coding systems are mandatory for filling out healthcare documents such as the Patient Summary (PS). PS cannot be automatically generated from the patient's EHR data, because of the sensitivity of its content. For this reason it needs to be validated by a General Practitioner (GP), who is the sole responsible of this document. The literature shows that the practice of coding is recognized as a difficult task for GPs and it often generates coding errors and misspecifications of clinical data. To overcome this issue, a support system based on standardized and formalized coding rules for the domain of application is proposed, to facilitate a more accurate coding process without breaking the law.</p>
      </abstract>
      <kwd-group>
        <kwd>coding rules</kwd>
        <kwd>patient summary</kwd>
        <kwd>coding support systems</kwd>
        <kwd>reference terminology</kwd>
        <kwd>rule-based systems</kwd>
        <kwd>ICD</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In adopting the European Union (EU) directive on cross-border care and healthcare
semantic interoperability, especially related to the Patient Summary (PS), most
European Countries are regulating the coding systems use in the frame of federated and
interoperable EHRs, making some of them mandatory for compiling healthcare
documents. In compiling the PS, data related to health conditions cannot be automatically
generated from those available in the GP’s EHR, because the GP is fully responsible
for its content and has to validate it. Nonetheless, since coding is proved to be a
difficult task, an automated coding support system (CSS) can be of help without
infringement of the law. The need for a centralized management of coding systems and
processes by means of a rule-based supporting tool is motivated by a number of critical
issues reported in the literature about the use of coding systems at different levels. In
proposing a CSS, it is important to consider the harmonization and integration of
medical terminologies used by domain experts to ensure information interoperability
and the full understanding of the meaning conveyed, thus avoiding the proliferation of
non-integrated and heterogeneous terminologies within EHRs. Furthermore, GPs
massively use natural language to record health conditions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], in particular
comorbidities, thus generating unstructured and uncoded data, mainly because they do not
know how to properly use coding systems and consider coding as an excessively
time-consuming activity. This work proposes a methodology for the creation of a CSS
that will be initially experimented for the Italian PS use case.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related works</title>
      <p>
        During the last twenty years much effort has been spent on the development of
support systems for the semi-automatic editing of healthcare documents with structured
data. Some of these tools have been tested on the coding of causes of death, which are
generally coded from death certificates using the International Classification of
Disease 10th revision (ICD-10). In particular, two software tools have been developed to
help this type of coding: MICAR-ACME [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], developed by US National Center for
Health Statistics, and more recently IRIS [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] developed by a European consortium.
These tools served as the basis for the development of other national support systems
for coding causes of death, such as the Italian one [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. However, since death
certificates already provide structured information, the issue of processing natural language
is relatively trivial, although the mortality coding rules by themselves are complex.
      </p>
      <p>
        In addition, automated coding tools based on Natural Language Processing (NLP)
have been developed [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]. Recently, a recommending system for ICD-10-CM
(Clinical Modification) coding starting from SNOMED CT (Systematized Nomenclature of
Medicine - Clinical Terms) annotated health records has been also developed as a
consequence of the World Health Organization (WHO) – International Health
Terminology Standards Development Organization (IHTSDO) harmonization effort [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Under the same framework, a further, ongoing evolution is the development of a
common ontology between ICD-11 and SNOMED CT [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Nonetheless, only few
systems solve coding tasks using a set of hand crafted expert rules, as in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] which
focused on the ICD-9-CM. Finally new methodologies use ontologies and automated
reasoning to provide and support fast and incremental classification of medical
terminologies or classification systems, as in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] where the Snorocket reasoner has been
developed to support SNOMED CT ultrafast classification.
      </p>
    </sec>
    <sec id="sec-3">
      <title>A coding support system for the Patient Summary</title>
      <p>
        According to the EU Guidelines, the PS is “the minimum set of information needed
to assure healthcare coordination and the continuity of care” [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Member States
adopted them often adding further clinical information, as in Italy, where a Prime
Minister’s Decree (to be issued) contains all the reference elements to be
implemented to allow for interoperability among regional EHR systems, recognizing the critical
role of the PS. According to it, PS reference elements, tagged as mandatory or
optional, can be reported as free text or by using dedicated coding systems. The application
of a CSS to ease the compilation of the PS allows for the selection of recommended
codes to be assigned to the information required within the PS. Because of its highly
structured content, the PS could be well coded using rules, reducing the variability of
natural language free texts to interoperable codes.
      </p>
      <p>
        To implement a challenging automated support system for coding health conditions
in clinical documents such as the PS, a four-step methodology is proposed: (see Fig. 1
for an overview of the process):
 Analysis of the epSOS project1 results and specifications and study of the
automated ICD-10 coding rules for morbidity and comorbidity2, to verify features useful to
guide the automated morbidity, procedures and interventions coding in the PS use
case. This step will produce standardized coding rules based on general guidelines
defined by qualified institutions (e.g. WHO) and described by the literature [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ];
 Design of an algorithm that applies coding rules to produce candidate codes and
assess their accuracy, and implemented in a suitable computable formal language
for representing rules and the domain. The suitability of rule-based languages (e.g.,
OWL + SWRL)3, and Task Network languages for the representation of guidelines
(e.g., ASBRU)4 will be analyzed;
1 epSOS Project: http://www.epsos.eu
2 WHO, ICD-10, vol. 2. Instruction Manual, 2010.
3 SWRL W3C recommendations: http://www.w3.org/Submission/SWRL/
4 http://www.openclinical.org/gmm_asbru.html
 Creation and use of complementary tools to support the transition from the
specialized and natural language used by GPs in their EHRs to the coding language: a
cross reference terminology of structured technical and lay terms, to be used as
intermediate between the natural language and the concepts of the international
coding systems; and finally transcoding tables to manage the different versions and
revisions of a coding system (e.g. ICD-9-CM to ICD-10) or to map between
different systems (e.g. SNOMED CT to ICD-10);
 Composition of the abovementioned tools to build a web service-based CSS.
      </p>
      <p>The accuracy assessment of candidate codes proposed by the CSS is up to the GP,
who, as mentioned above, has the full responsibility of PS clinical content.</p>
      <p>
        In particular, the use of a rule-based approach with respect to other NLP ones (e.g.
Support Vector Machines and Hidden Markov models) allows a better translation of
rules for coding patient summaries using ICD9-CM. Those rules, similar to those
defined by WHO for coding mortality, should be made explicit and translated also in
a computable way. Furthermore, the creation of the cross reference terminology is
based on existing terminological tools, such as the ICD-10 Alphabetical Index5;
consumer-oriented medical vocabularies (e.g., the ICMV [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]); ICD-11 narrower terms6
that include synonyms and quasi-synonyms; and a dictionary for NLP, created from a
database of 295,000 EHRs [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Finally, a mapping to some major standardized coding
systems will be performed. See Fig. 2 for an example of clinical data coding in the PS
using the CSS.
      </p>
      <p>In the example above it is shown as the clinical information in input written in
natural language by the GP is processed in order to arrive, as output, with the correct
coding to be included in the PS. Another coding example is shown in Fig.3:
5 WHO, ICD-10, vol. 3. Alphabetical Index, Italian version, 2014.
6 ICD-11 beta draft available at: http://apps.who.int/classifications/icd11/browse/f/en</p>
      <p>
        This use case shows how the CSS could filter the clinical information contained in
GPs’ database in order to include in the PS only the relevant information, (i.e.
congiuntivite cronica “chronic conjunctivitis”) those referring to a chronic disease, not only
to single events (as “mal di testa” or “dolori addominali”). The PS, as stated above,
contains only a standardized set of basic medical data including only the most
important clinical facts required to ensure safe and secure healthcare [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Discussion and Conclusions</title>
      <p>The present paper aims at proposing an experimental methodology for the
development of a rule-based CSS, with an initial experimentation in Italy, that will allow to
develop: i) a web service to directly support natural language text coding, and ii) a set
of rules in an open format, to be embedded also in third-party software.</p>
      <p>With respect to the limitations produced by manual coding, the use of a sound
rulebased CSS presents consistent advantages (that are common to rule-based NLP
methods): (i) it requires the adoption of internationally updated standard systems and
standardized methodology for the accurate coding of health conditions; (ii) it could
significantly reduce the coding time and costs, requiring GPs interaction only to
choose among and validate the recommended codes; (iii) it improves the quality of
coding by reducing the variability due to different subjective interpretations,
especially in the case of comorbidities. By using the proposed CSS it is also possible to
measure how often the GP is able to find the right candidates codes among those suggested
by the system. On the other hand, some limitations need to be considered, mainly
related to the complexity of the domain: (i) it may be necessary to formalize a huge
amount of rules to represent the number of possible situations; (ii) maintenance and
updating of the knowledge base and computational costs of the system can be high
with the risk of inefficiency.</p>
      <p>This pilot study will allow considerations also on the economic aspects, to be
compared with the cost of training and maintaining up to date the large number of
GPs that need to write and code patient summaries.</p>
      <p>Although developed for the Italian PS, the proposed methodology could be further
adapted to other European Countries.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work is supported by the following projects: “Realizzazione di Servizi della
Infrastruttura Nazionale per l’Interoperabilità per il Fascicolo Sanitario Elettronico”
(prot. 7626) funded by the Agency for Digital Italy (AgID) and by the
SemanticHealthNet Expert Agreement (prot. 0008991).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cardillo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiaravalloti</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasceri</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Assessing ICD-9-CM and ICPC-2 Use in Primary Care. An Italian Case Study</article-title>
          . In: Kotsova,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Grasso</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.),
          <source>Digital Health 2015 (DH '15)</source>
          , pp.
          <fpage>95</fpage>
          -
          <lpage>102</lpage>
          . ACM New York - USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Israel</surname>
            ,
            <given-names>R. A.</given-names>
          </string-name>
          :
          <article-title>Automation of mortality data coding and processing in the United States of America</article-title>
          .
          <source>World Health Stat Q</source>
          .
          <volume>43</volume>
          (
          <issue>4</issue>
          ):
          <fpage>259</fpage>
          -
          <lpage>62</lpage>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Johansson</surname>
            ,
            <given-names>L. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pavillon</surname>
          </string-name>
          , G.:
          <article-title>IRIS: A language-independent coding system based on the NCHS system MMDS</article-title>
          .
          <article-title>In: WHO-FIC NETWORK MEETING (</article-title>
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Istituto</surname>
          </string-name>
          <article-title>Nazionale di Statistica - ISTAT: Metodi e software per la codifica automatica e assistita dei dati</article-title>
          .
          <source>Tecniche e Strumenti. N.4</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chiaravalloti</surname>
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guarasci</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lagani</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pasceri</surname>
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trunfio</surname>
            <given-names>R.:</given-names>
          </string-name>
          <article-title>A Coding Support System for the ICD-9-CM standard</article-title>
          ,
          <source>In: the IEEE International Conference on Healthcare Informatics (ICHI2014)</source>
          , pp.
          <fpage>71</fpage>
          -
          <lpage>78</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Friedman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shagina</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lussier</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hripcsak</surname>
          </string-name>
          , G.:
          <source>Automated Encoding of Clinical Documents Based on Natural Language Processing. JAMIA</source>
          .
          <volume>11</volume>
          :
          <fpage>392</fpage>
          -
          <lpage>402</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Campbell</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brear</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scichilone</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>White</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giannangelo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carlsen</surname>
            ,
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solbrig</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fung</surname>
            ,
            <given-names>K.W.</given-names>
          </string-name>
          :
          <article-title>Semantic interoperation and electronic health records: context sensitive mapping from SNOMED CT to ICD-10</article-title>
          . Stud Health Technol Inform.
          <volume>192</volume>
          :
          <fpage>603</fpage>
          -
          <lpage>7</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Rodrigues</surname>
            <given-names>J. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schulz</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rector</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spackman</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Üstün</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chute</surname>
            <given-names>C. G.</given-names>
          </string-name>
          , Della Mea V.,
          <string-name>
            <surname>Millar</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Persson K</surname>
          </string-name>
          . B.:
          <article-title>Sharing ontology between ICD 11 and SNOMED CT will enable seamless re-use and semantic interoperability</article-title>
          .
          <source>Stud Health Technol Inform</source>
          .
          <volume>192</volume>
          :
          <fpage>343</fpage>
          -
          <lpage>6</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Farkas</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szarvas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Automatic construction of rule-based ICD-9-CM coding systems</article-title>
          .
          <source>BMC Bioinformatics</source>
          .
          <volume>9</volume>
          (
          <issue>Suppl 3</issue>
          ):S10 (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Metke-Jimenez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lawley</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Snorocket 2.0: Concrete Domains and Concurrent Classification</article-title>
          .
          <source>In: the OWL Reasoner Evaluation Workshop (ORE</source>
          <year>2013</year>
          ), pp.
          <fpage>32</fpage>
          -
          <lpage>38</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <article-title>eHealth Network of the European Union: Guidelines on minimum/non-exhaustive patient summary dataset for electronic exchange in accordance with the cross-border directive</article-title>
          <year>2011</year>
          /24/EU (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Frattura</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gongolo</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Munari</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Identification and coding of the main condition using ICD: suggested workflows</article-title>
          .
          <source>In: WHOFIC NETWORK Annual Meeting</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Cardillo</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamilin</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Serafini</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>A Methodology for Knowledge Acquisition in Consumer-Oriented Healthcare</article-title>
          . In:
          <article-title>Knowledge Discovery, Knowledge Engineering and Knowledge Management Communications in Computer and Information Science</article-title>
          , Vol.
          <volume>128</volume>
          , pp.
          <fpage>249</fpage>
          -
          <lpage>261</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>