<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Ontologies in support of data mining based on associated rules: a case study in a diagnostic medicine company</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lucélia P. Branquinho</string-name>
          <email>luceliabranquinho@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maurício B. Almeida</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renata M.A. Baracho</string-name>
          <email>renatabaracho@eci.ufmg.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Information Science, Federal University of Minas Gerais (UFMG) - Belo Horizonte - MG -</institution>
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>A well-known alternative to identify hidden standards is the use of data mining techniques. In order to obtain more efficiency in data mining, ontologies have been used to improve the representation in specialized knowledge domains. Here, we apply ontologies in a dataset of a diagnostic medicine company, which concerns to viral human hepatitis, with the aim of obtaining the best correlations between the laboratory tests prescribed by physicians and the real occurrences of diseases. Our preliminary findings show that the use of ontologies provides reduction in the number of attributes in the pre-processing phase, then improving the performance of data mining process as a whole.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>generalization and the consequent reduction of the number of attributes to be mined via
the identification of similarities between laboratories tests, considering the relations
mapped in the viral hepatitis ontology called HVO.</p>
      <p>The next sections will be organized as follows: section 2.1 provides a brief
description on the use of ontologies in data mining; section 2.2, describes the
construction of prototype of a viral hepatitis ontology; section 2.3 explains how
generalization of terms will be applied in the data mining pre-processing phase as per
association rules; section 3 details the results obtained with the proposed model. Finally,
section 4 showcases final considerations.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Method</title>
      <p>
        In some fields, such as Biomedicine, specialized communities have been developing and
publishing, since the 1990s, a series of ontologies to aid in representing and retrieving
informational [
        <xref ref-type="bibr" rid="ref5">Perez-Rey et al. 2004</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>2.1. Ontologies and data mining</title>
      <p>Knowledge extraction, generally referenced in literature as Knowledge Discovery in
Database (KDD) should be grouped into three phases: pre-processing, DM and
postprocessing. Pre-processing, which is relevant for our goals in this paper, comprises the
collection, organization and treatment of data, while DM involves algorithms and
techniques to search for knowledge.</p>
      <p>Ontologies have been used to increase the relevance of the patterns discovered
through the mining techniques. One of the techniques in which ontologies are being
utilized is mining through association rules, which display the correlation between sets
of items in series of data or transactions. [Ferraz 2008].</p>
      <p>The advantage of pruning restrictions is to exclude information in which users
are not interested in since the beginning. Every general rule should be able to replace a
number of specific rules by means of generalization processes. Whenever this approach
is feasible, a semantic improvement of the mined association rules and a reduction in the
cardinality of the set rules will simultaneously take place.</p>
    </sec>
    <sec id="sec-4">
      <title>2.2. Viral Hepatitis Ontology Construction</title>
      <p>
        The development of the viral hepatitis ontology was based mapped clinical analysis
laboratory tests for diagnosing human viral hepatitis considering LOINC, OGMS
ontology [
        <xref ref-type="bibr" rid="ref7">Scheuermann et al 2009</xref>
        ], IDO ontology [
        <xref ref-type="bibr" rid="ref1">Cowell and Smith 2006</xref>
        ], DOID
ontology [Lynn Schriml, 2009] and FMA ontology [FMA 2012]. It describes the clinical
picture throughout the disease cycle by mapping terminological items that encompass
diseases, their causes, their manifestations and diagnosis.
      </p>
      <p>We associate the laboratory tests with the viral infectious disease to enable the
generalization of the attributes to be mined, as proposed in Figure 1. In the triage phase
of Hepatitis C, for example, four specific tests may be requested for the virus
identification and another eight unspecific tests may be ordered for monitoring liver
functions. This situation may be generalized without having denominated each
laboratory tests as an attribute for data mining.</p>
      <p>The ontology for hepatitis was created at this moment of our experiment so that
we could test its application in our computational architecture. We are aware that some
improvements in modeling are in order, for example: i) "has symptom" and "is
observed" are not ontological relations; ii) "An axioms like Hepatitis C subclass Of
hasSymptom some Jaundice can be falsified by one single patient who has Hepatitis C
but no jaundice"; iii) instead of "no symptom, and following OGMS, we should think in
use "healthy organism". Such improvements which will be part of our future work in the
following of the research.</p>
    </sec>
    <sec id="sec-5">
      <title>2.3. Generalization of terms in the data mining pre-processing phase</title>
      <p>Our study makes use of ontologies, reasoners and Jena software to promote pruning and
filtering (generalization) of data from the list of laboratory tests collected from the
diagnostic medicine company’s database. When analyzing the relationships shared by
the terms, one might identify which laboratory tests are related to which disease and
stage. Therefore, the similarity between terms is considered as a means to generalize the
attributes in the pre-processing phase. Figure 2 depicts the proposed model.</p>
      <p>Based on HVO ontology, we consider relationships with the disease and with its
features and also utilizing the Jena6 tool, as well as inference rules. So, we were able to
obtain more general terms to represent laboratory test groups associated with viral
hepatitis diagnosis.</p>
    </sec>
    <sec id="sec-6">
      <title>3. Results</title>
      <p>The development of ontologies along with the use of inference mechanisms during the
pre-mining phase has reduced the number of attributes to be mined by the association
rules algorithm, namely, the Apriori algorithm. It reduced the amount of laboratory tests
related to the direct diagnosis of the hepatitis virus, and also the number of unspecific
tests for assessment over the course of diseases.</p>
      <p>
        Based on knowledge obtained from the development of HVO, a list of laboratory
test orders was collected from the company’s database containing at least one test
directly related to a hypothetical hepatitis virus diagnosis. Laboratory test orders that
complied with the previously described rule were selected during three months, January,
February and March 2015, totaling 34440 occurrences (Table 1). Test applications,
which are complementary to the diagnosis, are distributed in collections made on the
organization's service units, conveyed by laboratory partners throughout Brazil. In this
sample, the occurrences featured 465 different laboratory tests. Considering the data of
service units (support = 0.2), laboratory patterns (support = 0.02) and confidence 0.75
was executed the ARules package [
        <xref ref-type="bibr" rid="ref3">Hasher 2007</xref>
        ] in R Language to extract the
association rules, we obtained the results presented below.
6 Available: &lt; https://jena.apache.org/&gt;.
      </p>
      <p>Considering the same database obtained by reduction of the number of attributes,
with the use of ontology, applied to 439 different laboratory tests and with the same
support value and confidence was executed again the algorithm Apriori . For base units
were obtained 258 and 115 rules for laboratory partners, we reached a reduction of 50%
of the resulting association rules, as shown in table 2.</p>
      <p>The attributes were categorized considering the relationship between the
modeled tests through equivalence axioms as showed in the example below:
hvo:laboratory_diagnostic_process_ hepatitis_A equivalent to: hvo:laboratory_testing_encounter
And (is_composed_of some (‘Laboratory test'
and (diagnoses only 'hepatitis A'))) and (is_composed_of min 0 ('laboratory test'
and (diagnostic_evaluation some Liver)))</p>
      <p>In this equivalence axiom (described by existential restrictions), a part of the
detailed diagnosis of the disease process is comprised of at least one medical application
(in this case Class HVO: laboratory_testing_encounter), which is composed of
complementary examinations (OGMS : laboratory test) for disease diagnostic (doid:
hepatitis a) and can also be a laboratory test for evaluating the state of the health (HVO:
diagnostic evaluation) of a liver (FMA: liver).</p>
      <p>Considering the limitation of further tests and the disease is possible to identify
relationships between them and, thus, promote the generalization and its representation
in single attribute, in this case, a type of viral hepatitis, as shown in Table 3.</p>
      <p>Table 3 shows the laboratory tests (LT) prescribed to patients by the doctor and
sent to the medical diagnostic laboratory. With the generalization attributes through the
method was represented axiom disease in which some complementary tests are
associated, in this case, a type of hepatitis. The other complementary tests were
maintained and makes up the list of attributes analyzed by mining technique Apriori
which extracted association rules related to viral hepatitis.</p>
      <p>Therefore, our findings suggested that it is possible to reduce the rules resulting
from data mining by reducing the possibilities of combining attributes. As a second, we
found that the generalization of the terms enables results with a greater significance,
since it can guide the post-mining phase analysis process.</p>
    </sec>
    <sec id="sec-7">
      <title>4. Discussion and conclusions</title>
      <p>The application ontology developed here strongly represents the LOINC tests for viral
hepatitis, as we understand that this classification is sufficient for assessing the
relationship of association rules. The extension of OGMS, DOID, FMA and IDO brings
7Identification of the laboratory test
8LOINC 1834-1 - Alpha-1 Fetoprotein – Laboratory test unspecified viral hepatitis
9 LOINC 61151-7 - Albumin – Laboratory test unspecified viral hepatitis
10 LOINC 5196-1 - Hepatitis B virus surface Ag
11 LOINC 1742-6 - Alanine aminotransferase – Laboratory test unspecified viral hepatitis
12 LOINC 5179-7 - Hepatitis A virus Ab.IgG
13 LOINC 13950-1 - Hepatitis A virus Ab.IgM
laboratory tests closer to the diagnosis cycle and, consequently, promotes the
identification of the correlations between lab tests.</p>
      <p>It is relevant to highlight two points. Firstly, the relevance of patterns extracted
by means of techniques that identify semantic similarity between terms is highly
dependent of the ontology construction and validation. Therefore, it is fundamental that
the domain ontologies being used be validated by specialists. Secondly, the KDD
approach requires greater reach of the algorithms and cannot be restricted to “is-a” and
“part-of" relations, which reinforces the use of formal semantics of ontologies.</p>
      <p>In future work, it is intended to promote the enrichment of the ontology with
new concepts and equivalence between complementary tests and disease for greater
generalization of attributes.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Cowell</surname>
            ,
            <given-names>L.G</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Infectious Disease Ontology</article-title>
          . In: Infectious Disease Informatics. Sintchenko V, editor. New York: Springer;
          <year>2010</year>
          . pp.
          <fpage>373</fpage>
          -
          <lpage>395</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Ferraz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <article-title>Ontology in association rules</article-title>
          . Available: &lt;http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3786067/&gt;.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Hasher</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hornik</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Grun</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchta</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <article-title>Introduction to arules - A computational environment for mining association rules and frequent item sets</article-title>
          ,
          <year>2007</year>
          . Available: &lt;http://cran.r-project.org/web/packages/arules/vignettes/arules.pdf&gt;.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Manda</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCarthy</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bridges</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <article-title>Interestingness measures and strategies for mining multi-ontology multi-level association rules from gene ontology annotations for the discovery of new GO relationships</article-title>
          . Available: &lt; http://www.ncbi.nlm.nih.gov/pubmed/23850840 &gt;.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Perez-Rey</surname>
            ,
            <given-names>D</given-names>
          </string-name>
          ; Maojo,
          <string-name>
            <given-names>V</given-names>
            ;
            <surname>Garcia-Remesal</surname>
          </string-name>
          , M;
          <string-name>
            <surname>Alonso-Calvo</surname>
            ,
            <given-names>R. Biomedical</given-names>
          </string-name>
          <string-name>
            <surname>Ontologies</surname>
          </string-name>
          .
          <source>In: Proceedings of the 4th IEEE Symposium on Bioinformatics and Bioengineering</source>
          , p.
          <fpage>207</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Piatetsky-Shapiro</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fayyad</surname>
            ,
            <given-names>U. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>From data mining to knowledge discovery: an overview</article-title>
          . Available: &lt;http://dl.acm.org/citation.cfm?id=
          <volume>257942</volume>
          &gt;.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Scheuermann</surname>
            ,
            <given-names>R. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Werner</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <article-title>Toward an Ontological Treatment of Disease</article-title>
          and Diagnosis.
          <year>2009</year>
          . Available: &lt;http://www.ncbi.nlm.nih.gov/pmc/articles/PMC3041577/&gt;.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Schriml</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2008</year>
          ). Available:&lt; http://do-wiki.nubic.northwestern.edu/dowiki/index.php/Main_Page&gt;
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>