<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Process model for data mining in health care sector</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Diego Roa - María del Pilar Villamil</string-name>
          <email>df.roa34@uniandes.edu.co</email>
          <email>{df.roa34,mavillam}@uniandes.edu.co</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juan Diego Arboleda Oracle</string-name>
          <email>juan.arboleda.tabares@oracle.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bogota</institution>
          ,
          <country country="CO">Colombia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Los Andes University</institution>
          ,
          <addr-line>Bogota</addr-line>
          ,
          <country country="CO">Colombia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents a process model to guide the data mining projects in the health care sector. The process model (PMH) is a specialization of CRISP-DM methodology and presents di erent issues associated to the data analysis and management. This proposal was validated in order to address real problems related to health care in Colombia. The results show that it is possible to establish new hypothesis about the clinical data, and revalidate these a rmations using the proposed process model.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data mining</kwd>
        <kwd>process model</kwd>
        <kwd>healthcare</kwd>
        <kwd>PMH</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In the health care sector, there are many opportunities to
apply data mining. Some of them are related to the
improvement of the quality control in health care. In
particular, analysis to detect and diagnose diseases, predict the
responses of the organism to speci c treatments and to identify
epidemiological pro les, are relevant themes for the health
care community.</p>
      <p>
        There are many methodologies to tackle data mining
opportunities such as CRISP-DM[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or the virtuous cycle of data
mining [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. All of them are designed to improve the
success of data mining projects. These methodologies are used
in many sectors such as nancial, pharmaceutical or health
care industries. However, there are very speci c
characteristics associated to these sectors that can be used to customize
the process and to improve the quality and e ectiveness of
these kind of projects.
      </p>
      <p>This paper presents a process model to guide the data
mining process in the health care sector. This process model
allows to reduce the costs and resources used in data
mining projects. The process model was evaluated by analyzing
49,000,000 individual register of health care (RIPS) obtained
from di erent sources, such as HMOs and the Minister of
Social Protection in Colombia from 2003 to 2006. The
evaluation goal was to compare treatments among Health
Maintenance Organizations (HMOs), as well as verifying the
compliance to standards of evidence gained from the scienti c
method (EBM- Evidence-based medicine); which will
support the quality process of health services. The analysis
was focused on hypertension diagnosis and allows to
evidence similarities between the national guidelines and the
health service. Speci cally it was possible to identify the
use of captopril an Angiotensin-Converter enzyme inhibitor
medicament in the hypertension treatment. This
medicament is cheaper according to other medicines of this type.
One hypothesis is that it is pre-scripted for economical
reasons. However, patients with this kind of treatment, returns
to the healthcare institution with complications in the
hypertension disease. This kind of complications increase the
illness costs. Finally, validations about this process model
enable new opportunities to establish public health policies
in Colombia.</p>
      <p>This paper is organized as follows. Section 2 describes
problems related to the data mining process in healthcare.
Section 3 presents Health care data management characteristics
and the PMH the Process Model proposed for data mining in
Healthcare sector. Section 4 exposes the validation method
of the proposed process. Finally, Section 5 concludes the
paper and presents other research issues.</p>
    </sec>
    <sec id="sec-2">
      <title>2. DATA MINING AND HEALTH CARE</title>
      <p>There are many studies that evidence the relevance of data
mining techniques in the health care sector. These studies
are associated to the treatment of patients and generally,
to the identi cation of best practices in the treatments of
speci c diseases.</p>
      <p>
        Some works such as [
        <xref ref-type="bibr" rid="ref15 ref16 ref17 ref6 ref9">6, 16, 17, 15, 9</xref>
        ] show di erent categories
of problems related to the health care sector that are solved
using data mining techniques. Some of them presents the use
of association rules, sequential patterns or clustering in the
prediction, analysis and monitoring of patient's treatments.
On the other hand, there are studies such as [
        <xref ref-type="bibr" rid="ref1 ref7">1, 7</xref>
        ] that
propose new algorithms to solve these issues.
      </p>
      <p>
        Abidi and Stolba in [
        <xref ref-type="bibr" rid="ref10 ref14">14, 10</xref>
        ] describe the relevance of
identifying clinical guides based on the individual registers of
health care. These guides support medical tasks,
increasing the quality of service of medical centers. However, these
works are focuses in structuring clinical guidelines, and they
do not emphasize in the methodology used to realize these
projects.
      </p>
      <p>
        Although some of the works such as [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] use the CRISP-DM
methodology and others [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] the virtuous cycle data mining
to improve the process quality, there are some
characteristics associated to the speci c domain that will be used to
reduce the number of incidentals that may arise in a data
mining project. These issues enhance the opportunity to
use methodologies that explicitly include clinical concepts,
problems related to the health care sector and moreover,
that supports the selection process of data mining techniques
based on the speci c characteristics of this sector.
The afore mentioned issues motivate the realization of PMH,
a speci c process model for health care context, that will be
presented in the section 3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. DATA MINING GUIDE FROM THE HEALTH</title>
    </sec>
    <sec id="sec-4">
      <title>CARE POINT OF VIEW</title>
      <p>This section presents a brief description about health care
context in subsection 3.1, with the purpose to highlight the
opportunity to provide new guides associated to data mining
applications to improve di erent kind of projects in health
care. Subsection 3.2 describes a new process model to guide
these projects.</p>
    </sec>
    <sec id="sec-5">
      <title>3.1 Health care data management</title>
      <p>The health care sector evidences challenges w.r.t. data
management because of data characteristics such as volume,
quality, availability, accessibility, and the relevance of the
decisions involved during the process analysis.</p>
      <p>Moreover, there are clinical guidelines that describe a set of
steps to treat a speci c disease. From these guides, it is
possible to determine the service e ciency, the time involved in
the treatment, and typical practices such as treatments and
medications. This kind of information provides important
elements to the decision's maker.</p>
      <p>According to the volume of data analyzed, it is important to
highlight that all clinical cases are relevant during process
analysis. This is contrary to other sectors. In other domains
a rule is meaningful when the support of the data is relatively
high, whereas in the health care domain, analysis involving
mortalities will be accounted for although the number of
cases will not be signi cant, statistically speaking.
The ideas mentioned before motivates the development of
the process model for healthcare (PMH), which is described
in the next sections.</p>
    </sec>
    <sec id="sec-6">
      <title>3.2 PMH overview</title>
      <p>The knowledge in data mining projects frequently remains
in few people like consultants and experts in speci c
domains. For this reason, the experiences and processes
cannot be reusable and applied to similar projects in health
care. Furthermore, it is necessary to know di erent
guidelines, references and standards related to quality of service,
with the purpose of understand the main characteristics of
the health care domain.</p>
      <p>As a result, this paper presents PMH (Process Model for the
health care context). This process is an specialization of the
CRISP-DM methodology proposed in Colombian health care
context, based on the veri cation carried out in the
pharmacological and non-pharmacological treatments in
Colombia's health care institutions (IPS), through the application
of data mining techniques on RIPS les. This process allows
to reduce time and resources with respect to develop mining
projects in this domain without a speci c knowledge.
This guide proposes seven steps in an iterative way tacking
into account health care context. The following paragraphs
provide a description about the di erent steps involved in
this process model. The numbers used in this description
correspond to the number used in the gure 1.
I. Scope De nition of the exercise. This step allows to
de ne the business problem to be analysed. Several
fundamental aspects must be clari ed in this step: what, how, and
why the assessment is done, as well as de ning the criteria
for the success of the exercise. Some of the typical
questions proposed by experts are focused on problems to
validate the e ectiveness of pharmacological treatments in the
emergency room according to IPS's best practices, others in
the control and monitoring of chronic diseases according to
the standards. In this step, these questions will be
contextualized according to the service(hospitalizations, urgencies,
procedures or medical appointments), and the diagnosis to
be monitoring.</p>
      <p>II. Selection of the reference guide. This step II
consists on the selection of the reference guide(s) to evaluate the
question proposed in the rst step; for example the
medical treatment used for a diagnosis. Currently, it is possible
to use the expert advise to validate the quality of service
in health care. On the other hand, there is speci c
manuals such as standards, protocols, and clinical guidelines
proposed by governments and organizations that can be used
as a reference guide.</p>
      <p>
        The standards are evidence based on references used to
evaluate the quality and performance of services, while protocols
are documents that describe the rules of action depending
on a speci c circumstance [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Usually, protocols are
speci c documentation de ned by each IPS. Also, the clinical
guidelines are systematically developed statements to assist
practitioners and patient decisions about appropriate health
care for speci c clinical circumstances [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        About clinical guidelines, the World Health Organization
(WHO) presents di erent guides related to the diagnosis,
treatment and prevention of speci c diseases such as obesity,
malnutrition or diabetes. Furthermore, di erent countries
have established national standards to treat a disease. For
example, Colombia has the 412 resolution which suggests
the set of activities and procedures that should be used in
public health diseases [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>Although clinical guidelines vary in content, they have
essentially the following structure:</p>
      <p>Clinical guideline structure
0. Authors
1. Introduction
2. Disease detection
3. Diagnosis
4. Classi cation and Tracking
5. Disease evaluation
6. Non-pharmacological treatment
7. Pharmacological treatment
8. Disease complications
9. Disease special situations
10. Hospital treatment
11. Emergency treatment
12. Clinical guidelines future review recommendations
13. Bibliography
In the structure above, the interest lies (mainly) in points 4
to 11. The quality control proposed is based on the
comparison between treatments with a speci c admission diagnosis
and an established clinical guide diagnosis.</p>
      <p>The suggestion is to choose the reference guide that has been
established in the national policies or regulations, because
international clinical guidelines may not have validity in a
speci c country, or may not be applied in certain IPS
because of socio-economic or epidemiological factors.
III. Identi cation of information sources. This step
consists of the identi cation of useful information sources
according to the scope of the project and the selected
reference guides. This identi cation depends on the data quality
and availability. These issues will be tacking into account to
decide the use of these sources during the analysis process.
In the healthcare sector, there are di erent sources that can
be obtained and used for the development of data mining
projects.</p>
      <p>Generally, countries have an individual healthcare register
corresponding to every hospitalization, urgency or
procedure associated to a patient. For example, the United States
has the Electronic Health Record (EHR) which includes
demographics, medical history, medication and allergies,
immunization status and observations (among others). On
the other hand, Colombia has the Individuals Registers of
Health Care (RIPS in Spanish), that provides information
related to the delivery of health services and demographic
variables.</p>
      <p>There are other kind of information sources related to
national behavior such as national survey or naming standards.
Some national surveys contains demographic and health
information that can be used to support the data mining
process. Furthermore, the WHO de nes the CIE10 standard.
This standard de nes the classi cation and organization of
diseases based on a unique code that represents the category
and the speci c a ection.</p>
      <p>Some naming standards are associated to speci c health
interventions available for each country. For example, the
Australian Classi cation of Health Interventions (ACHI)
contains all the procedures that are realized by HMOs in the
country. The Unique Procedures Classi cation in healthcare
(CUPS in Spanish) is the Colombian classi cation system for
this information.</p>
      <p>IV. Selection and preparation of healthcare
information. It is necessary to have the support of health experts
to select the information that is highly relevant to solve the
proposed problem in step I. Each disease presents di erent
characteristics. For example, there are diseases like prostate
cancer or pre-eclampsia in which sex is not a determining
factor. The rst a ects men, and the latter applies only
for pregnant women. In chronic diseases like hypertension,
time is a signi cant variable. Its treatment is based on
monitoring the patient periodically with formulated procedures
and medicines according to a speci c order and to patient's
evolution over time. On contrary, appendicitis treatment is
considered relatively short. In this case the time variable
is not relevant. These considerations must be taken into
account in selection step.</p>
      <p>The selected information follow a data cleaning process.
Generally, health information has problems related to data
management such as replication of records. Moreover, medicines
data management proposed a new challenge to data mining
experts. Usually, this data does not have a standard to be
lled. For example, the medicine "amoxycillin" can be lled
as "amox" or "amoxycilin". In these cases, the similarity
word analysis can be used to solve the issue.</p>
      <p>The corresponding discretization of continuous data and the
standardization of information must be made, necessary
procedures for the execution of mining algorithms, which should
be in line with own business rules of the selected diagnosis.
For example, in the case of Alzheimer's disease, age
categories should be created after 40 years, being consistent
with the characteristics of vulnerable populations, and the
evolution of the disease over time. Di erent from
Appendicitis disease, where age ranges should be used much broader,
since it is a disease that can occur at any age. The
complexity of both the discretization and the standardization
of data may depend largely on the amount of selected
information sources and the absence of the use of ontologies
for the uni cation and standardization of medical terms and
concepts.</p>
      <p>An statistical analysis of this step is necessary for physicians
because it's important to know the data percentage that
must be cleaned and the problems that arise the datasets.
V. Information adjustment and preliminary
analysis. In this step is important to analyse the resulting dataset
of the previous phase. This analysis concerns to identify
the main characteristics of the data and the discovering of
new variables that are relevant for an speci c situation. In
healthcare, variables such as hospitalization window, total
cost of treatment and patient satisfaction are relevant in
many situations. For example, to solve questions like, what
is the most expensive treatment?, which is the one with the
lowest satisfaction?, which one represents a lower rate of
days of staying? these variables are relevant in this context.
The preliminary analysis is based on descriptive statistics.
The objective is to review the statistical data distribution in
order to avoid biased results. In this step can be identi ed
how many men or women are involved in the dataset, how is
the age distribution or if the treatment is based signi cantly
on drugs rather than procedures. The health experts
evaluate the results and if necessary, this step is review again.
VI. Selection and implementation of data mining
algorithms. This step consists in the identi cation of data
mining algorithms to achieve the objectives proposed. In
healthcare, there are probable classi cations of the mining
problems. The rst, is related to the analysis of treatments,
which is feasible to predict the organism response against
speci c procedures. Next, the monitoring and evolution
of patients in infectious and chronic disease. The latter,
is based on verify the proper provision of health services.
For example, for the last classi cation there situations in
which is appropriate to know the percentage of compliance
of treatments with respect to a clinical guidelines or which
are the implications of meeting/failing the clinical
guidelines in terms of costs. The data mining algorithms that
are proposed are association rules, sequential patterns and
clustering techniques to solve speci c problems in healthcare
sector.</p>
      <p>Association Rules: this technique identi es the cause-
effect relations between variables. It is possible to characterize
an speci c group using clustering techniques, and determine
the behaviour of an speci c cluster using association rules.
In healthcare, it's interesting to discover relations among
events. For example, if a patient's disease evolves to a
chronic phase, the probability to have a decease based on
speci c characteristics can be found. Moreover, In the
context of quality control, the association rules allow us to
discover rules that may or may not correspond to established
treatments in clinical guidelines It is recommendable to use
this technique in the diagnosis of acute illness. This kind
of illness usually are applied during a patient's admission to
the IPS and does not require periodic monitoring to ensure
new procedures or drugs depending on the patient progress.
Furthermore, association rules could be used in problems
that analyze the treatment's behaviour.</p>
      <p>The possible way to do in quality control in the health care
it is as follows:
Two data sets are created, one for drug treatment
information and one for non-drug information for each patient with
the same admission diagnosis, and with the same method of
admission. It is important that the information of the
admission method is not mixed, since the in each treatments
can be di erent. Thus, given an X diagnosis in
hospitalization, we have the following data set of medications:
For this example, an association rule would be: "In the
pharmacological treatment of a patient with an X diagnostic, if
medicines M30 and M90 are provided, then M70 medicine
will also be provided."
This technique required some parameters to be customized:
The support and the con dence. The rst one is the
proportion of transactions in the data set which contain the
itemset:
supp(A) = occurrence(A)</p>
      <p>size(dataset).</p>
      <p>The second one is de ned as the conditional probability (
P (BjA)): the occurrence of A, given the occurrence of B:
conf (A ) B) = supp(A S B)
supp(A).</p>
      <p>Sequential Patterns: This technique detects cause-e ect
relations considering the time periods in which transactions
occurred. In health care context, there are many situations
that involves periodic controls and patient's monitoring. As
mentioned before, standards, clinical guidelines and
protocols specify sequences of treatments that should be applied
in a explicit situation. For this reason, it is feasible to
compare the patient's procedures versus reference guides using
sequential patterns. Chronic diseases like hypertension or
diabetes requires that patients return to the IPS several
times with the same diagnosis. This kind of problems
involves many variables in each time period. The
recommendation is to use sequential patterns to analyze the evolution
of patient's health based on the treatments applied.
Based on the concept that, at any given time, a clinical event
is the formulation of one or more drugs or procedures we
propose the creation of two data sets, drugs and procedures
datasets. In this case the records should be grouped by
patient, ordering clinical events in ascending order by date
and time.
All clinical events from a patient arranged in order, can be
seen together as a sequence, where each event corresponds to
a set of drugs or procedures. Figure 4 shows the sequences
found in previous patients.
Clustering: to found groups of elements with similar
characteristics. In healthcare, is very common to analyze
populations based on speci c characteristics. Nevertheless, it's
possible to use this technique in the validation of right
treatments to right people. The reference guides de ne speci c
treatments for people with Speci c demographic
characteristics. In this case, the use of clustering determines subsets
based on procedures and demographic information.
In addition, there are situations in which data quality is
poor. For this reason, its impossible to analyze particular
issues in the dataset because of the con dence of data. The
suggestion is to use clustering techniques to generalize the
main characteristics of an speci c group. To perform these
analysis, a data set must be created which includes patient
information such as gender, age, marital status and race,
(among others), as well as drugs and/or procedures
provided, grouped into treatments found in the previous
section, as shown in gure 5:
As in previous algorithms, it is important to analyze the
number of people supporting each cluster, before making
any conclusions. It is also essential to understand, that a
cluster represents a very small percentage of the population
does not necessarily implies that should be discarded. It all
depends on the clinical context that is being evaluated and
the criteria of the medical expert.</p>
      <p>VII. Result Validation and Impact Analysis. This
nal steps concerns to tunning up the data mining model
based on the health expert feedback and the results achieved.
A detailed study of the results is made by a board of medical
experts in the area supported by a technical group of people
to determine whether or not to repeat some of the above
steps, possibly with a change in strategy or range, or by a
re nement of the data used, to get speci c conclusions.</p>
    </sec>
    <sec id="sec-7">
      <title>4. VALIDATION</title>
      <p>This proposal is validated based on the PMH process and
the product obtained by applying this steps in the solution
of a real problem in Colombia's health care sector.
Our proposal is a specialization of the CRISP-DM
methodology. The CRISP-DM process has been validated in several
domains and modi ed during more than a decade, based on
the application of the process in many data mining projects.
In that sense, we re ned the CRISP phases in order to
improve the knowledge and special issues of the health care
sector, reducing the time and resources that have to be used
for understand this particular domain.</p>
      <p>The PMH process incorporates the diagnosis and reference
guide selection in the business understanding phase of
CRISPDM, and presents the selection criteria for these steps.
Furthermore, in the other phases is possible to understand the
principal problems associated with data quality in health
care, and to determine which algorithms should be used to
solve several problems in this sector.</p>
      <p>The next sections are focused on the product validation. It
applies the PMH process in the quality control of Colombia's
health care sector and validate the results based on the
experts criteria. The development of this exercise consists in
two iterations of the process. The speci cs steps of PMH
are described below.</p>
    </sec>
    <sec id="sec-8">
      <title>4.1 Problem description</title>
      <p>Hypertension is a chronic disease that a ects 20% of people
in the world. This is considered the rst cause of morbility
and the most representative disease related to cardiovascular
a ections. For this reason, the objective of this exercise was
to evaluate the pharmacological and non-pharmacological
treatment for this disease in Colombia.
plemented using the IBM Intelligent Miner 8.1 data mining
tool.</p>
      <p>The gure 6 shows the results of the mining model.</p>
    </sec>
    <sec id="sec-9">
      <title>4.2 First iteration</title>
      <p>
        Step I consists on the description of the problem according to
the context, related to the diagnosis and type of disease that
is going to be tackled. In this case, it is relevant to analyze
the characteristics of hypertension. This is a chronic disease,
generally asymptomatic and it requires continuous medical
assistance. The typical complications of hypertension are
related to cardiac failures. This complications implies
hospitalizations, urgencies and complex procedures.
Based on the health experts support, there are di erent
reference guides related to hypertension, but its treatment may
di er from one country to another. For this reason, the
selection of the reference guide was focused on the "Clinical
guideline for hypertension disease" proposed by the
Association of Faculties of Medicine of Colombia [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>As mentioned in section III, Colombia has the Individual
Registers of Health Care. The objective of this data is
related to bill the delivery of health services made by the
Colombian's IPS. The information was collected from HMOs
and the Minister of Social Protection from years 2003 to
2006 (49,000,000 of individual register of health care).
Furthermore,the RIPS data are lled based on the CIE10 and
CUPS standards. For this reason, these information is also
taken into account.</p>
      <p>According to the expert opinion, the relevant information
for this analysis is described in table 2:</p>
      <p>Variables
Principal diagnosis
IPS Identi cation
Date
Sex
Department
Type of medical service
Medicine
Procedures
Length of stay
decease diagnosis
related diagnosis
The statistical analysis results for the dataset shows that
a 77% of the registers correspond to procedures, 17% are
medical appointments and a 6% of the data to medicines
prescriptions. Moreover, a 67% correspond to men and a
33% to women. On the other hand, a 16% of the patients
return to the IPS for health controls.</p>
      <p>In the pharmacological treatment of hypertension, we want
to analyze the most relevant medication sequences. In this
case, the time variable is highly relevant because we want
to trace the prescription of medicines for a speci c disease.
For this reason, we used sequential patterns, according to
the ideas presented in subsection 3.2. Our model was
imThe main result is the use of captopril. More than 40% of
the patients were medicated with this medicine. Based on
health experts opinions, the captopril is a medicine used in
mono-therapy treatments and is prescripted for economical
reasons. Other important conclusion is related to data
quality of medicines; a speci c naming standarisation was used
to resolve the problem. In the same way, several manual
process was realised to mitigate replication of records
problem.</p>
    </sec>
    <sec id="sec-10">
      <title>4.3 Second iteration</title>
      <p>Based on the results of the rst iteration, a second iteration
of the steps suggested in PMH was performed. Therefore,
the health experts proposed to analyze if the health system
may incur in higher costs because of the prescription of
captopril to the patients. Using the RIPS data, we want to
determine the complications related to these patients.
According to the PMH, it is necessary to prepare the data
that will be used by the model. For this reason, the data
presents demographic information and relevant aspects about
the evolution of a patient's disease.</p>
      <p>In this case, new variables have to be include based on the
health expert recommendation. This variables are
associated to the evolution to a chronic phase in the clinical
history of patients. For this reason, were introduced the date
of the patient's complication, the associated diagnosis, the
number of hospitalizations or procedures before and after the
complication of hypertension. The next step of the
methodology, proposes a preliminary analysis of the information. In
this case, the information consists on all the records of the
patients that have su ered this disease.</p>
      <p>To describe the complication of a patient, the sequential
clustering technique was used. As described in section 3,with
this technique it is possible to nd clusters of patients with
similar sequences and characteristics.</p>
      <p>The results of the second iteration shows that patients with
similar sequences associated to the use of captopril have
complications such as chronic cardiac failures, hypertensive
crisis or heart attacks.</p>
      <p>In general, it is important to analyze that the use of
captopril is pre-scripted for economical reasons. This strategy
is useful in a short term period, but in a long term, we can
observe that patients with this kind of treatment, returns
to the healthcare institution with complications in the
hypertension disease. This kind of complications increase the
illness costs.</p>
    </sec>
    <sec id="sec-11">
      <title>5. CONCLUSIONS AND FUTURE WORK</title>
      <p>This paper proposes a process model to guide the data
mining process in the health care sector. It suggests a set of
iterative and facultative steps to improve the results of the
mining process. This process model was evaluated using the
analysis of quality of service for the treatment of
hypertension in Colombia. The results shows that it is possible to
establish new hypothesis about the datasets, and revalidate
this a rmations using the proposed process model. At the
same way, these results evidence some facilities provided to
the data mining expert to guide their process, specially
associated to the knowledge about healthcare context such as
data sources, reference guides and data mining techniques.
An exhaustive validation of the process model is considered
as future work, in terms of a formal comparison between
the use of CRISP-DM and PMH. At the same time, new
kind of question from the expert point of view will be
interesting to resolve using this process model. In particular,
the identi cation of an epidemiological pro le for Colombian
population.</p>
    </sec>
    <sec id="sec-12">
      <title>6. ACKNOWLEDGMENT</title>
      <p>The authors appreciate the support of Jose Abasolo,
professor at Los Andes University, which provides the initial ideas
and suggestions for the development of this article.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>E. A.M.</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Fu</surname>
          </string-name>
          .
          <article-title>Privacy preserving distributed learning clustering of healthcare data using cryptography protocols</article-title>
          .
          <source>In Computer Software and Applications Conference Workshops (COMPSACW)</source>
          ,
          <source>2010 IEEE 34th Annual</source>
          , pages
          <volume>140</volume>
          {
          <fpage>145</fpage>
          , july
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Chapman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clinton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kerber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Khabaza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Reinartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Shearer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Wirth</surname>
          </string-name>
          .
          <article-title>Crisp-dm 1.0 step-by-step data mining guide</article-title>
          .
          <source>Technical report</source>
          ,
          <article-title>The CRISP-DM consortium</article-title>
          ,
          <year>August 2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] A. colombiana de Facultades de Medicina.
          <article-title>Gu a cl nica para la hipertension arterial</article-title>
          . http://www.redsalud.gov.cl/archivos/guiasges/ hipertension arterial primaria.
          <source>pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>eHow</given-names>
            <surname>Health</surname>
          </string-name>
          .
          <article-title>De nition of clinical protocol</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>X.</given-names>
            <surname>Fang</surname>
          </string-name>
          .
          <article-title>Are you becoming a diabetic? a data mining approach</article-title>
          .
          <source>In Fuzzy Systems and Knowledge Discovery</source>
          ,
          <year>2009</year>
          . FSKD '
          <volume>09</volume>
          . Sixth International Conference on, volume
          <volume>5</volume>
          , pages
          <fpage>18</fpage>
          {
          <fpage>22</fpage>
          , august
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Harleen</surname>
          </string-name>
          and
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Wasan</surname>
          </string-name>
          .
          <article-title>Empirical study on application of data mining techniques in health care</article-title>
          .
          <source>Journal of computer science 2</source>
          , pages
          <fpage>194</fpage>
          {
          <fpage>200</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ji</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <article-title>A novel abnormal ecg beats detection method</article-title>
          .
          <source>In Computer and Automation Engineering (ICCAE)</source>
          ,
          <year>2010</year>
          The 2nd International Conference on, volume
          <volume>1</volume>
          , pages
          <fpage>47</fpage>
          {51, february
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>B. M.J.A.</surname>
          </string-name>
          and
          <string-name>
            <surname>L. G.S.</surname>
          </string-name>
          <article-title>Mastering data mining</article-title>
          .
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>L. N.</surname>
          </string-name>
          <article-title>Predicting the risk of future hospitalization</article-title>
          .
          <source>In Database and Expert Systems Applications (DEXA)</source>
          ,
          <source>2010 Workshop on</source>
          , pages
          <volume>120</volume>
          {124, september
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S. N.</given-names>
            and
            <surname>T. A. M.</surname>
          </string-name>
          <article-title>The relevance of data warehousing and data mining in the eld of evidence-based medicine to support healthcare decision making</article-title>
          .
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Nevine</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Malek</surname>
          </string-name>
          .
          <article-title>Data mining for cancer management in egypt case study: Childhood acute lymphoblastic leukemia</article-title>
          .
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>N. I.</surname>
          </string-name>
          <article-title>of Health. About clinical practice guidelines</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>[13] M. of Social Protection. Resolucion</source>
          <volume>412</volume>
          . http://mps.minproteccionsocial.gov.co/pars/cajaherram/documentos/Biblioteca/CompendioNormativo/ resolucion 412 00.pdf,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A. S. S.</given-names>
            <surname>Raza</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Raza</surname>
          </string-name>
          .
          <article-title>A case for supplementing evidence base medicine with inductive clinical knowledge: Towards a technology-enriched integrated clinical evidence system</article-title>
          .
          <source>In Proceedings of the Fourteenth IEEE Symposium on Computer-Based Medical Systems, CBMS '01</source>
          , pages
          <fpage>5</fpage>
          {, Washington, DC, USA,
          <year>2001</year>
          . IEEE Computer Society.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>N. R. T.</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Jian</surname>
          </string-name>
          .
          <article-title>Introduction to the special issue on data mining for health informatics</article-title>
          .
          <source>SIGKDD Explor</source>
          . Newsl.,
          <volume>9</volume>
          :
          <issue>1</issue>
          {
          <fpage>2</fpage>
          ,
          <string-name>
            <surname>June</surname>
          </string-name>
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>van Driel</surname>
            <given-names>M. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>C. K.</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. P. P.</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. J. A.</surname>
          </string-name>
          , and
          <string-name>
            <surname>B. H. G.</surname>
          </string-name>
          <article-title>A new web-based data mining tool for the identi cation of candidate genes for human genetic disorders</article-title>
          .
          <source>Eur J Hum Genet</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <volume>57</volume>
          {
          <fpage>63</fpage>
          +,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. L.</given-names>
            <surname>Henning</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. W. W.</surname>
          </string-name>
          , and
          <string-name>
            <given-names>O.</given-names>
            <surname>Jorn</surname>
          </string-name>
          .
          <article-title>Drug exposure side e ects from mining pregnancy data</article-title>
          .
          <source>SIGKDD Explor</source>
          . Newsl.,
          <volume>9</volume>
          :
          <fpage>22</fpage>
          {
          <fpage>29</fpage>
          ,
          <year>June 2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>