<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Improving clinical record visualization recommendations with Bayesian stream learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Position Paper</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pedro Pereira Rodrigues</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Claudia Dias</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ricardo Cruz-Correia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CINTESIS - University of Porto</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Medicine of the University of Porto</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>LIAAD - INESC Porto</institution>
          ,
          <addr-line>L.A.</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Clinical record integration and visualization is one of the most important abilities of modern health information systems (HIS). Its use on clinical encounters plays a relevant role in the e cacy and efciency of healthcare. However, integrated HIS of central hospitals may gather millions of clinical reports (e.g. radiology, lab results, etc.). Hence, the clinical record must manage a stream of reports being produced in the entire hospital. Moreover, not all documents from a patient are relevant for a given encounter, and therefore not visualized during that encounter. Thus, the HIS must also manage a stream of events of visualization of reports, which runs in parallel to the stream of documents production. The aim of our project is to provide the physician with a recommendation of clinical reports to consider when they log in the computer. Our approach is to model relevance as the probability that a given document will be accessed in the current time frame. For that, we design a data stream management system to process the two streams, and Bayesian networks to learn those probabilities based on document, patient, department and user information. One of the biggest challenges to the learning problem, so far, is that no negative examples are produced by the stream (i.e. there are no record of documents not being visualized) leading to a one-class classi cation problem. The aim of this paper is to clearly present the setting and rationale for the approach. Current work is focused on both the stream processing mechanism and the Bayesian probability estimation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The identi cation of clinically relevant information should enable an
improvement in user interface design and in data management. However, it is di cult to
identify what information is important in daily clinical care, and what is used
only occasionally. The main problem addressed by this project is how to
estimate the relevance of healthcare information in order to anticipate its usefulness
at a speci c point of care. In particular, we want to estimate the probability
of a piece of information being accessed during a certain time interval, taking
into account the type of data, the context where it was generated and is needed
and the type of users who access it, and to use this probability to prioritize the
information.
1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Relevance of clinical documents</title>
      <p>
        In the healthcare domain, and especially in critical and acute care, the age of data
is one of the factors often used to assess data relevance, making new information
more relevant to the current search. Some authors have categorized old data as
data at least three days old. However, in a previous study, the authors have
examined for how long are clinical documents used by health professionals in
a hospital environment, and how this is associated with document content and
the context of information request [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Those results show that some clinical
reports are still used after one year regardless of the context in which they
were created, although signi cant di erences exist in reports created in distinct
encounter types. The authors conclude that the usage of past patient data (data
from previous hospital encounters) varied signi cantly according to the setting
of healthcare and document content, which contradicts the de nition of old data
used in previous studies. Hence the need to de ne better rules for recommending
documents in encounters.
1.2
      </p>
    </sec>
    <sec id="sec-3">
      <title>Setting</title>
      <p>
        The current setting is a central hospital that has several departmental
information systems that produce clinical documents (e.g. radiology reports and lab
results) that might be relevant for the practice of healthcare. The access to this
documents is better achieved by a centralized information system that integrates
all the di erent departmental systems, aggregating the documents that are most
relevant for the current encounter [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These sources can be modeled as data
streams. A data stream is an ordered sequence of instances that can be read
only once or a small number of times [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], using limited computing and storage
capabilities. Hence, there are two data streams being produced in parallel:
{ a stream of documents being produced; each element is a new document (or
a new version of an existing one); and
{ a stream of visualization events; each element is an event of visualization of
a previously produced document.
      </p>
      <p>The rst stream is gathered by integrating documents from heterogeneous
information systems, so we might end up with syncronization issues. In the
current setting, the central hospital information system receives an increasing
rate of 200+ documents per hour (as seen in Figure 1), relative to the daily 5300+
patients in the hospital (including inpatients, outpatients and emergency rooms)
which need to be processed. This created a pool of 8M+ documents (produced
since 2004 from 400K+ patients). However, due to constant document revisions,
\only" 2.9M+ are actually available for visualization (active documents), whilst
5M+ are previous versions.
0
5
2
0
0
2
0
5
/rsouH 150
t
n
e
couDm 100</p>
      <p>● VPirsoudaulcizeadtioDnoscuinm2e0n1ts0/2011
0 ●
2004
●
2005
●
2006
●
2007
Year
●
2008
●
2009
●
2010</p>
      <p>The second stream is controlled in the centralized integrating information
system, tracking information on visualizations performed by 4850+ users. Older
documents tend to be less visualized in encounters but, as seen in Figure 1,
nearly half of the visualizations in 2010 and 2011 targeted documents older than
January 2009, so old documents cannot be discarded. During an encounter with
a patient, the pool of documents that one of the currently active 2375 users
might access is 537.9 (average number of active documents per active patient).
Clearly, the user cannot be presented with a list of 530+ documents, so a ranking
is needed to prioritize the most relevant.
1.3</p>
    </sec>
    <sec id="sec-4">
      <title>Learning problem</title>
      <p>
        Overall, this is one clear setting of data streams in medical scenarios, on which
machine learning techniques can be applied to improve healthcare. In order to
select relevant documents for visualization in a precise encounter, we need to
take care of both the age and the information related to those documents.
Either way, this setting is an uncertain one. Thus, the best way to model this
relevance is to estimate the probability that the document is going to be
visualized in the current time period. In machine learning, uncertainty is usally well
modeled using Bayesian approaches [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Therefore, we seek to develop Bayesian
networks to estimate the probability of a certain clinical document being
visualized in the near future, and use this probability to rank the list of possible
documents related to the current encounter. The fact that no negative examples
are present (only visualization events) creates a harder setting for learning, so we
are aiming at Bayesian stream learning for one-class classi cation [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The
oneclass classi cation problem is di erent from the conventional binary/multi-class
classi cation problem in the sense that the negative class is either not present
or not properly sampled. This is, at least in medical settings and at the best of
our knowlege, uncharted research territory.
1.4
      </p>
    </sec>
    <sec id="sec-5">
      <title>Aim and outline</title>
      <p>The aim of this work is to extend the health information system responsible for
visualization of clinical documents at point of care, providing the physician with
a recommendation of clinical reports to consider when in the presence of a
patient. Basically, we propose to study how to estimate the relevance of healthcare
information in a particular setting, aiming to use this to create adaptive user
interfaces. Speci cally, our objectives are:
{ to collect and prepare log data from hospital information systems usage, to
feed our two data streams;
{ to study the factors associated with the relevance of clinical documents;
{ to de ne an algorithm to estimate the relevance of a particular document
for a precise encounter;
{ to implement an adaptive user interface based on a ranked list of
recommended documents per encounter.</p>
      <p>The paper is organized as follows. Next section presents background
knowledge on electronic health records, learning from data streams and Bayesian
networks in healthcare. Then, section 3 presents our approach for i) the data stream
processing, ii) the estimation strategy, iii) the incremental learning and iv) the
recommendation generator. Section 4 ends the exposition with future work.
2</p>
      <p>
        Background
This work is related with three di erent areas of research: medical
informatics, especially devoted to electronic health records; Bayesian learning from data
streams; and one-class classi cation (check [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] for a survey on this problem).
2.1
      </p>
    </sec>
    <sec id="sec-6">
      <title>Electronic Health Records</title>
      <p>
        The practice of medicine has been described as being dominated by how well
information is collected, processed, retrieved, and communicated [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Patient
records, the patient and published evidence are the three information sources
needed to practice evidence-based medicine [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>
        Currently in most hospitals there are great quantities of stored digital data
regarding patients, in administrative, clinical, lab or imaging systems. An
important challenge is to guarantee the optimal conditions for health professionals to
access clinical data while hospital information systems are still being developed.
Although great advances have been made over the years, on-demand access to
clinical information is still inadequate in many settings, contributing to
duplication of e ort, excess costs, adverse events, and reduced e ciency. Although it is
widely accepted that full access to integrated electronic health records (EHRs)
and instant access to up-to-date medical knowledge signi cantly reduces faulty
decision making resulting from lack of information [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], there is still very little
evidence that life-long EHRs improve patient care [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Shapiro et al. found that, although emergency department doctors believe
their patients would bene t from longitudinal records, they only try to obtain
such data in 10% of the cases [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Furthermore Hripcsak et al. described access
rates to WebCIS in the emergency department [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], which indicated that data
generated before the current emergency visit are accessed often, but by no means
in a majority of times (5% to 20% of the encounters), even when the user was
noti ed of the availability of such data.
      </p>
      <p>
        Cruz-Correia et al. have done several pilot studies to analyse for how long
clinical documents are useful for health professionals in a hospital environment,
bearing in mind document content and the context of the information request.
The results show that some clinical reports are still used one year after creation,
regardless of the context in which they were created, although signi cant di
erences existed in reports created during distinct encounter types. The median-life
of reports by the type of encounter during which they were created is 1.7 days
for emergency, 3.9 days for inpatient and 27.7 days for outpatient encounters.
They conclude that the usage of patients past information (data from previous
hospital encounters), varied signi cantly according to the setting of healthcare
and content [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Also, the amount of digital data produced in the medical imaging
departments has increased rapidly in recent years, due mainly to two factors: 1) greater
use of additional diagnostic procedures, resulting in a greater number of tests
produced and, 2) an increase in the quality of the examinations, which is
translated into a greater number of images acquired. Despite great development in
digital storage technologies, the cost and e ort required to maintain this
information online throughout its life cycle can be considerable.</p>
      <p>The management of information in these systems is usually implemented
using Hierarchical Storage Management (HSM) solutions. This type of solution
enables the implementation of various layers which use di erent technologies
with di erent speeds of access, corresponding to di erent associated costs.
However, the solutions which are currently implemented in Picture Archiving and
Communication System (PACS) use simple rules for information management,
based on variables such as the time elapsed since the last access or the date of
creation of information, not taking into account the likely relevance of
information in the clinical environment.</p>
      <p>Classifying the relevance of information based only on the time elapsed since
the date of acquisition is clearly ine cient. It is expected that the need to consult
an examination at a given time will be dependent on several factors beyond the
date of the examination, such as type of examination and the patient's pathology.
Thus, a system that uses more factors to identify the relevance of information at
a given time would be more e cient in managing the information that is stored
in fast memory and slow memory.
2.2</p>
    </sec>
    <sec id="sec-7">
      <title>Machine learning from data streams</title>
      <p>
        What distinguishes current data from earlier one are automatic data feeds. We do
not just have people who are entering information into a computer. Instead, we
have computers entering data into each other [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Thus, there are applications
in which the data is modeled best not as persistent tables but rather as transient
data streams.
      </p>
      <p>
        A data stream is an ordered sequence of instances that can be read only once
or a small number of times using limited computing and storage capabilities.
The data elements in the stream arrive online, being potentially unbounded in
size. Once an element from a data stream has been processed it is discarded or
archived. It cannot be retrieved easily unless it is explicitly stored in memory,
which is small relative to the size of the data streams. These sources of data are
characterized by being open-ended, owing at high-speed, and generated by non
stationary distributions [
        <xref ref-type="bibr" rid="ref8 ref9">8,9</xref>
        ].
      </p>
      <p>
        In online streaming scenarios, predictions are usually followed by the real
label value in a short future (e.g., prediction of next value of a time series).
Nevertheless, there are also scenarios where the real label value is only available
after a long term, such as predicting one week ahead electrical power
consumption [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Learning techniques which operate through xed training sets and
generate static models are obsolete in these contexts. Faster answers are usually
required, keeping an anytime data model and enabling better decisions, possibly
forgetting older information.
      </p>
      <p>
        The sequences of data points are not independent, and are not generated
by stationary distributions. We need dynamic models that evolve over time and
are able to adapt to changes in the distribution generating examples [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. If the
process is not strictly stationary (as most of real-world applications), the target
concept may gradually change over time. Hence data stream mining is an
incremental task that requires incremental learning algorithms that take drift into
account [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Hulten et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] presented desirable properties for data stream learning
systems. Overall, they should process examples at the rate they arrive, use a single
scan of data and xed memory, maintain a decision model at any time and be
able to adapt the model to the most recent data. Successful data stream learning
systems were already proposed for both prediction [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and clustering [
        <xref ref-type="bibr" rid="ref19 ref3">3,19</xref>
        ]. All
of them share the aim to produce reliable predictions or clustering structures.
2.3
      </p>
    </sec>
    <sec id="sec-8">
      <title>Machine learning in healthcare</title>
      <p>
        The application of data mining and machine learning techniques to medical
knowledge discovery tasks is now a growing research area. These techniques vary
widely and are based on data-driven conceptualizations, model-based de nitions
or on a combination of data-based knowledge with human-expert knowledge [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        The de nition of clinical decision support systems is now a major topic since
it may help the diagnosis, the prognosis of rate of mortality, the prognosis of
quality of life, or even treatment selection. However, the complicated nature of
real-world biomedical data has made it necessary to look beyond traditional
biostatistics [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] without loosing the necessary formality. For example, naive
Bayesian approaches are closely related to logistic regression [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Hence, those
systems could be implemented applying methods of machine learning [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], since
new computational techniques are better at detecting patterns hidden in
biomedical data, and can better represent and manipulate uncertainties [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        Traditional statistical methods require that the model structure is given and
only probabilistic information is learned from biomedical evidence, in the form
of data, whereas machine-learning approaches enable that both the structure of
the models and the probabilistic information are evidence-based[
        <xref ref-type="bibr" rid="ref14 ref16">14,16</xref>
        ]. Bayesian
approaches have an extreme importance in these problems as they provide a
quantitative perspective [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Moreover, Bayesian statistical methods allow
taking into account prior knowledge when analyzing data, turning the data analysis
into a process of updating that prior knowledge with biomedical and health-care
evidence [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. However, only after the 90's we may nd evidence of a large
interest on these methods, namely on Bayesian networks, which o er a general and
versatile approach to capturing and reasoning with uncertainty in medicine and
health care [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Given their improved management of uncertainties, Bayesian networks have
been successfully applied in healthcare domains [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Bayesian networks can be
seen as an alternative to logistic regression where statistical dependence and
independence are not hidden in approximating weights, rather explicitly
represented by links in a network of variables [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. They describe the distribution
of probabilities of one set of variables, making possible a two-fold analysis: a
qualitative model and a quantitative model, presenting two types of information
for each variable.
      </p>
      <p>
        On a general basis, a Bayesian network represents a joint distribution of one
set of variables, specifying the assumption of independence between them, with
the inter-dependence between variables being represented by a directed acyclic
graph. Each variable is represented by a node in the graph, and is dependent of
the set of variables represented by its ascendant nodes; a node X is a ascendant
of another node Y if exists a direct arc from X to Y [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. To give more
representational power to the relations represented by the arcs of the graph, it is necessary
to associate values to it. The matrix of conditional probability is given for each
variable, describing the distribution of probabilities of each variable given its
ascendant variables.
      </p>
      <p>
        After the qualitative and quantitative models are constructed, the next step,
and one of the most important, is how to calculate the new probabilities when
new evidence is introduced in the network. This process is called inference and
works as follows. Each variable has a nite number of categories greater than
or equal to two. A node is observed when there is knowledge about the state
of that variable. The observed variables have a huge importance because with
conditional probabilities they de ne the prior probabilities of the non observed
variables. With the joint probabilities we can calculate the marginal probabilities
of each unobserved variable, adding for all categories the probabilities that the
variable is in the desired state [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
3
      </p>
      <p>Optimizing the visualization of the clinical record
This section exposes our approach to the problem, presenting the data stream
management system, the relevance estimation strategy, the Bayesian learning
model, and the documents recommender system.
3.1</p>
    </sec>
    <sec id="sec-9">
      <title>Data stream management system</title>
      <p>Currently, a lot of patient information is accessible to healthcare professionals
at the point of care. In some cases, the amount of information is becoming
too large to be readily handled by humans or to be e ciently managed by
traditional storage algorithms. Most Hospital information systems record the
actions performed by users { the log le. These logs are kept for audit purposes
but can give insights into the information needs of healthcare professionals in a
particular situation. The study of these logs should allow us not only to describe
how the systems were used, but may also be useful to predict future use of the
system and of the data items it contains. To this latter task, there are two data
streams being produced in parallel: the stream of documents being created, and
the stream of visualization events.</p>
      <p>
        The rst stream is gathered by integrating documents from heterogeneous
information systems, which should be modeled according to the insert-delete
or turnstile model [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], allowing that observations might be updated or deleted
by future events (some clicnical documents are subject of validation and
revision, deactivating the previous versions of that document). The second one is
controlled in the centralized integrating information system, but it only tracks
information on visualizations, so it should be modeled according to the
insertonly or time series model [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], since the event of visualization is not subject of
deletion. However, given the predictive task of our system, we could consider the
accumulative or cash-register model [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], where each observation is an increment
to a given sum, i.e. the counters of the number of visualizations of each document
in the current time period.
3.2
      </p>
    </sec>
    <sec id="sec-10">
      <title>Estimating relevance of clinical documents</title>
      <p>
        By applying regression methods or other modeling techniques it is possible to
identify which factors are associated with the usage or relevance of patient data
items. These factors and associations can then be used to estimate data relevance
in a speci c future time interval. We expect to nd considerable di erences in
the relevance or median-life of information depending on several factors, such
as the origin of the data (e.g. lab department, emergency department), the age
of patient, and the current patient diagnosis. We also expect there to be an
exponential decay in the use and thus relevance of each patient's data over
time [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        As previously discussed, we rely on the Bayesian networks ability to model
uncertainty to estimate the probability of a given clinical document to be
visualized. After de ning the set of patient variables P and the set of other factors F
that are associated with that relevance, we shall create a Bayesian network with
jjP jj + jjF jj + 2 nodes, where the remaining nodes are the age of the document
and the class (visualized or not). Given the discrete characteristics of variables in
the Bayesian network, age of document needs to be categorized into contiguous
intervals. According to previous work [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], this categorization should be de ned
with exponential intervals, to model the relevance decay with time. Moreover,
they can be de ned as percentiles of visualization events in previous data.
3.3
      </p>
    </sec>
    <sec id="sec-11">
      <title>Incremental learning the Bayesian network</title>
      <p>When a document is visualized (i.e. a new observation is created in the
visualization stream) a learning example can be created, computing the age of
the document with real creation and visualization dates. This learning
example can thus be fed to the Bayesian network. However, there are no records of
non-visualization of documents, so negative learning examples are never created.
Sophisticated approaches shall be tested to solve this issue:
1. for each positive example representing a visualization event, create as many
negative examples for the time periods elapsed since the last visualization
(or creation); or
2. de ne a purely one-class classi cation learner.</p>
      <p>From our setting, we can easily extract information on previous visualization
of a given clinical document (the log le includes this data) so we will follow the
rst approach.
3.4</p>
    </sec>
    <sec id="sec-12">
      <title>Generating the recommendations</title>
      <p>The goal is to achieve the following setup. A doctor-patient encounter (e.g.
emergency, outpatient consultation, inpatient consultation) requires visualization of
clinical documents. These could be documents from the same patient (patient's
medical history) or documents from patients with similar characteristics or
diagnosis. Given the huge amount of available clinical documents, the information
system should list only the most relevant for that encounter.</p>
      <p>Several approaches to rank the recommendations could be followed, and
certainly they will be studied. At rst, we shall consider the probability of a clinical
document being visualized as the single rank variable. However, if we need to
include documents from di erent patients, we might end-up biasing the results by
considering erroneous visualizations. Naively, the probabilities of visualization of
all possible documents need to be computed everytime the system needs to list
the recommendations, which will turn the process infeasible given the streaming
setup. There are, at least, three possible ways to solve this issue:
1. at every visualization event, probability of visualization of only that
document is updated;
2. at every visualization event, probability of visualization of all documents is
updated;
3. at every ranked list request, relevance of that patient's or similar patient's
documents is updated;</p>
      <p>This is the least solved part of the system, and future work is expected.
Nevertheless, we feel that it is important to keep this target in mind because it
may condition the possible paths that previous modules are going to traverse.
4</p>
      <p>Future steps and expected impact of the system
Future work is concentrated on: a) de ning a data stream management system,
b) de ning the factors that in uence the relevance of clinical documents, c)
build learning models to estimate the relevance of a single document for a given
encounter, d) generate recommendations based on a ranking, and e) develop
and test the prototype with real data. To our knowledge, the use of machine
learning techniques to support graphical user interfaces and management storage
systems in healthcare information systems is novel and could be an important
contribution to science and likely to be incorporated into commercial products in
the future. The results of this research are also likely to have an important impact
on the quality of healthcare by further increasing the usability and intelligence
of existing information systems.</p>
    </sec>
    <sec id="sec-13">
      <title>Acknowledgments</title>
      <p>This work is supported by Portuguese Foundation for Science and Technology
(FCT) under projects OPTIM (PTDC/EIA-EIA/099920/2008) and KDUDS
(PTDC/EIA-EIA/98355/2008).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Barnett</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Computers in medicine</article-title>
          .
          <source>JAMA: the journal of the American Medical Association</source>
          <volume>263</volume>
          (
          <issue>19</issue>
          ),
          <volume>2631</volume>
          (
          <year>1990</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Clamp</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keen</surname>
          </string-name>
          , J.:
          <article-title>Electronic health records: Is the evidence base any use? Medical Informatics and the Internet in Medicine 32(1</article-title>
          ),
          <volume>5</volume>
          {
          <fpage>10</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cormode</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muthukrishnan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhuang</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          :
          <article-title>Conquering the divide: Continuous clustering of distributed data streams</article-title>
          .
          <source>In: Proceedings of the 23rd International Conference on Data Engineering (ICDE</source>
          <year>2007</year>
          ). pp.
          <volume>1036</volume>
          {
          <issue>1045</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cruz-Correia</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vieira-Marques</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferreira</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almeida</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wyatt</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Costa-Pereira</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Reviewing the integration of patient data: how systems are evolving in practice to meet patient needs</article-title>
          .
          <source>BMC Medical Informatics and Decision Making</source>
          <volume>7</volume>
          (
          <issue>1</issue>
          ),
          <volume>14</volume>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cruz-Correia</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wyatt</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dinis-Ribeiro</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Costa-Pereira</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Determinants of frequency and longevity of hospital encounters' data use</article-title>
          .
          <source>BMC Medical Informatics and Decision Making</source>
          <volume>10</volume>
          ,
          <issue>15</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dick</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>The Computer-based Patient Record: An Essential Technology for HealthCare</article-title>
          . National Academy Press (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gama</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Medas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>P.P.</given-names>
          </string-name>
          :
          <article-title>Learning with drift detection</article-title>
          . In: Bazzan,
          <string-name>
            <given-names>A.L.C.</given-names>
            ,
            <surname>Labidi</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the 17th Brazilian Symposium on Arti cial Intelligence (SBIA</source>
          <year>2004</year>
          ).
          <source>Lecture Notes in Arti cial Intelligence</source>
          , vol.
          <volume>3171</volume>
          , pp.
          <volume>286</volume>
          {
          <fpage>295</fpage>
          . Springer Verlag,
          <source>Sa~o Luiz</source>
          , Maranha~o,
          <string-name>
            <surname>Brazil</surname>
          </string-name>
          (
          <year>October 2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gama</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>P.P.</given-names>
          </string-name>
          :
          <article-title>Data stream processing</article-title>
          . In: Gama,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Gaber</surname>
          </string-name>
          , M.M. (eds.)
          <source>Learning from Data Streams - Processing Techniques in Sensor Networks, chap. 3</source>
          , pp.
          <volume>25</volume>
          {
          <fpage>39</fpage>
          . Springer Verlag (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Gama</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Sebastia~o,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Rodrigues</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.P.</surname>
          </string-name>
          :
          <article-title>Issues in evaluation of stream learning algorithms</article-title>
          .
          <source>In: Proceedings of the 15th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (KDD</source>
          <year>2009</year>
          ). pp.
          <volume>329</volume>
          {
          <fpage>337</fpage>
          . ACM Press, Paris, France (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Guha</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meyerson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mishra</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Motwani</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Callaghan</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Clustering data streams: Theory and practice</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>15</volume>
          (
          <issue>3</issue>
          ),
          <volume>515</volume>
          {
          <fpage>528</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Hripcsak</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sengupta</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilcox</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Green</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Emergency department access to a longitudinal medical record</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          <volume>14</volume>
          (
          <issue>2</issue>
          ),
          <volume>235</volume>
          {
          <fpage>238</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Hulten</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spencer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Domingos</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Mining time-changing data streams</article-title>
          .
          <source>In: Proceedings of the Seventh ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>97</volume>
          {
          <fpage>106</fpage>
          . ACM Press (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Khan</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Madden</surname>
            ,
            <given-names>M.G.</given-names>
          </string-name>
          :
          <article-title>A survey of recent trends in one class classi cation</article-title>
          .
          <source>In: Arti cial Intelligence and Cognitive Science - 20th Irish Conference. Lecture Notes in Computer Science</source>
          , vol.
          <volume>6206</volume>
          , pp.
          <volume>188</volume>
          {
          <fpage>197</fpage>
          . Springer (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Lucas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Bayesian analysis, pattern analysis, and data mining in health care</article-title>
          .
          <source>Current Opinion in Critical Care</source>
          <volume>10</volume>
          ,
          <volume>399</volume>
          {
          <fpage>403</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lucas</surname>
          </string-name>
          , P.,
          <string-name>
            <surname>van der Gaag</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>Hanna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Bayesian networks in biomedicine and health-care</article-title>
          .
          <source>Arti cial Intelligence In Medicine</source>
          <volume>30</volume>
          ,
          <volume>201</volume>
          {
          <fpage>214</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Mitchell,
          <string-name>
            <surname>T.M.:</surname>
          </string-name>
          <article-title>Machine Learning</article-title>
          .
          <source>McGraw-Hill</source>
          , international edn. (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Muthukrishnan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Data Streams: Algorithms and Applications</article-title>
          . Now Publishers Inc, New York, NY (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>P.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gama</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A system for analysis and prediction of electricity load streams</article-title>
          .
          <source>Intelligent Data Analysis</source>
          <volume>13</volume>
          (
          <issue>3</issue>
          ),
          <volume>477</volume>
          {496 (
          <year>June 2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>P.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gama</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedroso</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          :
          <article-title>Hierarchical clustering of time-series data streams</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>20</volume>
          (
          <issue>5</issue>
          ),
          <volume>615</volume>
          {627 (May
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Schurink</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lucas</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoepelman</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bonten</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Computer-assisted decision support for the diagnosis and treatment of infectious diseases in intensive care units</article-title>
          .
          <source>Lancet Infectious Diseases</source>
          <volume>5</volume>
          ,
          <issue>305</issue>
          {
          <fpage>312</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Shapiro</surname>
            ,
            <given-names>J.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gathers</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kannry</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kushniruk</surname>
            ,
            <given-names>A.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kuperman</surname>
          </string-name>
          , G.:
          <article-title>Survey of emergency physicians to determine requirements for a regional health information exchange network</article-title>
          . In: AMIA Spring Congress (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wyatt</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wright</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Design should help use of patients' data</article-title>
          .
          <source>Lancet(British edition)</source>
          <volume>352</volume>
          (
          <issue>9137</issue>
          ),
          <volume>1375</volume>
          {
          <fpage>1378</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>