<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Integrating Data Analysis Tools for Better Treatment of Diabetic Patients</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Svetla Boytcheva</string-name>
          <email>svetla.boytcheva@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Galia Angelova</string-name>
          <email>angelov@adiss-bg.com</email>
          <email>galia@lml.bas.bg</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhivko Angelov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitar Tcharaktchiev</string-name>
          <email>dimitardt@gmail.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Adiss Lab Ltd.</institution>
          ,
          <addr-line>Sofia</addr-line>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Information and Communication Technologies, Bulgarian Academy of Sciences</institution>
          ,
          <addr-line>Sofia</addr-line>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Medical University Sofia, University Specialized Hospital for Active Treatment of Endocrinology</institution>
          ,
          <addr-line>Sofia</addr-line>
          ,
          <country country="BG">Bulgaria</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Proceedings of the XIX International Conference “Data Analytics and Management in Data Intensive Domains” (DAMDID/RCDL'2017)</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <fpage>230</fpage>
      <lpage>237</lpage>
      <abstract>
        <p>This paper presents the construction and usage of an anonymous Diabetes Register for patients in Bulgaria. The Register is generated automatically from outpatient records submitted to the Bulgarian National Health Insurance Fund in 2010-2014 and continuously updated using outpatient records for 2015-2016. The construction relies on advanced automatic analysis of free text information as well as on Business Analytics technologies for storing, maintaining, searching, querying and analyzing data. Original frequent pattern mining algorithms enable to find patterns and sequences taking into account temporal information. The paper discussed the software environment as well as experiments in frequent pattern mining that enable knowledge discovery in the very large repository underlying the Register (currently 262 million pseudonymized outpatient records submitted to the Bulgarian National Health Insurance Fund in 2010-2016 for more than 5 mln citizens yearly). The claim is that the synergy of modern analytics tools transforms a static archive of clinical patient records to a sophisticated software environment for knowledge discovery and prediction.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Medicine is known as a Data Intensive Domain: due to
the recent penetration of the Information and
Communication Technologies (ICT) in all areas of our
society, a rapidly increasing amount of medical data is
produced by the healthcare sector, on the one hand, and
by biomedical research on the other hand. In the
healthcare sector, ICT applications support health
diagnostics, development and maintenance of medical
Electronic Health Records, telemedicine and telecare,
patient administration, almost all aspects of healthcare
management and healthcare delivery as well as medical
education and training. In biomedical research, progress
in deeper understanding of medical phenomena is sought
by construction of big data models: e.g. virtual
physiological human, models of brain, in computational
genetics and so on. Public access to health information is
changing the relationship between the patients and the
health institutions that are responsible for care delivery.
The monitoring and control function of patient
organizations is facilitated by the modern ICT tools as
well. Today we are still in an early phase of a long-term
technological and social shift that will be implied by
advancing further the ICT fundamentals and tools.</p>
      <p>In this paper we present the integration of various
ICT tools for automatic generation of a Diabetes Register
for Bulgarian patients. The huge amount of clinical data,
underpinning a repository of Outpatient Records (ORs),
enabled to construct interfaces that support both
monitoring functionalities (oriented to the health
management authorities) and research-oriented
functionalities for knowledge discovery. The monitoring
functionalities are based on business analytics while
research tools use data mining and pattern search. The
software environment includes also components for
automatic analysis of free texts in Bulgarian. These
components facilitate the Register generation and its
update because they deliver values of clinical tests and
lab data which are described as unstructured text only.</p>
      <p>This paper is structured as follows. Section 2
overview related work in several areas that are relevant
to the subject: Diabetes registers, Natural Language
Processing (NLP) for clinical narratives, Business
Intelligence (BI) and analytics, Frequent Pattern Mining
(FPM). Section 3 presents the experimental study context
and summarizes the developments during the last 3-4
years (because the Register was built iteratively). Section
4 presents relevant achievements in automatic analysis of
clinical narratives in Bulgarian language. Section 5
discusses recent algorithms for frequent pattern mining
and presents experiments related to knowledge discovery
in the Register repository. Section 6 contains the
conclusion and plans for future work.</p>
      <sec id="sec-1-1">
        <title>2.1 Diabetic Registers</title>
        <p>
          There are several nation-wide Diabetes Registers in the
world, e.g. in Denmark [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], Sweden [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], Norway [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
Registers explicate the number of patients who are
diagnosed with Diabetes and provide good monitoring
and control. Constructing registers is expensive and
burdening the patients as well as the medical experts with
additional administrative work. Furthermore, in some
countries chronic disease management is not recognized
as a part of general medical practice. As for the
construction, most medical experts agree that Registers
are a must since Diabetes is a chronic disease with
significant social consequences. Electronic patient
registration systems are proposed like the one in Ireland
[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] (but it is not implemented yet). It is interesting to
mention that in Sweden, during the Diabetic Register
development phase 2001-2005, the registration rate of
patients gradually increased and reached 75% which in
2010 still remains stable and is one of the highest in the
country [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Thus infrastructure construction is a critical
issue but data collection and update are further problems
that can be solved only by persistency and diligence.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>2.2 NLP of Clinical Narratives</title>
        <p>
          Usually automatic analysis of clinical narratives is
implemented partially: only fragments of the text are
considered. The phrases, selected as “interesting”, are
typically picked up due to the presence of a word or an
entity which are considered “significant”. This approach
for shallow analysis is called “Information Extraction”
(IE). IE from clinical texts matures only recently but its
accuracy gradually improves and often exceeds 90% [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
The review [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] stated in 2008 that “current applications
are rarely applied outside of the laboratories they have
been developed in, mostly because of scalability and
generalizability issues”. Today, however, this is valid for
languages other than English because, with the active
contribution of numerous research groups in the USA,
NLP for English clinical narratives has much better
performance at present. Comprehensive language
resources exist for English, such as UMLS [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] as well as
tools like KnowledgeMap Concept Identifier [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] which
processes clinical notes and returns CUIs (Concept
Unique Identifiers) for the recognized UMLS terms.
Another important tool is the public NegEx system
which identifies and interprets negations in English
clinical texts [
          <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
          ]. We also mention the open-source
cTAKES1 (clinical Text Analysis and Knowledge
Extraction System) and the Health Information Text
Extraction (HITEx) system [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. A recent study [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
enumerates the advantages to incorporate NLP for
English in medical systems: it systematically links
several terms to a concept using databases that
standardize health terminologies; avoids manual work
for searching term variations; increases the number of
patients in the considered cohorts and thus increases the
1 http://ctakes.apache.org/
sensitivity of the recognition. Despite the NLP
limitations, the conclusion is that NLP engines are
powerful components ready for integration in medical
data mining and – due to improvements expected in the
future, e.g. more accurate mappings of terms to medical
concepts – the importance of NLP as a valuable
supporting technology will grow.
        </p>
        <p>
          Here we consider NLP for Bulgarian clinical text. No
comprehensive resources exist for Bulgarian medical
language; the International Classification of Diseases
ICD-10 is the only terminological resource which is
available in electronic format. Our experience shows that
within 2-3 years one can achieve good performance in
separate extraction tasks. We apply software prototypes
developed some years ago that are gradually improved.
The most useful tools are a drug extractor (it finds in the
free text the drug name, dosage, frequency and route of
admission [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]) as well as an extractor of numeric values
of lab data and clinical tests [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
2.3 Big Data, Business Intelligence Tools
Big Data usually designates a massive volume of
structured and unstructured data, too large or too
dynamic to be processed by traditional software tools and
techniques. The popular "3Vs" features of Big Data were
first introduced by Gartner (previously META group):
"high Volume, high Velocity, and/or high Variety" [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
Wikipedia is an example for big data consisting of
unstructured texts, images and hyperlinks. Big data
analytics is the process of collecting, organizing and
analyzing big data to discover useful information.
Business Intelligence tools analyze big data of
enterprises in order to provide historic, present and
predicted views to the business processes. Predictive
analytics for establishment of trends is the preferred
functionality in contrast to databases that deal with data
items and extract subsets of data values. Visualization is
an important feature of BI tools because they show
generalizations and tendencies in one screen [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
Another necessary feature is the speed of processing
since big data often appear in real time.
        </p>
        <p>
          In our project we use a BITool which stores data in
n-dimensional cubes and explores multi-dimensional
data i.e. hyperplanes [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The user can split the dataset
into groups of objects with similar features. If temporal
dimension is included the user can track changes of
object characteristics over time by animation. BITool
enables the discovery of similar situations over time
when a search pattern is specified for a particular period.
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>2.4 Frequent Pattern Mining</title>
        <p>
          There are two principal tasks in pattern search: frequent
pattern mining (FPM) where the events (objects) are
considered as unordered sets, and frequent sequence
mining (FSM). Approaches for solving the FPM task
vary from the naïve BruteForce and Apriori algorithms,
where the search space is organized as a prefix tree, to
Eclat algorithm that uses tidsets directly for support
computation by processing prefix equivalence classes
[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Most FPM and FSM methods do not consider
contextual information about extracted patterns. They
usually build a (huge) prefix tree. Most FPM algorithms
generate all possible frequent patterns (FPs).
Summarized information for data relations can be
extracted as maximal frequent itemsets (MFI) in order to
reduce redundancy and decrease significantly the
number of FPs for post-analysis. All classic algorithms
for FPM can be modified for MFI search.
        </p>
        <p>
          We have proposed a novel algorithm for mining sets
of events in order to identify strong co-occurrence of
patterns [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. It is a cascade data mining approach for
FPM enriched with context information which aims at
the discovery of complex relations between medical
events with respective timestamps. Experiments with
this approach are presented in Section 5 to illustrate the
functionality of the Diabetes Register as a research tool.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3 Experimental Study</title>
      <sec id="sec-2-1">
        <title>3.1 Principal Objective</title>
        <p>
          A pseudonymized Register of diabetic patients was
generated in 2015 from the Outpatient Records, collected
by the Bulgarian National Health Insurance Fund
(NHIF), in compliance with all legal requirements for
safety and data protection [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ]. The usual patient
registration process was kept without burdening the
medical experts with additional paper work. NHIF is the
only obligatory Insurance Fund in Bulgaria so we note
that working with ORs ensures 100% registration of all
patients who contacted the healthcare system at all
(however there are Bulgarian citizens who are not
insured and some others who have ORs but are not
properly diagnosed with Diabetes). The data repository,
underpinning the Register, currently contains more than
262 mln pseudonymised ORs submitted to the NHIF in
2010-2016 for more than 7.3 mln Bulgarian citizens
(more than 5 mln yearly), including 483,836 diabetic
patients. In Bulgaria ORs are produced by General
Practitioners (GPs) and Specialists from Ambulatory
Care whenever they contact patients. Despite the primary
accounting purpose ORs summarize sufficiently the case
and motivate the requested reimbursement. They are
semi-structured files with predefined XML-format.
Many indicators in the Register copy the structured data
submitted to NHIF in ORs: (i) date and time of the visit;
(ii) pseudonymized personal data, age, gender; (iii)
pseudonymised visit-related information; (iv) diagnoses
in ICD-10; (v) NHIF drug codes for medications that are
reimbursed; (vi) a code if the patient needs special
monitoring; (vii) a code concerning the need for
hospitalization; (viii) several codes for planned
consultations, lab tests and medical imaging.
        </p>
        <p>ORs contain also important values presented in free
text fields: glycated haemoglobin (HbA1c), body mass
index (BMI), weight, blood glucose and blood pressure
etc. These values are essential for a Diabetic Register so
2 https://www.drugs.com/drug-class/incretin-mimetics.html
they are extracted automatically from four XML fields:
(i) Anamnesis: summarizes case history, previous
treatments, often family history, risk factors; (ii) Status:
summary of patient state, height, weight, BMI, blood
pressure etc.; (iii) Clinical tests: values of clinical
examinations and lab data listed in arbitrary order; (iv)
Prescribed treatment: codes of drugs reimbursed by
NHIF, free text descriptions of other drugs. Integration
of large scale text analysis is a real novelty in this field.</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2 Analytics Using BITool</title>
        <p>Today the system BITool supports the Diabetes Register
at the University Specialized Hospital for Active
Treatment of Endocrinology ″Acad. Ivan Penchev″,
Medical University – Sofia (this Hospital was authorized
by the Bulgarian Ministry of Health to host the Register
of diabetic patients in Bulgaria). BITool’s functionalities
enable the monitoring of significant indicators like
glycated hemoglobin (HbA1c) and blood glucose values.
In this way the Register achieves its objective: to provide
an adequate monitoring strategy for diabetic patients and
to improve the healthcare and quality of life for the
patients and their families. Two examples illustrate the
services. Figure 1 shows the number of diabetic patients
in the dimensions age-gender (at certain moment). Here
BITool operates on the structured information from the
NHIF archive: patient pseudonym, age and gender.
Further statistics of this kind might concern explorations
of diabetic patients per region code, types of diabetes and
diabetes complications, per GPs, per types of medication,
according to frequency of visits etc.</p>
        <p>Figure 2 Reduction of HbA1c levels after application
of incretin2 based drugs
4 NLP for Bulgarian Clinical Narratives
Design and implementation of software for automatic
extraction of patient-related entities from a Big Data
collection is a quite challenging task. One needs to scale
up existing research prototypes to process millions of
patient records, coping with noisy and missing data, and
still providing reliable results. Some numeric entities
refer to key risk factors for development of Diabetes
Mellitus (levels of glycated hemoglobin HbA1c and
blood glucose) and cardio-vascular diseases (high blood
pressure). Unfortunately in the Bulgarian clinical
practice these values are usually documented in free text
paragraphs, presented in a huge variety of formats, so
their automatic identification is difficult. We note that
according to some studies, today more than 80% of the
patient-related clinical information is stored as free text
in the Electronic Health Record systems.</p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] we proposed a hybrid method for automatic
generation of grammar rules for IE from clinical data.
The experiments were made and evaluated over
approximately 9.5 million of ORs. Here we cite only the
evaluation of blood pressure extraction from the ORs of
about 1,800,000 patients with arterial hypertension for 3
year period: all available values are about 38.3 million
and the extraction was performed with precision 92%
and recall 98%. The variety of recording formats and
explanations written by thousands of medical
professionals require constant evaluation of grammar
coverage and extraction accuracy in general. Some of the
main advantages of the proposed method, beyond its
reliable performance and good precision in text mining,
are the modularity, extensibility, and scalability.
5 Research in Frequent Pattern Mining
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>5.1 Contextual Information</title>
        <p>Most FSM and FPM approaches do not use contextual
information about extracted patterns. These algorithms
extract general templates but do not answer the major
question whether they are influenced in some way by the
context and whether they are valid in various aspects.
Existing methods which search for patterns using
contextual information are based on attributes that are
organized into hierarchical structures and on attributes’
generalizations and specializations.</p>
        <p>
          Context information is organized as attributes of
itemsets and tidsets. Attributes may have different
organization - structured or unstructured. This enables to
explore the context-dependent templates. Rabatel et al.
[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] propose an approach in marketing domain taking
into account not only the transactions that have been
made but also various attributes associated with
customers like age, gender etc. Attributes have a
hierarchical structure ((), () ) and
explore patterns at different levels of attributes
abstraction – lattice  (Figure 4). Traditional methods
consider only the top level [∗,∗] - for any age and
regardless of gender, i.e. without attributes. Rabatel et al.
designed the algorithm Gespan and made experiments
with about 100,000 product descriptions from
amazon.com.
        </p>
        <p>H(age)</p>
        <p>*
young
(y)
old
(o)</p>
        <p>H(gender)</p>
        <p>*
male
(m)
female
(f)
[y,*]
[y,m]</p>
        <p>[*,*]
[o,*]
[y,f]
[*,m]
[o,m]</p>
        <p>H
[*,f]
[o,f]</p>
        <p>
          Ziembiński [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] proposes a new approach for
extracting small contextual models from smaller
collections of data that later are summarized in
generalized models using information from contextual
models with common information. This approach applies
a metrics for measuring distance of context models. All
values for similarity assessment are normalized in the
range between 0.0 and 1.0. Attribute values are
considered identical if the similarity function returns 1.0.
In the opposite case the result is 0.0. This approach
allows extracting patterns for data that would otherwise
have to be dropped out of the templates because of its
dispersion and low frequency.
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>5.2 Experimental Setup</title>
        <p>We apply a retrospective analysis for patients from the
Diabetes Register with Diabetes Type 2. The period of
interest is two years preceeding the onset of the Diabetes
Type 2, i.e. the so called prediabetes condition. In order
to illustrate the potential of contextualized FPM we
present results in searching comorbidities for patients in
prediabet condition. Text mining modules are used to
convert raw text descriptions to structured event data.
Text
Mining
Structured
Information
Processing</p>
        <p>Outpatient</p>
        <p>Records</p>
        <sec id="sec-2-4-1">
          <title>Preprocessing</title>
          <p>Structured
Information
Processing</p>
        </sec>
        <sec id="sec-2-4-2">
          <title>Data Analysis</title>
          <p>Association</p>
          <p>Rules
Generation
Data
Modeling</p>
          <p>Context
Information
Processing</p>
        </sec>
        <sec id="sec-2-4-3">
          <title>Prediction &amp; Prevention Models</title>
          <p>Comorbidity</p>
          <p>Analysis</p>
          <p>Risk Factors
visit to a doctor as a single event. For each patient   ∈ 
an event sequence of tuples 〈, 
〉 is
  ∈ 
generated:  (  ) = (〈 1,  1〉, 〈 2,  2〉, … , 〈  
,    〉),  =
̅1̅,̅̅̅. Let ℰ be the set of all possible events and  be the
set of all possible timestamps. Let  = {  1,  2, … ,   }
be the set of all diseases ICD-103 codes, which we call
items. Each subset  ⊆</p>
          <p>is called an itemset. We define
a projection function : (ℰ ×  )
→ 2 : (
(  )) =
 (  ) = ( 1i,</p>
          <p>2i, … ,    ), such that for each patient
the projected time sequence contains only the
first occurrence (onset) of each disorder recorded in
3 International Classification of Diseases and Related
Health Problems 10th Revision. http://apps.who.int/
collection</p>
          <p>after
〈
,</p>
          <p>be the set of all itemsets in our
projection
in
the</p>
          <p>format
〉. We shall call 
a database. We are
looking for itemsets  ⊆ 
with frequency (sup() )
above given . Let
itemsets, i.e. ℱ = { |  ⊆   sup() ≥ }
ℱ denote the set of all frequent
. A
frequent itemset  ∈ ℱ</p>
          <p>is called maximal if it has no
frequent supersets. Let ℳ denote the set of all maximal
frequent
ℱ, ℎ ℎ
itemsets,
  ⊂ }
i.e.</p>
          <p>ℳ = { |  ∈ ℱ  ∄  ∈
. Let 2 denote the power set (set
of all subsets) of itemset . Then each subset of  ∈ ℱ
is
also a frequent itemset, i.e. ∀  ∈ 2   ℎ
  ∈
ℱ . For each item</p>
          <p>∈  we define the set called pidset:
 (id) = {  | 〈  ,  (  )〉 ∈  
∈  (  )}.</p>
          <p>To study the nature of comorbidities we need to
investigate the context in which they occur. Therefore we
add
some
semantic
attributes to
each
event
–
demographics of patients, age and gender, treatment,
status, lab data and etc.</p>
          <p>We define a set of attributes of interest  =
{ 1,  2, … ,   }. Context Q for some patient   ∈ 
defined as the set of attribute-value pairs from patient
is
profile information:</p>
          <p>(  ) = {〈 1,  1〉, 〈 2,  2〉, … , 〈  ,   〉}.</p>
          <p>In order to decrease the number of possible values of
attributes we apply some aggregation of data. For
instance age value is categorized according to the World
Health Organization (WHO) standard age groups. Data
for body
mass index (BMI) are also
categorized
according to the</p>
          <p>WHO4 standard
classification
underweight, normal weight, overweight, obesity.</p>
          <p>For some data concerning demographic information,
like region ID we have large number of distinct values.
For such data we add also some additional properties
concerning background information for the region – e.g.
whether it is south, north, west, east, central, northwest
etc., and mountain, river, sea, thermal spring, urban
region etc. For status and clinical test data we take the
worst value for the period, according to the risk factors
deinition.</p>
          <p>In primary interest for Diabetes Type 2 are BMI,
glycated haemoglobin, blood pressure (RR – Riva Roci),
blood glucose, HDL-cholesterol.</p>
          <p>From  (  ) we generate a feature vector  (  ) =
( 1 ,  2 , … ,   ), where each attribute   ∈ 
with  
possible values is represented by  
consecutive
positions in the vector. For the set of maximal frequent
itemset ℳ</p>
          <p>
            with cardinality |ℳ| = K we have K classes
of comorbidities. We apply classification of multiple
classes in order to generate rules for each comorbidity
class. We use large scale
multi class classification
because we deal with a big database and a large group of
comorbidity classes. We use Support Vector Machines
(SVM) and optimization based on block minimization
method described by Yu et al. [
            <xref ref-type="bibr" rid="ref24">24</xref>
            ].
4 WHO, BMI Classification http://apps.who.int/bmi/
index.jsp?introPage=intro_3.html
2013-2014 were excerpted from the Diabetes Register
when, as we assume, these patients were in a pre-diabetes
condition. The idea of this experiment is to check
whether we can successfully discover risk factors for
these patients looking only at their ORs in 2013 and
2014. Then, maping our hypotheses to the real data for
2015, we test whether our approach is reasonable. (We
note that due to the relatively short period of observation
and lack of data about mortality, at the moment we
cannot follow diabetes development in longer periods.)
          </p>
          <p>In the Register each OR, corresponding to a single
visit, cointains up to 4 diagnoses encoded in ICD-10.
Some diagnoses are presented by 4-sign encodings, i.e.
in a more specific way, while others use the more general
3-sign encoding. Due to the hierarchical organization of
ICD-10 we shall analyse individually two collections:
the original one, that is more specific (with 4-sign codes
- see Example 1) and we shall generalise also all
diagnoses to more general classes (with 3-sign codes
see Example 2). The examples present collections of
diagnoses for a patient with ID 2196365.</p>
          <p>Example 1:
L94.1,</p>
          <p>L94,
M51.1, M33.9}</p>
          <p>Example 2:
I(2196365)={I10,</p>
          <p>M10.9,</p>
          <p>M10, K76.9, K76,
code I10).</p>
          <p>M06.9,</p>
          <p>G57.9,</p>
          <p>Z00.8,</p>
          <p>H53,
I(2196365)={I10, M10, K76, L94, M06, M51,
M33, H53, Z00, G57}</p>
          <p>For some patients, the available ORs contain no
information about certain attributes of the context
information (Table 2). It is well known that missing data
in</p>
          <p>medical documentation is inevitable. Thus some
attribute values are replaced by the value NA, which is
considered as the most general value.</p>
          <p>For example the context information for the patient
with ID 2196365 is:
 (2196365) = {〈,
〈,
〈
_, 6.39
03〉, 〈,</p>
          <p>58〉, 〈, 1
29.32〉, 〈ℎ
〉, 〈ℎ _ℎ, 1.15
1, 
〉,
〉,
〉}
AGE
15-44
45-59
60-69
70-89
Figure 6 Age of the patients in the support set of
"MFI#12"</p>
          <p>HDL CHOLESTEROL
low high
21% 23%
medium</p>
          <p>56%</p>
          <p>Data about HbA1c are available only for 3 out of 453
patients, that is why we consider this attribute as a more
general value ANY. But we note that the lack of HbA1c
measurements is not surprising because tests for HbA1c
are made when the Diabetes is diagnosed (and this has
happened in 2015 for the selected patient cohort).</p>
          <p>Data for blood glucose are available only for 30% of
these patient and for 50% of them the values were high.</p>
          <p>Deeper analyses reveal medical arguments why
higher risk exist especially for the patients in the support
set of MFI#12: Z00 I10 M51 #SUP: 453. The diagnose
5 http://usbale.com/Register_Diabetes.htm
with ICD-10 code M51 (Thoracic, thoracolumbar, and
lumbosacral intervertebral disc disorders) means that the
patients have lower motor activity and sedentary
lifestyle, which causes obesity, overweight, higher
values of cholesterol and blood pressure and therefore
increases the risk of developing Diabetes. Actually this
has happened in 2015. We note that in general the
ICD10 diagnose M51 is not considered risky for Diabetes.
But our algorithm reveals this unknown and latent
interrelationship.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>6 Conclusion and Future Work</title>
      <p>In this paper we present a software environment for
collection and processing of Big Data in medicine - a
Data Intensive Domain. The Diabetes Register has been
developed stepwise and its research functionality is still
under construction. We believe that the integration of
various technologies is the proper way to approach the
challenges of large-scale information processing because
the integration ensures flexible multi-functionality and
enables reuse of results.</p>
      <p>The nation-wide Diabetes Register of Bulgaria is now
visible in Internet5 together with some public statistical
information. We plan to develop the Register further as a
predictive and preventing tool using the synergy of
advanced technologies which enable to discover risk
groups of patients that have predisposition to various
socially-significant diseases. We have shown here that
the present software environment is mature enough to
identify patients with complexes of risk factors for
development of Diabetes, e.g. risks like: family history
(relatives with Diabetes); obesity; arterial hypertonia
(RR&gt;140/90); low physical activity; giving birth to a
baby with weight more than 4 kg or gestational Diabetes;
established impaired fasting glycaemia or impaired
glucose tolerance; other states of insulin resistance (e.g.
acanthosis nigricans, a specific hyperpigmentation of the
skin that might be due to endocrine disorders);
HDLcholesterol≤0.90 mmol/l or triglycerides≥2.2 mmol/l
(≥2.82 mmol/l according to ADA); diagnosed polycystic
ovarian syndrome, a cardio-vascular disease, or mental
disorders etc. These risk factors are explicated in the
patient-related documents either by values of clinical
tests or by keywords and typical phrases that describe the
factor. The patients with predisposition suffer from
disorders and syndromes, diagnosed by various medical
specialists in various time periods, but without any
chance to establish connections between the medical
doctors – e.g. a connection between a Psychiatrist and a
Cardiologist that have consulted the patient. Elaborating
further the analytics facility of the Register will provide
functionality to monitor patient status over time, in the
context of all available information, and to issue alerts
for coincidence of risk factors that open the door to
Diabetes and other chronic diseases. In this way we
believe that in the foreseeable future it will become
possible to identify the Bulgarian citizens who have
predisposition to develop Diabetes Mellitus.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>The research presented here is partially supported by the
grant 02/4 SpecialIZed Data MIning MethoDs Based on
Semantic Attributes (IZIDA), funded by the National
Science Fund in 2017–2019. The support of Medical
University – Sofia, the Bulgarian Ministry of Health and
the National Health Insurance Fund is acknowledged.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Carstensen</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al.:
          <article-title>The Danish National Diabetes Register: Trends in incidence, prevalence and mortality</article-title>
          .
          <source>Diabetologia</source>
          .
          <volume>51</volume>
          (
          <issue>12</issue>
          ),
          <fpage>2187</fpage>
          -
          <lpage>2196</lpage>
          (
          <year>2008</year>
          ). doi:
          <volume>10</volume>
          .1007/s00125-008-1156-z
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Hallgren</given-names>
            <surname>Elfgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. M.</given-names>
            ,
            <surname>Grodzinsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Törnvall</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          :
          <article-title>The Swedish National Diabetes Register in clinical practice and evaluation in primary health care</article-title>
          .
          <source>Prim. Health Care Res. Dev</source>
          .
          <volume>17</volume>
          (
          <issue>6</issue>
          ),
          <fpage>549</fpage>
          -
          <lpage>558</lpage>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .1017/S1463423616000098
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Cooper</surname>
            ,
            <given-names>J. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thue</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Claudi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Løvaas</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carlsen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sandberg</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The Norwegian Diabetes Register for Adults - an overview of the first years</article-title>
          .
          <source>Norsk Epidemiologi</source>
          .
          <volume>23</volume>
          (
          <issue>1</issue>
          ),
          <fpage>29</fpage>
          -
          <lpage>34</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>O</given-names>
            <surname>'Mullane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>McHugh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Bradley</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. P.</surname>
          </string-name>
          :
          <article-title>Informing the development of a national diabetes register in Ireland: a literature review of the impact of patient registration on diabetes care</article-title>
          .
          <source>Inform. Primary Care</source>
          .
          <volume>18</volume>
          (
          <issue>3</issue>
          ),
          <fpage>157</fpage>
          -
          <lpage>68</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Hallgren</given-names>
            <surname>Elfgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.M.</given-names>
            ,
            <surname>Törnvall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Grodzinsky</surname>
          </string-name>
          ,
          <string-name>
            <surname>E.</surname>
          </string-name>
          :
          <article-title>The process of implementation of the diabetes register in Primary Health Care</article-title>
          .
          <source>Int. Journal of Qual. Health Care</source>
          .
          <volume>24</volume>
          (
          <issue>4</issue>
          ),
          <fpage>419</fpage>
          -
          <lpage>424</lpage>
          (
          <year>Aug 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Meystre</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kipper-Schuler</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hurdle</surname>
            ,
            <given-names>J. F.</given-names>
          </string-name>
          :
          <article-title>Extracting Information from Textual Documents in the Electronic Health Record: A Review of Recent Research</article-title>
          .
          <source>IMIA Yearbook of Medical Informatics</source>
          , pp.
          <fpage>138</fpage>
          -
          <lpage>154</lpage>
          . (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>[7] UMLS, the Unified Medical Language System</article-title>
          . https://www.nlm.nih.gov/research/umls/
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Denny</surname>
            ,
            <given-names>J. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Irani</surname>
            ,
            <given-names>P. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wehbe</surname>
            ,
            <given-names>F. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smithers</surname>
            ,
            <given-names>J. D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spickard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The KnowledgeMap Project: Development of a Concept-Based Medical School Curriculum Database</article-title>
          .
          <source>In: AMIA Annu Symp Proc.</source>
          , pp.
          <fpage>195</fpage>
          -
          <lpage>199</lpage>
          . (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Chapman</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bridewell</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hanbury</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cooper</surname>
            ,
            <given-names>G. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchanan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>A Simple Algorithm for Identifying Negated Findings and Diseases in Discharge Summaries</article-title>
          .
          <source>Univ. of Pittsburgh</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Gindl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          : Negation Detection in
          <source>Automated Medical Applications. TUW</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>[11] HITEx Manual: https://www.i2b2.org/software/ projects/hitex/hitex_manual.html</mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Liao</surname>
            ,
            <given-names>K. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cai</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Savova</surname>
            ,
            <given-names>G. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murphy</surname>
            ,
            <given-names>S. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karlson</surname>
            ,
            <given-names>E. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ananthakrishnan</surname>
            ,
            <given-names>A. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gainer</surname>
            ,
            <given-names>V. S.</given-names>
          </string-name>
          et al.:
          <article-title>Development of phenotype algorithms using electronic medical records and incorporating natural language processing</article-title>
          .
          <source>British Med</source>
          . J.,
          <volume>350</volume>
          (
          <issue>1</issue>
          ): h1885 (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Boytcheva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Shallow Medication Extraction from Hospital Patient Records</article-title>
          .
          <article-title>Studies in Health Technology and Informatics</article-title>
          . vol.
          <volume>166</volume>
          , pp.
          <fpage>119</fpage>
          -
          <lpage>128</lpage>
          . IOS Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Tcharaktchiev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boytcheva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelov</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zacharieva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Completion of Structured Patient Descriptions by Semantic Mining</article-title>
          .
          <source>Studies in Health Technology and Informatics</source>
          , vol.
          <volume>166</volume>
          , pp.
          <fpage>260</fpage>
          -
          <lpage>269</lpage>
          . IOS Press (
          <year>2011</year>
          ). doi:
          <volume>10</volume>
          .3233/978-1-
          <fpage>60750</fpage>
          -740-6-260
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Laney</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>3D Data Management: Controlling Data Volume, Velocity, and Variety</article-title>
          . META Group Research Note,
          <volume>6</volume>
          ,
          <issue>10</issue>
          (
          <year>2001</year>
          ) https://blogs.gartner.com/doug-laney/files/2012/ 01/ad949-3D
          <article-title>-</article-title>
          <string-name>
            <surname>Data-Management-ControllingData-Volume-Velocity-</surname>
          </string-name>
          and-Variety.pdf
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>[16] Top 238 Business Analytics Tools. Predictive Analytics Magazine (Feb</source>
          <year>2012</year>
          ). http://www.predictiveanalyticstoday.com/topbusiness-intelligence-tools/
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Angelova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nikolova</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , Angelov, Zh.:
          <article-title>Embedding language technologies in a data analytics tool</article-title>
          .
          <source>Advances in Bulgarian Sciences</source>
          , pp.
          <fpage>29</fpage>
          -
          <lpage>42</lpage>
          .
          <article-title>National Centre for Information and Documentation (</article-title>
          <year>2016</year>
          ). ISSN:
          <fpage>1314</fpage>
          -
          <lpage>3565</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Nasreen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Azam</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shehzad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Naeem</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghazanfar</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          :
          <article-title>Frequent Pattern Mining Algorithms for Finding Associated Frequent Patterns for Data Streams: A Survey</article-title>
          .
          <source>Procedia Computer Science</source>
          ,
          <volume>37</volume>
          ,
          <fpage>109</fpage>
          -
          <lpage>116</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Boytcheva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelov</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tcharaktchiev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Mining Comorbidity Patterns Using Retrospective Analysis of Big Collection of Outpatient Records</article-title>
          .
          <source>Health Inf Sci Syst. Journal</source>
          , Springer (
          <year>2017</year>
          ). ISSN:
          <fpage>2047</fpage>
          -
          <lpage>2501</lpage>
          (to appear)
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Tcharaktchiev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zacharieva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boytcheva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , et al.
          <article-title>Building a Bulgarian National Registry of Patients with Diabetes Mellitus</article-title>
          .
          <source>Journal of Social Medicine</source>
          .
          <volume>2</volume>
          ,
          <fpage>19</fpage>
          -
          <lpage>21</lpage>
          (
          <year>2015</year>
          )
          <article-title>(in Bulgarian)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Boytcheva</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelova</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelov</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tcharaktchiev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Text Mining and Big Data Analytics for Retrospective Analysis of Clinical Texts from Outpatient Care</article-title>
          .
          <source>Cybernetics and Information Technologies</source>
          ,
          <volume>15</volume>
          (
          <issue>4</issue>
          ),
          <fpage>58</fpage>
          -
          <lpage>77</lpage>
          (
          <year>2015</year>
          ). doi:
          <volume>10</volume>
          .1515/cait-2015
          <source>-0055</source>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Rabatel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bringay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poncelet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Mining sequential patterns: a context-aware approach</article-title>
          .
          <source>Advances in Knowledge Discovery and Management</source>
          , pp.
          <fpage>23</fpage>
          -
          <lpage>41</lpage>
          . Springer (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Ziembiński</surname>
            ,
            <given-names>R. Z.</given-names>
          </string-name>
          :
          <article-title>Accuracy of generalized context patterns in the context based sequential patterns mining</article-title>
          .
          <source>Control and Cybernetics</source>
          .
          <volume>40</volume>
          ,
          <fpage>585</fpage>
          -
          <lpage>603</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>H. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hsieh</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>K. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C. J.:</given-names>
          </string-name>
          <article-title>Large linear classification when data cannot fit in memory</article-title>
          .
          <source>ACM Transactions on Knowledge Discovery from Data (TKDD)</source>
          ,
          <volume>5</volume>
          (
          <issue>4</issue>
          ),
          <volume>23</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>