<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Application of medical data classification methods for a medical decision support system*</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>EC-leasing Company</institution>
          ,
          <addr-line>125, Varshavskoe highway, Moscow, 117587, Russian Federation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Research University Higher School of Economics</institution>
          ,
          <addr-line>11, Pokrovsky Boulevard, Moscow, 101000, Russian Federation</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1956</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>Decision support systems (DSS) allow us to help the doctor in making diagnoses to the patient, also medical DSS help to assess the need for a particular examination of the patient. In this article methods of medical data classification are considered, these methods are the part of the medical DSS. The paper includes investigation of data classification methods as hierarchical cluster analysis, k-means analysis and discriminant analysis. The selected methods are implemented using the example of cardiological data. A hypothesis is put forward that it is possible to determine the presence or absence of tuberculosis in a person from cardiological data by using data classification methods. Such indicators as sensitivity and specificity evaluate the effectiveness of the methods. In addition, ROC and AUC are presented. Thus, the DSS will be able to determine a certain degree of probability to assume the presence of tuberculosis in a person. The doctor will decide on the need for additional examinations depending on the values obtained,</p>
      </abstract>
      <kwd-group>
        <kwd>Decision Support System</kwd>
        <kwd>Data Analysis</kwd>
        <kwd>Telemedicine</kwd>
        <kwd>Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Currently, the creation of decision support systems (DSS) is relevant, and this
direction is also developing in the field of medicine. DSS allow us to help the doctor
in making diagnoses to the patient. In addition, with the help of these systems, it is
possible to determine the need for various examinations for the patient [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The use of
the medical DSS for doctors will prevent patients from being sent to expensive
additional examinations, which are not always safe [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The paper discusses methods of data analysis that will be implemented in the
medical DSS in order to help doctors. The paper implements such methods as
hierarchical cluster analysis, k-means analysis and discriminant analysis.</p>
      <p>The parameters of electrocardiogram (ECG) were used as experimental data. This
data is depersonalized.</p>
      <p>With implementing the methods of medical DSS a hypothesis is put forward about
the possibility of predicting the presence or absence of tuberculosis by ECG
parameters. The sample contains a nominal variable (tb), which reflects the presence
of diagnosed tuberculosis in a person (tb = 1) or its absence (tb = 0). The
experimental data collected ECGs recorded in people with a confirmed form of
tuberculosis in the second stage of the disease.</p>
      <p>
        Currently ETU “LETI” under the leadership of Professor Kalinichenko A. N. is
doing the similar studies. However, investigations of ETU “LETI” have a direction
different from this work. They research the detection of signs of cardiac disease in
ECG using machine learning [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The sample used in the study is divided into training and test samples, where the
mathematical model is created on the training one, and the quality of the obtained
model is evaluated on the test one. As a result, a model and an accuracy value of the
correct prediction of belonging to the group are obtained for each of the considered
methods.</p>
      <p>An approach with training of DSS methods based on medical data will make it
possible to make an early diagnosis of the patient's health condition. This means that
it is possible to assume with a certain degree of probability that a person has signs of
tuberculosis or not according to the recorded ECG. If possible signs of the disease are
detected, this patient should be sent to get a more detailed examination together with a
pulmonologist.</p>
      <p>The purpose of this work is to implement classification methods to the medical
SPR for early diagnosis to determine the presence or absence of tuberculosis signs. In
accordance with this goal, the following tasks were identified: to compile descriptive
statistics of the initial experimental data, to investigate and apply methods for
classification on experimental data, and to formulate a conclusion.</p>
      <p>
        The performance of the methods is evaluated using sensitivity and specificity
indicators. Sensitivity is the percentage of correctly classified "ill" people, and
specificity is the percentage of correctly classified "healthy" people [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In addition, a
ROC is constructed for the results of the methods and the area under the curve (AUC)
is calculated. The ROC curve is a tool for assessing diagnostic ability, representing a
graph where the sensitivity and specificity values in the range from 0 to 1 are taken as
axes [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
2
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Materials and methods</title>
      <sec id="sec-2-1">
        <title>Materials</title>
        <p>General information of the person and parameters of his cardicycle were taken to
study the methods of data classification and subsequent verification of the proposed
hypothesis about the possibility of determining the presence or absence of
tuberculosis in a person from cardiological data.</p>
        <p>
          A cardiocycle (or cardiac cycle) is a period of blood circulation generated by the
cyclic activity of the heart. The measurement unit for this periodicity is one cardiac
cycle. The length of the cardiocycle is the period of cardiac contractions [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The
elements of the cardiocycle are presented below (Fig. 1).
        </p>
        <p>The main elements for the ECG analysis are the start and end time of the elements
of the cardiocycle, as well as the PQ interval, the QRS complex, the ST segment, the
QT interval and the P wave stands out especially among the elements of the
cardiocycle.</p>
        <p>The data is presented in the form of a table, where each row corresponds to one
ECG. Also in the same row in columns contains non-personally-identifying
information about the person. Below is a table with parameters for analysis and
explanations to them (Table 1).</p>
        <p>The total sample consists of 5928 registered ECGs. Below is a table with general
information about people from the data sample (Table 2).</p>
        <p>The distribution of the number of people according to their age classification is
also presented (Fig. 2). In the sample by age, there is a bias towards people over 18
years old. This is explained by the fact that ECG registration of persons under 18
years old is possible in the presence of a parent and with his permission, so there were
few persons of younger groups in the collected data.</p>
        <p>All</p>
        <p>The following table includes the main values of the number of ECGs for different
groups (Table 3).
l
l</p>
        <p>A</p>
        <p>It should be noted that not all parameters of the cardiocycle were calculated for all
ECGs, so observations with partially uncalculated parameters were automatically
discarded methods implementation.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Methods</title>
        <p>
          This section describes methods of data analysis and provides brief information on
them [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>
          Cluster analysis is used to separate the original data into groups (clusters) that are
amenable to interpretation so that the elements of one group were similar in the
parameters, while elements from different groups should differ from each other [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
(Fig. 3).
        </p>
        <p>Hierarchical cluster analysis is used for relatively small numbers of observations.
During the analysis, initially each observation is located in its own cluster, then
neighboring clusters are combined in pairs until there are only two clusters left.</p>
        <p>
          K-means analysis allows you to divide an arbitrary data set into a given number of
groups so that the objects of the same cluster are close enough to each other, and the
objects of different groups do not intersect [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. In this case, the observation belongs to
the cluster to the center of which it is the closest.
        </p>
        <p>First, the center of the class is determined, then all objects within the specified
threshold value from the center are grouped.</p>
        <p>
          Discriminant analysis is a method of statistical analysis that allows you to divide
data into disjoint groups. This method allows us to identify the variables that affect
the separation, as well as their weight coefficients [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The result of performing a
discriminant analysis is a discriminant function that uses a nominal dependent
variable. Discriminant analysis is an alternative to multiple regression analysis.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        IBM SPSS Statistics 23 software was used in order to implement selected methods.
IBM SPSS Statistics is a statistical analysis platform with a set of functions [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>The use of the hierarchical cluster analysis method did not lead to significant
results. The number of observations in 5928 recorded ECGs was too large as a sample
for this method.</p>
      <p>Further, the number of observations in the sample was reduced to 50% of randomly
selected observations. As an assumption, a range was set for the number of classes:
there should be 2-3 clusters. This was done because a huge number of clusters are
obtained on this sample without this restriction. The values in two or three clusters
were chosen based on the following: we need to get two clusters with measurements
of people without tuberculosis and with tuberculosis. The possible number of three
clusters is taken to compare the results.</p>
      <p>The model was built, its variables were saved with the indication of belonging to
the cluster. A conjugacy table was constructed for two variables containing
information about the distribution of all variables into 2 and 3 clusters in order to
evaluate the performance of this method (Figure 4).</p>
      <p>According to the table almost all observations fell into the first cluster, it also
observed when the data divided into two clusters and into three clusters. It is also
seen that the second and third clusters in both divisions are very small relative to the
first cluster. If we compare the decision to divide into two clusters or three, we can
conclude that the two-cluster solution is the most stable.</p>
      <p>Further, for the sake of clarity of the obtained solution an analysis of the averages
was carried out, the part of the resulting picture is presented below (Fig. 5).</p>
      <p>The target variable – the variable of the presence or absence of tuberculosis in this
dimension (tb). During the application of hierarchical clustering it was found that the
average values of clusters are 0.24 and 0.73, where 0 is “healthy” and 1 is “sick”. You
can also pay attention to other parameters, for example, taller people fell into the
“healthy” cluster, and almost all smokers fell into the “sick” cluster. It is worth noting
the average weight of the subjects in the second cluster – 138 kg, which is quite a lot.</p>
      <p>When analyzing this model in detail we conclude that hierarchical clustering is not
suitable for working with this sample of medical data.</p>
      <p>When implementing the k-means method two clusters are initially set: people
without a diagnosis of tuberculosis and people with diagnosed tuberculosis (parameter
tb=0 and tb=1, respectively). There are many observations in the medical dataset, so
10 iterations were set for the method to work.</p>
      <p>The centers of the two clusters obtained are presented below (Fig. 6).</p>
      <p>We also obtained an estimate based on Fischer statistics on the significance of the
parameter in the differentiation of clusters (Fig. 7).The figure below reveals an
example that shows that the target variable tb is significant, as is weight, height, and
smoking. The most significant parameters among the parameters of the cardiocycle
are the R wave, S wave and the QRS complex.</p>
      <p>In addition, the k-means method obtained results is similar to the hierarchical
clustering method: outputs data on the number of observations in clusters are 2331
observations in the first cluster and 30 observations in the second cluster. The
obtained values coincided with the values for the number of observations when
dividing into two clusters during hierarchical clustering. These calculations were
obtained by randomly selecting 50 % of all observations.</p>
      <p>When using the k-means method with the same parameters on a full sample the
following division was obtained by the number of observations in clusters: 4707 and
56 observations, respectively. Thus, the result was obtained that one cluster is
dominated by data when clustering into two groups.</p>
      <p>Two additional parameters were created using of the k-means clustering method:
indicating the number of the membership cluster and the distance to its center.</p>
      <p>Next a graphical illustration of the results of this method was constructed: the
grouping variable is the cluster number, the differentiating variable is the distance to
the cluster center. The figure below, as well as the line on the cluster, shows the
median value (Fig. 8).</p>
      <p>Based on the results of the analysis it can be concluded that the parameters of the R
wave were the most significant parameters in clustering by this method. Below is the
spread of the R wave indicator depending on the presence or absence of tuberculosis
(Fig. 9).</p>
      <p>The figure shows that the variations of this indicator differ depending on the
presence or absence of tuberculosis, but visually almost half of the values of the
indicator are the same both in the presence of tuberculosis and in its absence.
According to the results of the clustering analysis the k-means indicator is the most
significant when divided into groups. This leads to the conclusion that the model is
not sufficiently accurate using the k-means method.</p>
      <p>
        Then the discriminant method was implemented and investigated. The sample was
first divided into training and test samples: 60% and 40%, respectively [
        <xref ref-type="bibr" rid="ref12 ref13">12-13</xref>
        ]. For
implementation the method of forced inclusion of variables was used and grouping
was performed by the variable of the presence or absence of tuberculosis tb.
      </p>
      <p>A table "Group statistics" was obtained with an indication of the average values of
each parameter, its standard deviation by group. The inequality of the mean and
standard deviation does not prove that these variables are distinctive features of the
selected clusters.</p>
      <p>The figure below shows the calculated values of the variables. Parameters whose
values in the table are greater than 0.05 can be excluded later for analysis purposes.</p>
      <p>From the figure below it can be seen that there are parameters that are insignificant
when divided into groups, for example, p_da, t_da and others. Thus, they can be
excluded when composing the equation (Fig. 10).</p>
      <p>The coefficients of the canonical discriminant function were also obtained to create
the equation (Fig. 11).</p>
      <p>The accuracy of the division into clusters is determined by the distance between
the average values of the discriminant function in the studied clusters. The greater the
distance, the better the groups are separated. The values of the centroids of the groups
are as follows: -0.376 and 1.242.</p>
      <p>You can determine the quality of the model based on the results of the
classification at the following table (Table 4).</p>
      <p>In the training sample the sensitivity is 71.4% and the specificity is 82.4%. In the
test sample the sensitivity is 71.5% and the specificity is 83.1%. This shows good
accuracy of this model.</p>
      <p>In addition, a ROC curve was constructed, the area under the curve of which was
0.853 (Fig. 12).</p>
      <p>Using the table with the coordinates of the curve points the threshold value for the
final discriminant equation 0.4511434 was selected. At this threshold the sensitivity is
76.4% and the specificity is 76.5%.</p>
      <p>The threshold value was selected from the points of the coordinate ROC. The
sensitivity and specificity values were selected so that the sum of sensitivity and
specificity was the maximum.
Three methods were implemented: hierarchical cluster analysis, k-means analysis, and
discriminant analysis.</p>
      <p>Analysis of the hierarchical cluster method showed that this method is not suitable
for analyzing large datasets. Even with the usage of reduction in the number of
observations in the sample it was not possible to obtain acceptable results.</p>
      <p>Analysis of the k-means method showed that this method can be used for
classification problems into two clusters "sick" and "healthy", but the accuracy of this
method did not show high results. The parameters that were identified by the method
as the most significant have not significant differences in the spread between
"healthy" and "sick". For more accurate operation of this method it is necessary to
filter out the least significant parameters and continue a more detailed study.</p>
      <p>The discriminant method analysis allowed us to obtain a discriminant equation
with sensitivity and specificity values of 76.4% and 76.5% respectively. Based on the
selected sensitivity and specificity values, a threshold was selected for working with
the discriminant equation.</p>
      <p>In order to implement the best method in terms of sensitivity and specificity in
medical DSS it should be tested on a larger sample size. Also for better accuracy in
predicting the probability of a person having second-stage tuberculosis,
crossvalidation should be performed.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In this paper several classification methods for working in the medical DSS were
investigated. The idea of creating a medical DSS is as follows: according to the ECG
parameters the trained methods determine the degree of probability of the presence of
tuberculosis of the second type in the examined person, whose open symptoms are
practically not observed. Thus, the system will help in the early stages of the disease
to determine the presence of it on the ECG.</p>
      <p>Medical data were used as experimental data in order to train DSS methods.
Medical data includes parameters calculated from an electrocardiogram, as well as
general impersonal parameters about its owner (height, weight, age, etc.). The
experimental sample consisted of more than five thousand electrocardiograms. Also a
hypothesis was put forward and tested about the possibility of determining the
presence or absence of signs of tuberculosis in a person by the parameters of an
electrocardiogram. This hypothesis was confirmed.</p>
      <p>Three methods were implemented: hierarchical cluster analysis, k-means analysis,
and discriminant analysis. The program for statistical data processing IBM SPSS
Statistics was used to carry out the work.</p>
      <p>Of the methods considered in this paper the most suitable for the classification
problem with a nominal target variable on the example of the study of medical
experimental data was the method of discriminant analysis. This method is similar to
regression analysis, which will be studied further, as well as other classification
methods that are not included in this work. In the future the model should be refined
to obtain a higher accuracy of the medical DSS.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Sim</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , et al.:
          <article-title>Clinical decision support systems for the practice of evidence-based medicine</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          ,
          <volume>8</volume>
          .6,
          <fpage>527</fpage>
          -
          <lpage>534</lpage>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Golub</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , et al.:
          <article-title>Delayed tuberculosis diagnosis and tuberculosis transmission</article-title>
          .
          <source>The international journal of tuberculosis and lung disease, 10.1</source>
          ,
          <fpage>24</fpage>
          -
          <lpage>30</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Nemirko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manilo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalinichenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Intellectual analysis of biomedical signals</article-title>
          .
          <source>Biotekhnosfera Journal</source>
          <volume>2</volume>
          (
          <issue>20</issue>
          ),
          <fpage>30</fpage>
          -
          <lpage>37</lpage>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Parikh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , et al.:
          <article-title>Understanding and using sensitivity, specificity and predictive values</article-title>
          .
          <source>Indian journal of ophthalmology, 56.1</source>
          ,
          <fpage>45</fpage>
          -
          <lpage>50</lpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fawcett</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>An introduction to ROC analysis</article-title>
          .
          <source>Pattern recognition letters, 27.8</source>
          ,
          <fpage>861</fpage>
          -
          <lpage>874</lpage>
          (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>B. I. N.</given-names>
          </string-name>
          , et al.:
          <article-title>Coronary artery motion during the cardiac cycle and optimal ECG triggering for coronary artery imaging</article-title>
          .
          <source>Investigative radiology, 36.5</source>
          ,
          <fpage>250</fpage>
          -
          <lpage>256</lpage>
          (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ott</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Lyman</surname>
            , Longnecker,
            <given-names>M. T.</given-names>
          </string-name>
          :
          <article-title>An introduction to statistical methods and data analysis. 7th edn</article-title>
          ., Brooks/Cole, Boston, United States (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kaufman</surname>
          </string-name>
          , L.,
          <string-name>
            <surname>Rousseeuw</surname>
            <given-names>P. J.:</given-names>
          </string-name>
          <article-title>Finding groups in data: an introduction to cluster analysis</article-title>
          . John Wiley &amp; Sons, New York, United States (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kanungo</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , et al.:
          <article-title>An efficient k-means clustering algorithm: Analysis and implementation</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          ,
          <volume>24</volume>
          .7,
          <fpage>881</fpage>
          -
          <lpage>892</lpage>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>McLachlan</surname>
            ,
            <given-names>G. J.</given-names>
          </string-name>
          :
          <article-title>Discriminant analysis and statistical pattern recognition</article-title>
          . John Wiley &amp; Sons, New York, United States (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Field</surname>
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Discovering statistics using IBM SPSS statistics</article-title>
          . 4th edn., SAGE Publications Ltd.,
          <string-name>
            <surname>London</surname>
          </string-name>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. Breiman L.:
          <article-title>Bagging predictors</article-title>
          .
          <source>Machine Learning</source>
          ,
          <volume>24</volume>
          ,
          <fpage>123</fpage>
          -
          <lpage>140</lpage>
          (
          <year>1996</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Fletcher</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ades</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kligfаild</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Exercise standards for testing and training: a scientific statement from the American Heart Association</article-title>
          .
          <source>Circulation Journal</source>
          ,
          <volume>128</volume>
          ,
          <fpage>873</fpage>
          -
          <lpage>934</lpage>
          (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>