<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Personal data segmentation based on conjugation index usage</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>P V Hripunov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D A Zherdev</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Pension Fund of the Russian Federation</institution>
          ,
          <addr-line>Shabolovka str. 4, Moscow, Russia, 119991</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Samara National Research University</institution>
          ,
          <addr-line>Moskovskoe Shosse 34, Samara, Russia, 443086</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>440</fpage>
      <lpage>442</lpage>
      <abstract>
        <p>The paper proposes a method for processing personal data that allows them to be divided into many segments or classes. The customer database is used as the source data. We use the indicator of conjugacy that has already proved the effectiveness in both recognition and clustering of data problems.</p>
      </abstract>
      <kwd-group>
        <kwd>Xk  x1 k </kwd>
        <kwd>x2 k </kwd>
        <kwd />
        <kwd />
        <kwd />
        <kwd />
        <kwd>x j k </kwd>
        <kwd />
        <kwd />
        <kwd />
        <kwd />
        <kwd>xM k </kwd>
        <kwd>k 1</kwd>
        <kwd>K</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Rk x j  
xT Q x
xjT xk j ,
j j</p>
      <p>k  1, K ,
Qk  Xk XTk Xk 
1 T</p>
      <p>X , k  1, K ,
k
vectors in the new dataset contain 300 features. All vectors in dataset must be the same size. For
example, if the first field size equals 12 symbols then the first 12 features of a vector contain the field
value and other 88 features are filled by zero value.When we processed the database, all personal data
was encrypted by summing up with some digital key.</p>
      <p>
        We use the similar procedure for clustering which was used in work [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for clustering of radar
images. At first step of the whole set we choose the two most "distanced" vectors. These vectors have
the minimal value of the correlation ratio and we can be labeled them as x1, xM .
      </p>
      <p>Then the algorithm from the remain set of vectors adds two new vectors ( x2 , xM 1 ). Each one of
these vectors must have the maximum of the correlation ratio:</p>
      <p>(x1T x2 )2</p>
      <p>R1,2 
RM 1,M </p>
      <p>,
x1 x2
(xTM 1xM )2</p>
      <p>,
xM 1 xM
with one vector was obtained at the first step. In the result received pairs of vectors x1, x2 and
xM 1, xM form the subspaces that were formed by matrices X1,2 and XM 1,M correspondingly. Then
using the remaining set next two vectors x3 , xM 2 that are closest to the subspaces are joined to
previously formed subspaces using computation ofconjugation index with a maximum value.</p>
      <p>Since the database contains a large number of vectors, process continues due finding of specific
number of vectors in both subspaces. For example, for such dataset the resulted subspace contains 15
vectors in both matrices Xk , Xl , which correspond to two subclasses.</p>
      <p>The procedure described above is repeated iteratively with all unlabeled vectors. Clustering is
continued until all the vectors will be specified to any of the subspaces.At the recognition stage with a
certain decision rule, the vector closest to one of the subclasses formed in thedescribed manner is
considered to belong to the class.
3. Results and discussion
In this paper, the problem of the determining possibility whether there is some given record in the
database is examined. After clustering process we can figure out the belonging of a vector to some
class. The subclass stores a small number of vectors in comparisonof the initial database.Therefore,
after the classification of the current vector, it will be easy to analyze data in a subclass and determine
isit possible to add a new value into the database.</p>
      <p>Thus, to verify the above assumption, we performed the experiment. From the database of 1041100
records there was performed the random selection of 1040 records five times. Each selection was
divided onto 80 subclasses, a subclass consists of 13 vectors. After the clustering procedure, the
generated vectors were classified.</p>
      <p>
        The testing vectors were formed using the existed in the dataset records with some modifications.
For example, there was simulated situation of incorrect handwritten letters conversion when the
personal data was filled in some document. As it was shown in the work [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] the problem of text
recognition is a difficult and can have many solutions. Figure 1 a, b presents the images of two
handwritten words. The word in Figure 1 a was correct converted by some letters recognition software
into “ADAM” as opposed to word in Figure 1 b that led to incorrect result: “ADRM”.
a)
      </p>
      <p>b)</p>
      <p>Figure 1. Examples of a) correct and b) incorrect handwritten letters conversion.</p>
      <p>
        In the classification experiment, 20 vectors of the type described above were tested. All vectors
were successfully classified based on the many-to-many approach [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] extending the possibilities of the
binary classification of the support subspaces algorithm [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Moreover the average value of
conjugation index was 0.95 for true defined class. This fact undoubtedly indicates the reliability of
using the conjugation index in problems of this kind. This is an advantage for following research of
such kind both with databases of a more complex type, with a larger field number, and for
classification using a whole database of one million or more records.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Yang</surname>
            <given-names>Y</given-names>
          </string-name>
          and
          <string-name>
            <surname>Guan J 2002 CLOPE:</surname>
          </string-name>
          <article-title>a fast and effective clustering algorithm for transactional data Proceedings of the eighth ACM SIGKDD international conference on Knowledge discovery and data mining 682-687</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Zhang</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ramakrishnan</surname>
            <given-names>R</given-names>
          </string-name>
          and
          <string-name>
            <surname>Livny M 1996 BIRCH</surname>
          </string-name>
          <article-title>: an efficient data clustering method for very large databases</article-title>
          <source>ACM Sigmod Record</source>
          <volume>25</volume>
          (
          <issue>2</issue>
          )
          <fpage>103</fpage>
          -
          <lpage>114</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>He</surname>
            <given-names>Z</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            <given-names>X</given-names>
          </string-name>
          and
          <string-name>
            <surname>Deng</surname>
            <given-names>S 2005</given-names>
          </string-name>
          <article-title>A cluster ensemble method for clustering categorical data</article-title>
          <source>Information Fusion</source>
          <volume>6</volume>
          (
          <issue>2</issue>
          )
          <fpage>143</fpage>
          -
          <lpage>151</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Huang Z 1997</surname>
          </string-name>
          <article-title>A fast clustering algorithm to cluster very large categorical data sets in data mining DMKD 3(8</article-title>
          )
          <fpage>34</fpage>
          -
          <lpage>39</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Fursov</surname>
            <given-names>V</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kozin</surname>
            <given-names>N 2007</given-names>
          </string-name>
          <article-title>Recognition through constructing the eigenface classifiers using conjugation indices</article-title>
          <source>IEEE Conference on Advanced Video and Signal Based Surveillance 465- 469</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Minaev</surname>
            <given-names>E</given-names>
          </string-name>
          and
          <string-name>
            <surname>Fursov</surname>
            <given-names>V 2016</given-names>
          </string-name>
          <article-title>Support subspaces method for fractal images recognition</article-title>
          <source>CEUR Workshop Proceedings</source>
          <volume>1638</volume>
          <fpage>379</fpage>
          -
          <lpage>385</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Zherdev</surname>
            <given-names>D A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kazanskiy</surname>
            <given-names>N L</given-names>
          </string-name>
          and
          <string-name>
            <surname>Fursov</surname>
            <given-names>V A</given-names>
          </string-name>
          <year>2015</year>
          <article-title>Object recognition in radar images using conjugation indices and support subspaces</article-title>
          <source>Computer Optics</source>
          <volume>39</volume>
          (
          <issue>2</issue>
          )
          <fpage>255</fpage>
          -
          <lpage>264</lpage>
          DOI: 10.18287/
          <fpage>0134</fpage>
          -2452-2015-39-2-
          <fpage>255</fpage>
          -264
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Bolotova</surname>
            <given-names>Y A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spitsyn</surname>
            <given-names>V G</given-names>
          </string-name>
          and
          <string-name>
            <surname>Osina P M 2017</surname>
          </string-name>
          <article-title>A review of algorithms for text detection in images and</article-title>
          videos
          <source>Computer Optics</source>
          <volume>41</volume>
          (
          <issue>3</issue>
          )
          <fpage>441</fpage>
          -
          <lpage>452</lpage>
          DOI: 10.18287/
          <fpage>2412</fpage>
          -6179-2017-41-3-
          <fpage>441</fpage>
          - 452
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Bishop Ch M 2006</surname>
          </string-name>
          <article-title>Pattern Recognition and Machine Learning</article-title>
          (New York: Springer) p
          <fpage>738</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>