<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>International Journal of Computer
Vision</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Recognition of Manuscript Tables in Computer Processing of Technical Transport Documentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Elena Y. Bursian</string-name>
          <email>bursianeu@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anton M. Demin</string-name>
          <email>ad2271@ya.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander P. Glukhov</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Emperor Alexander I, St. Petersburg State Transport, University</institution>
          ,
          <addr-line>St. Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Joint Stock Company Railway, Research Institute (JSC, "VNIIZhT")</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2000</year>
      </pub-date>
      <volume>33</volume>
      <issue>2</issue>
      <fpage>35</fpage>
      <lpage>43</lpage>
      <abstract>
        <p>The article discusses the process of recognizing handwritten characters presented in tables of technical railway documentation and test works of students of PGUPS. In the model under study, a skeleton graph is constructed for each recognizable region and the procedure for statistical processing of the characteristics of the branches of skeletal graphs is analyzed.Skeletal graphs are constructed for reference symbols and are dynamically replaced during recognition by skeletal graphs of recognized areas. The nature of dependencies between the components of skeletal graphs of reference symbols and dynamically added objects is investigated. In the process of automatic recognition of handwritten characters, the obtained statistical results are applied.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The task of automatic processing and optical
recognition of a pre-scanned or captured image is
relevant in many fields of activity. Verification of
scanned completed tables and forms, tests,
questionnaires, tests during distance learning of PSU
students, tables of technical railway documentation
requires handwriting recognition.</p>
      <p>Universal electronic document management systems
and individual specialized programs are based on the
principles of the general theory of recognition.
Monographs by V. N. Vapnik [Vap74],
A.Ya.Chervonenkis, Yu.I.Zhuravlyov[Zhu05], A.B.
Merkov[Mer14], V.V. Ryazanov, O.V. Senko.</p>
      <p>Researches of L.M. Mestetsky[Mes09], D. A
Gavrilov[Gav19],N. A.
Lomov[Lom16],Yu.V.Vizilter[Vis12], Ya.A. Furman
relate to practical application and have recognized
applied value. The main foreign works are presented by
R. C. Gonzalez,L.Lam [Lam95], C.Suen, T. Y. Zhang
[Zha84], C. Y. Suen, D. T.Lee [Lee82],H. Blum
[Blu67], R. O. Duda [Dud00], P. E. Hart, R. Shapiro,
L. Shapiro, G. Stockman, P. Viola and M. Jones
[Vio04],S. Rosset[Ros04], L.C.Molina [Mol02].
Currently, there are computer systems for recognizing
handwritten characters in almost noisy images:
ABBYY FormReader, OmnPage, CuneiForm,
ReadirisPro. Moreover, the use of these systems in the
recognition of specialized tables is not always
effective, since the recognized information in many
cases has a predetermined structure. Recognition of
handwritten characters in technical tables on
lowquality images is an urgent scientific and technical
task.</p>
      <p>Handwritten tables and test documents usually assume
that there are different density distributions for specific
characters, but they are unknown. During computer
processing, the characteristics of the sample can be
calculated and unknown.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Statement of the problem</title>
      <p>To build an automatic recognition system for
individual handwritten characters, a set of basic
recognizers is used with not always a high probability
of object recognition. Weak recognizers are grouped,
and a committee of classifiers is built using the
AdaBoost algorithm.</p>
      <p>When automatically checking test works, it is
permissible to assume that the classifier divides the
space of vectors of informative attributes X into two
sets X1 and X2, since the test usually assumes a single
answer. You can also automatically build a committee
of classifiers separately for each character.</p>
      <p>To build a committee of classifiers, you must first
construct many different basic recognizers. The
construction of basic recognizers by estimating the
parameters of multidimensional distribution densities
of the vector characteristics of handwritten characters
is an urgent task.</p>
    </sec>
    <sec id="sec-3">
      <title>2 Construction final classifier of basic recognizers and</title>
      <p>After scanning the document, the image is reduced to a
two-gradation view. Figure 1 shows a scanned image
of a table of railway documentation.</p>
      <p>The set of points P belonging to the skeleton
representation is also called the skeleton of the region
D. We can assume that the skeleton of the region is the
set of centers of maximal circles lying in the region
(Figure 2).</p>
      <p>Based on the region’s skeleton, for each recognized
It should be noted that a modern approach to
maintaining railway documentation requires the
mandatory introduction of electronic document
management [*].</p>
      <p>For each recognized area, a skeletal description is
constructed. The calculation of the characteristics of
the skeletal representation of the region is based on the
following definition.</p>
      <p>The point P belongs to the skeletal representation of the
domain D if and only if the following statement holds:
"() ⊂  &amp; ∄"+(+) ⊂ : "() ⊂
"+(+)&amp;" () ≠ "+(+),
where Br(P) is a circle centered at point P and radius r
[Dud00].
object, the characteristics of the loaded graph, called
the region’s skeleton graph, are calculated. The skeletal
graph of a recognized object can be represented as
follows (Figure 2, Figure 3).
For each branch of the skeletal graph, a vector of
informative characteristics is constructed, made up of
the slope coefficients of the edges of the skeleton graph
or directly the slope angles of the edges of the skeleton
graph. In the case when the symbols are written in one
hand, between the values of the slope angles of the
edges of the skeletal graph taken at equal intervals,
there is a statistical dependence.
The regression of the slope angles of the skeletal graph
of a symbol from the values of the corresponding
angles taken at previous intervals can be calculated
using multivariate regression analysis methods,
creating a system for determining unknown regression
coefficients.
We assume that == (+, … 8)is the vector of
regression factors, in the training set for the symbol
with number i the regression factors take the values:
+(?), … 8(?)and correspond to the angles of inclination
of the edges of the skeletal graph of the symbol,3(?)is
the response value. = = (3, +, … 8) – estimated
regression parameters,= = ((+), … (;))error vector,
The number of unknown regression parameters does
not exceed the number characters in the training set
k&lt;n.</p>
      <p>Considering that the angle of inclination of the edges of
the skeletal graph of a symbol taken at a certain
interval is a random variable, we can write its
conditional expectation in the following form.</p>
      <p>[|] = 3 + ++ … 88
An estimate of the regression parameters a can be
obtained by calculating the pseudoinverse matrix.</p>
      <p>= (ΦΦ=)G+Φ
where</p>
      <p>1
⎛ 1
Φ = ⎜</p>
      <p>…</p>
      <p>When a symbol is recognized, the vectors of the
informative characteristics of the symbol are compared,
in this case, the angles of inclination of the edges of the
skeletal graph, with the set of angles of inclination
obtained by regression of the angles of inclination
taken at previous levels.</p>
      <p>To classify the classifier, basic recognizers are also
used based on algorithms for calculating the Hausdorff
distances between sets of critical vertices of skeletal
graphs, calculating correlation functions for the angles
of inclination of edges of skeletal graphs, and
comparing them with threshold values of correlation
functions of coordinates of critical vertices of skeletal
graphs, where the vertices of skeletal graphs are
considered critical non-hanging peaks with degrees
other than two[Ros04], [Mol02].</p>
      <p>The final classifier is constructed using the AdaBoost
algorithm [Vio04]. The training sample is the set of
ordered pairs A = {(x1, y1), ... (xi, yi), ... (xn, yn)}, where
xi is the skeleton graph of the symbol, yiÎ{1, -1} if yi =
- 1, the skeleton graph corresponds to the symbol, yi =
1 otherwise, ht are the basic classifiers, tÎ{1,2 ... T}.
The initial distribution of elements of the set A is
initialized by the uniform distribution P1 = 1 / n. In the
next steps, for tÎ{1,2 ... T}, the coefficients αt and the
new distribution Pt+1(i) are calculated using the
distribution Pt(i) calculated in the previous step.
P =</p>
      <p>min
T∈{+,: ,…W}</p>
      <p>({ℎP() ≠ ?})
That is, εt is the probability of error of the classifier ht,
provided that the classifier ht is less error than other
classifiers on the distribution Pt.</p>
      <p>P =
1
2
ln (
1 − P
P</p>
      <p>)
P`+() =</p>
      <p>ec
bc(?)dGfcgc(hi ),
The probability distribution for the next step is updated
by the formula:
wherezt is the normalizing factor.</p>
      <p>The final classifier is built according to the formula:
(?) = (n</p>
      <p>P ℎP(?)).
=
Po+</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>When applying the AdaBoost algorithm for
recognizing individual handwritten characters in tables
of railway documentation or in test works of students
of PSUPS, a training set of skeletal graphs and a set of
basic classifiers with a volume of at least two dozen
elements are required.</p>
      <p>It is possible to develop basic classifiers and build the
final classifier, both for an arbitrary character, and for
individual characters with the division into two classes:
belonging to the class of the given character and
relation to other sets. The development of basic
classifiers based on various algorithms and heuristic
procedures is a laborious task.</p>
      <p>The use of multidimensional linear and polynomial
regression methods for the angles of inclination of the
edges of the skeleton graph of the symbol makes it
possible to construct a set of basic recognizers, since
when writing a symbol in case of deviation or
curvature of the sign, a regression trend is confirmed,
confirmed by experimental data.</p>
      <p>To obtain experimental data, an experimental set of
programs based on the Visual C ++ platform was
developed. The construction of the skeletal
representation of the symbol, the skeletal graph of the
symbol, and the angles of inclination of the edges of
the skeletal graph of the symbol was carried out.
The considered approach allows us to solve more
complex problems of text recognition. In particular,
recognition of handwritten railway documentation or
recognition of handwritten text allows for a large
number of mathematical symbols and its presentation
in the generally accepted TEX format.
[Vap74] Theory of pattern recognition (statistical
problemsof learning), V. N. Vapnik, A. Ya.</p>
      <p>Chervonenkis. Publishing House "Science",
Main Edition of the Physics and</p>
      <p>Mathematics Literature, M., 416 pp., 1974.
[Zhu05] Yu.I. Zhuravlev, V.V. Ryazanov, O.V.</p>
      <p>Senko. RECOGNITION Mathematical
methods.Software system.Practical
applications. - M .: FIZMATLIT -- 159 p.,
2005 .
[Zhu78] Yu. I. Zhuravlev.On the algebraic approach to
recognition and classification problems//
[Gav19] D. A.Gavrilov ,L. M.Mestetskiy,</p>
      <p>Semenov A. B. A method for aircraft
labeling in aerial and satellite images based
on continuous morphological
models /Programming and Computer
Software. —. — Vol. 45, no. 6. — P. 303–
310, 2019.
[Lom16] N. A. Lomov, L. M. Mestetskiy. Area of the
disk cover as an image shape
descriptor. Computer Optics, 40(4):516–525,
2016.
[Vis12]VisilterYu.V., Sidyakin S.V. Construction of
morphological spectrahalftone image //
Bulletin of computer and
informationTechnologies, N4, pp. 8-17,
2012.
[Lee82] D. T.Lee Medial axes transform of planar
shape // IEEE Transactions on Pattern
Analysis and Machine Intelligence. PAMI-4,
363-369p., 1982.
[Blu67] H. Blum. A transformation for extracting new
descriptors of shape // Models for the
Perception of Speech and Visual Form, MIT</p>
      <p>Press, 362-380p, 1967.
[Bul12] P.E. Bulavsky, M.N. Vasilenko, A.A.</p>
      <p>Kornienko, A.D. Khomonenko. Automation
of information support for railway managers
on the basis of the introduction of electronic
document management.News of St. Petersburg
University of Railways Messages, no. 2
(31),p. 116-118, 2012.
[Mol02] L.C.Molina, L.Belanche, A.Nebot. Feature
Selection Algorithms: A Survey And
Experimental Evaluation. Proceedings of the
2002 IEEE International Conference on Data
Mining, IEEE Computer Society, 306–313 p. ,
2002.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>