<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Static gestures classi cation using Convolutional Neural Networks on the example of the Russian Sign Language</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oleg Potkin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrey Philippovich</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Moscow Polytechnic University</institution>
          ,
          <addr-line>Moscow, 107023, Russian Federation, WWW home page:</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article describes the dataset developed for the classi er training, the algorithm for data preprocessing, described an architecture of a convolutional neural network for the classi cation of static gestures of the Russian Sign Language (the sign language of the deaf community in Russia) and represented an experimental data.</p>
      </abstract>
      <kwd-group>
        <kwd>deep neural networks</kwd>
        <kwd>convolutional neural networks</kwd>
        <kwd>classi cation</kwd>
        <kwd>gestures</kwd>
        <kwd>Russian Sign Language</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Human-computer interaction interfaces are diverse in their scope and
implementation: systems with console input-output, controllers with gesture control,
brain-computer interfaces [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] etc. Systems with the data input based on the
custom gestures recognition have gained wide popularity since 2010 after the
release of the contactless game controller "Kinect" from Microsoft. These
devices increase their market share to this day and become a part of everyday life
of di erent user categories. So, for example, the Volkswagen car manufacturer
introduced the multimedia system Golf R Touch Gesture Control for the car
multimedia system management using gestures [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], researchers from Cybernet
Systems Corporation have developed personal computer software that allows to
use hands gestures instead of the usual input devices [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Russian scientists
have demonstrated a technology of protection against spam bots based on the
CAPTCHA mechanism [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        To convert gestural commands into the control signal, a gesture classi cation
mechanism is needed, which in turn can be obtained from various devices: special
gloves de ning the coordinates of the joints [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], as well as 2D and 3D video
cameras. The approach using gloves has a signi cant drawback the user needs
to wear a special device connected to the computer. However, the approach
based on the concept of computer vision using video cameras is considered more
natural and less expensive.
      </p>
      <p>The present study demonstrates the system of the static gestures classi
cation for the Russian Sign Language based on the computer vision approach using
convolutional neural networks. The topic is relevant and represents a starting
point for researchers in the eld of gesture recognition.</p>
    </sec>
    <sec id="sec-2">
      <title>Dataset description for classi er training</title>
      <p>For training, cross-validation and testing of the classi er, a set of data was
developed1, containing 1042 images of hands in a certain gesture con guration
the class. In total, 10 classes are represented in the data set, each of which
corresponds to a certain static gesture of the Russian Sign Language.</p>
      <p>Fig. 1 shows sample images from each of the data set classes.</p>
      <p>Images were made with di erent light conditions and on a small range of
object distances from the camera (0.4 0.7 meters).</p>
      <p>Table 1 represents the image parameters from the dataset.
1 Dataset and the source code of the project are available on GitHub:
https://github.com/olpotkin/DNN-Gesture-Classi er</p>
      <p>
        Before sending the data to the classi er, it is necessary to perform image
transformations that will help to get rid of minor characteristics for the classi er,
which in turn will increase the performance of the classi er. This process is called
preprocessing. Image preprocessing reduces the number of unimportant details
(e.g. color information from RGB space). Practices and an e ciency of image
preprocessing techniques is demonstrated in the publications [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        In the developed system, the RGB color space was transformed into YCbCr
for skin segmentation [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and binarization was applied to the threshold values
(Fig. 2).
      </p>
      <p>The nal stage of processing is normalization. The pixel values are in the
range from 1 to 255, but for correct operation of the neural network, it is
necessary to input values from 0 to 1. To do this, performed a conversion using the
formula 1:</p>
      <p>N =</p>
      <p>P
Pmax
(1)
where N the pixel value after normalization, P the value of the normalized
pixel, Pmax the maximum value of the range.</p>
      <p>
        In order to minimize the over tting possibility of the classi er [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the
dataset is divided into 3 parts:
{ Training set contains 80% of the dataset (666 images).
{ Test set, contains 20% of the data set (209 images). Using for evaluation of
the model. These samples never used during training and cross-validation
processes.
      </p>
      <p>{ Cross-validation set, contains 20% of the training sample (167 images).</p>
      <p>
        Neural network architecture and the classi cation
results
In this section described CNN architecture on the abstract level. As a
benchmark model for an experiment LeNet-5 CNN architecture was taken. The main
disadvantage of LeNet-5 is over tting in some cases and no built-in mechanism
to avoid this [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. So, benchmark architecture was improved by adding dropout
layers [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Improved LeNet-5 architecture is shown on Fig. 3.
{ Input layer (Input). The size of the layer corresponds to the image size after
the preprocessing: 1 128 128.
{ Convolution layer (Conv1). Output size: 8 62 62. Convolution window
size: 5 5. Activation function ReLU (Recti ed Linear Unit). ReLU can
be described by the equation f (x) = max(0; x) and implements a simple
threshold transition at zero. Compared to the sigmoid activation function,
ReLU increases learning speed and classi er performance [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
{ Convolution layer (Conv2). Output size: 24 29 29. Convolution window
size: 5 5. Activation function ReLU.
{ Convolution layer (Conv3). Output size: 36 13 13. Convolution window
size: 5 5. Activation function ReLU.
{ Convolution layer (Conv4). Output size: 48 5 5. Convolution window size:
5 5. Activation function ReLU.
{ Convolution layer (Conv5). Output size: 64 3 3. Convolution window size:
3 3. Activation function ReLU.
{ Dropout layer 1, p = 0:25. The layer is constructed in such a way that each
neuron can fall out of this layer with probability p, therefore, other neurons
remain in the layer with probability q = 1p. The dropped neurons are not
included in the classi er training, that is, at each new epoch, the neural
network is partially changed. This approach e ectively solves the over tting
problem of the classi er [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
{ Fully connected layer 1. Size: 576.
{ Fully connected layer 2. Size: 1024.
{ Dropout layer 2, p = 0:25.
{ Fully connected layer 3. Size: 256.
{ Fully connected layer 4. Size: 128.
{ Output layer. Size: 10. Activation function softmax.
      </p>
      <p>The neural network is developed using the Keras framework based on
TensorFlow with the following hyper-parameter values after the optimization
procedures: the learning rate is 0.0001, the batch size is 128, the number of epochs
is 24, the metrics accuracy, the optimizer type Adam.</p>
      <p>Training and test procedures were performed on the benchmark classi er
(LeNet-5) with results shown in Table 2.</p>
      <p>Benchmark model is over tting because during the training process it shows
a way better result compared with test process.</p>
      <p>Results of improved architecture with optimization of hyper-parameters shown
in Table 3.
As a result of the research, the dataset consisting of more than a thousand
elements belonging to ten di erent classes of the Russian Sign Language was
developed and published. Algorithms and procedures for preliminary data
processing in Python 3.5 were developed. The classi er based on the convolutional
neural network using the Keras and TensorFlow frameworks was designed and
developed.</p>
      <p>The classi er demonstrates an accuracy of classi cation in 91.38% on the
test dataset that is better than result of benchmark classi er 87.08%. The
represented results of an accuracy are su cient to be a strong basis for further
research in this direction.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Potkin</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ivanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Neurointerface Neurosky: interaction with the hardware computing platform Arduino. 64th Open Student Scienti c</article-title>
          and Technical Conference, Moscow: Moscow State University of Mechanical Engineering, pp.
          <fpage>151</fpage>
          -
          <lpage>153</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2. Volkswagen: Gesture Control:
          <article-title>How to make a lot happen with a small gesture</article-title>
          . [http://www.volkswagen.co.uk/technology/comfort-and-convenience/gesturecontrol]
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Charles</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          , Glenn Beach,
          <article-title>Gene Foulk: A Basic Hand Gesture Control System for PC Applications</article-title>
          .
          <article-title>Cybemet Systems Corporation (</article-title>
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Shumilov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Philippovich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Gesture-based animated CAPTCHA</article-title>
          .
          <source>Information and Computer Security</source>
          , Vol.
          <volume>24</volume>
          Iss 3 pp.
          <fpage>242</fpage>
          -
          <lpage>254</lpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dipietro</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Angelo</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Sabatini</surname>
            , Dario,
            <given-names>P.</given-names>
          </string-name>
          <article-title>: urvey of Glove-Based Systems and their applications</article-title>
          .
          <source>IEEE Transactions on systems, Man and Cybernetics</source>
          , Vol.
          <volume>38</volume>
          , No.
          <issue>4</issue>
          , pp
          <fpage>461</fpage>
          -
          <lpage>482</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Vicen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Gonzalez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Tra c Sign Classi cation by Image Preprocessing and Neural Networks</article-title>
          .
          <source>Conference: Proceedings of the 9th international work conference on Arti cial neural networks</source>
          (
          <year>1970</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>A.A.</given-names>
            <surname>Vedenov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.V.</given-names>
            <surname>Zurin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.B.</given-names>
            <surname>Levchenko</surname>
          </string-name>
          :
          <article-title>Optical image preprocessing for neural network classi er system</article-title>
          .
          <source>In book: Parallel Problem Solving from Nature</source>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kussul</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baidyk</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kussul</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Neural network system for face recognition</article-title>
          .
          <source>Conference: Circuits and Systems</source>
          ,
          <year>2004</year>
          .
          <source>ISCAS '04. Proceedings of the 2004 International Symposium onVolume: 5</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Son</given-names>
            <surname>Lam</surname>
          </string-name>
          <string-name>
            <surname>Phung</surname>
          </string-name>
          , Bouzerdoum,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Chai</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.:</surname>
          </string-name>
          <article-title>A novel skin color model in YCbCr color space and its application to human face detection</article-title>
          .
          <source>Visual Information Processing Research</source>
          Group Edith Cowan University, Westem Australia, IEEE KIP, pp
          <fpage>289</fpage>
          -
          <lpage>292</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Brownlee</surname>
          </string-name>
          , J.:
          <article-title>What is the Di erence Between Test and Validation Datasets? Machine Learning Process</article-title>
          , [https://machinelearningmastery.com/di erence-testvalidation-datasets] (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Ahmed</surname>
          </string-name>
          El-Sawy,
          <article-title>Mohamed Loey: CNN for Handwritten Arabic Digits Recognition Based on LeNet-5</article-title>
          . In book:
          <source>Proceedings of the International Conference on Advanced Intelligent Systems and Informatics</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.E.:
          <article-title>Imagenet classi cation with deep convolutional neural networks</article-title>
          .
          <source>NIPS</source>
          pp.
          <volume>19</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Geo</surname>
            rey
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hinton</surname>
            , Nitish Srivastava,
            <given-names>Krizhevsky A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruslan</surname>
            <given-names>R</given-names>
          </string-name>
          . Salakhutdinov:
          <article-title>Improving neural networks by preventing co-adaptation of feature detectors</article-title>
          . Department of Computer Science, University of Toronto, pp.
          <fpage>1</fpage>
          -
          <lpage>18</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>