<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>and A. Chaudhary. Robust ges-
ture recognition using kinect: A compari-
son between DTW and HMM. Optik - In-
ternational Journal for Light and Electron
Optics</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Real-Time Vision Based System for Recognition of Static Dactyls of Albanian Alphabet</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Eriglen Gani</string-name>
          <email>eriglen.gani@fshn.edu.al</email>
          <email>eriglen.gani@fshn.edu.al Bruno Goxhi Department of Informatics Faculty of Natural Sciences University of Tirana bruno.goxhi@fshnstudent.info</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alda Kika</string-name>
          <email>alda.kika@fshn.edu.al</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics, Faculty of Natural Sciences, University of Tirana</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>11</fpage>
      <lpage>12</lpage>
      <abstract>
        <p>The aim of the paper is to present a realtime vision based system that is able to recognize static dactyls of Albanian alphabet. We use Kinect device, as an image receiving technology. It has simpli ed the process of vision based object recognition, especially for segmentation phase. Di erent from hardware based methods, our approach does not require that signers wear extra objects like data gloves. Two pre-processing techniques, including border extraction and image normalization have been applied in the segmented images. Fourier transform is applied in the resultant images which generates 15 Fourier coe cients representing uniquely that gesture. Classi cation is based on a similarity distance measure like Euclidian distance. Gesture with the lowest distance is considered as a match. Our system achieved an accuracy of 72.32% and is able to process 68 frames per second.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Sign language is used as a natural way of
communication between hearing impaired people. It is very
important for the inclusion of deaf people in society.
There exist a gap in communication between hearing
impaired people and hearing ones. It comes from
inability of hearing people to understand sign language.
To overcome this gap most of the times interpreters
can be used. The other, more comfortable solution is
usage of technology. Natural interfaces can be used to
capture the signs and understand their meaning by
using body positions, hand trajectories and head
movements. Using technology for catching, processing and
translating dactyls in an understandable form for non
deaf people, would help deaf ones integrate faster in
the society [GK16]. A real time dactyls translator
system would provide many facilities for this community.
Many countries have tried to develop real-time sign
language translator like: [Ull11, GK15c, TL11].
Unfortunately Albanian sign language (AlbSL) did not
get much focus as other languages.</p>
      <p>Deaf people in Albania used to communicate on the
way that is based on nger-spelled Albanian words
[ANA13]. Although not so e cient, dactyls play an
important role in this type of communication. They
form the bases of communication for deaf people.
Albanian dactyl alphabet is composed of 36 dactyls.
Among them 32 are static dactyls. 4 of them are
dynamics ones, which are obtained from consecutive
sequences of frames. The dynamic dactyls include (C,
E , SH and ZH) [ANA13]. Our work is focused only in
32 static dactyls.</p>
      <p>Two most widely used methods for building
realtime translator system are hardware based and vision
based [ZY14]. In hardware based method the
signers have to wear data gloves or some other marker
devices. It is not very natural to them. Vision
based methods are more challenging to be developed
but are more natural for deaf people. Two most
common problems include a)complex background and
b)illumination change [ZY14]. Sometimes it is hard
to distinguish human hands from other objects parts
of the same environment. Sometimes the shadow or
light e ects the correct identi cation of human hand.
Kinect sensor by Microsoft, has simpli ed the process
of vision based object recognition, especially the
segmentation phase. It o ers some advantages like:
provides color and depth data simultaneously, it is
inexpensive, the body skeleton can be obtained easily and
it is not e ected by the light [GK16]. We are using
Kinect sensor as a real-time image receiving
technology for our work.</p>
      <p>Our Albanian sign language translator system
includes a limited set of number signs and dactyls. In
the future other numbers, dynamic dactyls and signs
will be integrated by making this system usable in
many scenarios that require participation of deaf
people. One usage of the system includes a program in
a bar that could help the deaf people making some
orders by combining numbers and dactyl gestures.</p>
      <p>Till now there does not exist any gesture data set
for Albanian sign language. We are trying to built a
system that is able to translate static dactyls signs for
Albanian sign language and in the future it will be
extended to dynamic dactyls and other signs.
Creating and continuously adding new signs to an Albanian
gesture data set would help building a more reliable
and useful recognition system for our sign language.</p>
      <p>Section 1 gives a brief introduction. Section 2
summarizes some related works. The rest of the paper is
organized as follows. Section 3 presents an overview of
methodology and a brief description of each
methodology's processes. Section 4 describes the experimental
environment. Section 5 presents the experiments and
results. The paper is concluded in Section 6 by
presenting the conclusions and future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>Many researchers have followed di erent
methodologies for building sign language recognition systems.
They are categorized into several types based on
input data and hardware dependency. Signs, which are
mostly performed by human hands can be static or
dynamic. The sign language recognition systems are
categorized as hardware based or vision based.</p>
      <p>Many works have been done to integrate some
hardware based technologies to capture and translate sign
gestures, among them the most widely used are data
gloves. Authors at [Sud14] built a portable system for
deaf people using a smart glove capable of capturing
nger movements and hand movements. In general
data gloves achieve high performance but are
expensive and not a proper way to human-computer
interaction perspective [GK15b].</p>
      <p>Web cameras with an image processing system can
be used in vision based approaches. Research at
[SSKK16] presents a vision based methodology using
web cameras to recognize gesture from Indian sign
language. The system achieves high recognition rate.
Authors at [WKSE02] and [LGS08] use color camera to
capture input gestures and then SVM (Support Vector
Machine) and Fuzzy C-Means respectively to classify
hand gestures. Despite this, in general web cameras
generate low quality of images and have an inability to
capture other body parts. It is also hard to generalize
the algorithms for web cameras due to many di erent
shapes and colors of hands [GK15b].</p>
      <p>Kinect sensors by Microsoft has simpli ed the
process of vision based object recognition. It has many
advantages as: provide color and depth data
simultaneously, it is inexpensive, the body skeleton can be
obtained easily and it is not e ected by the light. Various
researchers are using Microsoft Kinect sensor for sign
language recognition as in [GK15c], [SB13], [VAC13].</p>
      <p>Vision based hand gesture recognition provides
more intuitive interaction with a system. It is a
challange task to identify and classify hand gesture.
Shape and movement play an important role in
gesture categorization. A comparison between two most
widely used algorithm for shape recognition is done at
[CBM07]. It compares Fourier descriptors (FD) and
HU moments in terms of performance and accuracy.
Algorithms are compared against a custom and a
reallife gesture vocabulary. Experiment results show that
FD is more e cient in terms of accuracy and
performance.</p>
      <p>Research at [BGRS11] addresses the issue of
feature extraction for gesture recognition. It compares
Moment In-variants and Fourier descriptors in terms
of in-variance to certain transformations and
discrimination power. ASL images were used to form gesture
dictionary. Both approaches found di cult to classify
correctly some classes of ASL.</p>
      <p>Authors at [BF12] compare di erent methods for
shape representation in terms of accuracy and
realtime performance. Methods that were used to
compare them include region based moments (Hu moments
and Zenike moments) and Fourier descriptors.
Conclusions showed that Fourier descriptors have the highest
recognition rate.</p>
      <p>Shape is an important factor for gesture recognition.
There exist many methods for shape representation
and retrival. Among them Fourier descriptors achieve
good representation and normalization. Authors at
[ZL+02] compare di erent shape signatures used to
derived Fourier descriptors. Among them: complex
coordinates, centroid distance, curvature signature and
cumulative angular function. Article concludes that
centroid distance is signi cantly better than other three
signatures.</p>
      <p>Sign language is not limited only in static
gesture. Majority of signs are dynamic ones. Research
at [RMP+15] proposed a hand gesture recognition
method using Microsoft Kinect. It uses two di
erent classi cation algorithms DTW and HMM by
discussing the pros and cons of each technique.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>Microsoft Kinect is used as a real-time image
retrieval. Kinect consists of an RGB camera, an IR
emitter, an IR depth sensor, a microphone array and a tilt.
The RGB camera can capture three-channel data in a
1280 x 960 resolution at 12 FPS or a 640 x 480
resolution at 30 FPS. In our work images consist of a 640 x
480 resolution at 30 FPS. The valid operating distance
of Kinect is approximately 0.8m to 4m [MSD16]. Due
to its advantages, it has simpli ed the process of vision
based object recognition, especially for segmentation
phase.</p>
      <p>Every pixel generated from Kinect device contains
information of their depth location layer and player
index. By using player index we focus only in pixels
that are part of human body [WA12]. In this way all
other pixels, not part of human body are excluded.
By applying a constant threshold we can obtain the
human hand, since it is the rst part of human body
towards the Kinect device [GK15a].</p>
      <p>In order to perform Fourier transform we have to
generate a centroid function which is based in hand
image contour. Theo Pavlidis is used as a hand contour
tracking algorithm [Pav12]. The segmented hand is
transformed in greyscale where each pixel is classi ed
as a white or a black one. After applying Theo Pavlidis
algorithm, the resultant image contains only border
pixels of human hand.</p>
      <p>Fourier descriptors can be derived from complex
coordinates, centroid distance, curvature signature or
cumulative angular function. In our case centroid
distance is used due to [ZL+02]. After locating the
center of white pixels in the image, we have calculated
the distance of every border pixels from it. It gives
the centroid function which represents two dimensions
area.</p>
      <p>The normalization process consists of extracting the
same number of pixel, equally distributed, among hand
border. Choosing a lower number of border pixels
decrease the system accuracy, while choosing a higher
number decrease the system performance. In our case
a number of 128 pixels has been chosen.</p>
      <p>Fourier descriptors are used to transform the
resultant image into a frequency domain. For each image,
only the rst 15 Fourier coe cients are used to de ne
them uniquely. Other Fourier coe cients do not
effect system accuracy. Every input gesture is compared
against a training data set using a similarity distance
measure like Euclidian distance. The gesture with the
lowest distance is considered as a match.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Experiment Environment</title>
      <p>Experiment environment used for implementing and
testing our real-time static dactyls recognition system
is composed of the following hardware: A notebook
with a processing capacity of 2.5 GHz, Intel Core-i5.
A memory capacity of 6 GB of RAM and a Windows 10
operating system with a 64-bit architecture. Microsoft
Kinect for Xbox 360 is used as a real-time image
retrieval technology. It generates 30 frames per second
and can be used as a RGB camera and also can provide
depth data.</p>
      <p>System was developed using .Net technology.
Kinect for Windows SDK 1.8.0.0 was used as library
between Kinect device and our application. It provides
a way to process Kinect signals. An overview of the
system architecture is given at Figure 2.</p>
      <p>Static Dactyls Recognition Application</p>
      <p>Kinect for Windows</p>
      <p>SDK
RGB Camera</p>
      <p>Depth Sensor</p>
    </sec>
    <sec id="sec-5">
      <title>Experiment and Results</title>
      <p>To test the proposed system, several experiments were
conducted. Each experiment is based on two aspects:
accuracy and computation latency. The rst
experiment measured the accuracy of correct identi cation
and classi cation of static dactyls. Our system is not
able to identify and classify dynamic dactyls. It is
based only in static ones. Firstly training data set is
created. It contains 320 dactyl gestures taken from two
di erent signers. Each gesture is performed 5 times
from each signers and is represented by 15 Fourier
coe cients. There are in total 4800 coe cients (15x320).
For real-time testing, 4 di erent signers were used.
Each signer performed 5 gestures for each dactyl sign.
In total they performed 640 experiments. Each
element in testing data set is compared against all
elements in training data set. The element with lowest
Euclidian distance is considered as a match. The
average recognition accuracy for each static dactyl is given
in the Table 1 and Table 2.</p>
      <p>For all static dactyls, the system achieves an average
accuracy rate of 72.32%. Results show that dactyls
with the highest accuracy rate are 'A', 'J', 'K' and
'V'. Their accuracy rate is above 95%. Dactyls with
the lowest accuracy rate are 'L', 'Y' and 'Z'. Their
accuracy rate is below 52%.</p>
      <p>Table 3 and Table 4 give information regarding
static dactyls confusion percentages. Some of
Albanian dactyls are easily confused with other dactyls
due to their similarity. Based on experimental results
dactyls "D", "E", "F", "N", "O" are more confused
ones.</p>
      <p>The second experiment deals with system
performance. We want to achieve a performance that allows
the system to be deployed in real-time. For every sign
we analyzed the time required for the following phases:
hand segmentation, hand contour tracing,
normalization, centroid function generation, Fourier
transformation and gesture classi cation. Table 5 summarize the
results.</p>
      <p>The system needs approximately 12 to 17 ms to
process a static dactyl. Most of the overall time is
consumed by hand segmentation and gesture classi
cation processes. They occupy approximately 82% of
total time. It can be deployed without any latency in
a real-time system that uses Microsoft Kinect.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>The aim of this paper is to built a real-time system
that is able to recognize static dactyls for Albanian
alphabet by using Microsoft Kinect. Albanian alphabet
is composed of 36 dactyls and 32 of them are static.
The static dactyls are used as inputs for our system.
Kinect device provides a vision based approach and is
used as an image retrieval technology. Its main
feature includes depth sensor. For every static dactyl, a
data set with 15 Fourier coe cients was built. In
total data set consists of 4800 coe cients. For testing
purpose, 4 di erent signers were used. Each of them
performed 5 times each of the static dactyls. A total
of 640 experiments were conducted. For classi cation
purpose a similarity distance measures like Euclidian
distance was used. Every element in testing data set
is compared against each element in training data set.
The element with the lowest Euclidian distance is
considered as a match. The system is tested against
accuracy and performance. Based on experiments results
the system achieves an accuracy rate of 72.32%. The
system needs to compute a static dactyl is 14.05 ms
in average. It can be deployed in a image receiving
technology that generates 68 frames per second.</p>
      <p>Future work consists of improving the overall
system performance and accuracy by applying a more
reliable data set. This can be done by including more
diverse signers who have high knowledge of Albanian
sign language. The future work also consist of adding
dynamic dactyls as well as other gestures of Albanian
sign language.
[ANA13]
[BF12]</p>
      <sec id="sec-6-1">
        <title>ANAD. Gjuha e Shenjave Shqipe 1.</title>
        <p>Shoqata Kombetare Shiptare e Njerezve qe
nuk Degjojne, 2013.</p>
        <sec id="sec-6-1-1">
          <title>Salah Bourennane and Caroline Fossati.</title>
          <p>Comparison of shape descriptors for hand
posture recognition in video. Signal, Image
and Video Processing, 6(1):147{157, 2012.
[BGRS11] Andre LC Barczak, Andrew Gilman,
Napoleon H Reyes, and Teo Susnjak.
Analysis of feature invariance and
discrimination for hand images: Fourier descriptors
versus moment invariants. In International
Conference Image and Vision Computing</p>
          <p>New Zealand IVCNZ2011, 2011.
[CBM07]</p>
          <p>Simon Conseil, Salah Bourennane, and
Lionel Martin. Comparison of fourier
descriptors and hu moments for hand
posture recognition. In Signal Processing
Conference, 2007 15th European, pages 1960{
1964. IEEE, 2007.
[GK15b]
[GK16]</p>
        </sec>
        <sec id="sec-6-1-2">
          <title>Eriglen Gani and Alda Kika. Identi kimi</title>
          <p>i dores nepermjet teknologjise microsoft
kinect. Buletini i Shkencave te Natyres,
20:82{90, 2015.</p>
        </sec>
        <sec id="sec-6-1-3">
          <title>Eriglen Gani and Alda Kika. Review on</title>
          <p>natural interfaces technologies for
designing albanian sign language recognition
system. The Third International Conference
On: Research and Education Challenges
Towards the Future, 2015.</p>
        </sec>
        <sec id="sec-6-1-4">
          <title>Archana S Ghotkar and Gajanan K</title>
          <p>Kharate. Dynamic hand gesture
recognition and novel sentence interpretation
algorithm for indian sign language using
microsoft kinect sensor. Journal of Pattern
Recognition Research, 1:24{38, 2015.</p>
        </sec>
        <sec id="sec-6-1-5">
          <title>Eriglen Gani and Alda Kika. Albanian</title>
          <p>sign language (AlbSL) number
recognition from both hand's gestures acquired
by kinect sensors. International Journal
of Advanced Computer Science and
Applications, 7(7), 2016.</p>
        </sec>
        <sec id="sec-6-1-6">
          <title>Yun Liu, Zhijie Gan, and Yu Sun.</title>
          <p>Static hand gesture recognition and
its application based on support
vector machines. In Software
Engineering, Arti cial Intelligence, Networking,
and Parallel/Distributed Computing, 2008.
SNPD'08. Ninth ACIS International
Conference on, pages 517{521. IEEE, 2008.</p>
        </sec>
        <sec id="sec-6-1-7">
          <title>MSDN. Kinect for windows sensor components and speci cations, April 2016.</title>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>Theodosios Pavlidis. Algorithms for graph</title>
        <p>ics and image processing. Springer Science
&amp; Business Media, 2012.
[SB13]</p>
        <p>
          Kalin Stefanov and Jonas Beskow. A
kinect corpus of swedish sign language
signs.
          <xref ref-type="bibr" rid="ref5">In Proceedings of the 2013</xref>
          Workshop on Multimodal Corpora: Beyond
Audio and Video, 2013.
[Sud14]
[TL11]
[Ull11]
[VAC13]
[WA12]
[ZL+02]
[ZY14]
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [GK15c]
          <article-title>[SSKK16] S Shruthi, KC Sona, and S Kiran Kumar</article-title>
          .
          <article-title>Classi cation on hand gesture recognition and translation from real time video using svm-knn</article-title>
          .
          <source>International Journal of Applied Engineering Research</source>
          ,
          <volume>11</volume>
          (
          <issue>8</issue>
          ):
          <volume>5414</volume>
          {
          <fpage>5418</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Bh</given-names>
            <surname>Sudantha</surname>
          </string-name>
          .
          <article-title>A portable tool for deaf and hearing impaired people</article-title>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Pedro</given-names>
            <surname>Trindade</surname>
          </string-name>
          and
          <string-name>
            <given-names>Jorge</given-names>
            <surname>Lobo</surname>
          </string-name>
          .
          <article-title>Distributed accelerometers for gesture recognition and visualization</article-title>
          .
          <source>In Technological Innovation for Sustainability</source>
          , pages
          <volume>215</volume>
          {
          <fpage>223</fpage>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Fahad</given-names>
            <surname>Ullah</surname>
          </string-name>
          .
          <article-title>American sign language recognition system for hearing impaired people using cartesian genetic programming</article-title>
          .
          <source>In Automation, Robotics and Applications (ICARA)</source>
          ,
          <year>2011</year>
          5th International Conference on, pages
          <volume>96</volume>
          {
          <fpage>99</fpage>
          . IEEE,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>In Image Information Processing (ICIIP)</source>
          ,
          <year>2013</year>
          IEEE Second International Conference on, pages
          <volume>96</volume>
          {
          <fpage>100</fpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jarrett</given-names>
            <surname>Webb</surname>
          </string-name>
          and
          <string-name>
            <given-names>James</given-names>
            <surname>Ashley</surname>
          </string-name>
          .
          <article-title>Beginning Kinect Programming with the Microsoft Kinect SDK</article-title>
          . Apress,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [WKSE02]
          <string-name>
            <given-names>Juan</given-names>
            <surname>Wachs</surname>
          </string-name>
          , Uri Kartoun, Helman Stern, and
          <string-name>
            <given-names>Yael</given-names>
            <surname>Edan</surname>
          </string-name>
          .
          <article-title>Real-time hand gesture telerobotic system using fuzzy c-means clustering</article-title>
          .
          <source>In Automation Congress</source>
          ,
          <source>2002 Proceedings of the 5th Biannual World</source>
          , volume
          <volume>13</volume>
          , pages
          <fpage>403</fpage>
          {
          <fpage>409</fpage>
          . IEEE,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Dengsheng</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Guojun
          <string-name>
            <surname>Lu</surname>
          </string-name>
          , et al.
          <article-title>A comparative study of fourier descriptors for shape representation and retrieval</article-title>
          .
          <source>In Proc. 5th Asian Conference on Computer Vision</source>
          . Citeseer,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Yanmin</given-names>
            <surname>Zhu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Bo</given-names>
            <surname>Yuan</surname>
          </string-name>
          .
          <article-title>Real-time hand gesture recognition with kinect for playing racing video games</article-title>
          .
          <source>In 2014 International Joint Conference on Neural Networks (IJCNN)</source>
          . Institute of Electrical &amp; Electronics
          <string-name>
            <surname>Engineers</surname>
          </string-name>
          (IEEE),
          <year>jul 2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>