<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Declarative and Classi er Gesture Recognition Method for Creating an E ective Feedback and Feedforward System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessandro Carcangiu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Deptment of Electrical and Electronic Engineering, University of Cagliari</institution>
          ,
          <addr-line>Cagliari 09123</addr-line>
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>For recognizing gestures in an interactive application, we could apply di erent approaches, which o er many advantages and disadvantages. On the one hand, Machine Learning techniques guarantee a high accuracy and robustness to noise, but they do support the creation of e ective feedback and feed-forward systems since they do not provide intermediate information. On the other hand, we can nd in the literature declarative and compositional methods, which allow gesture sub-parts identi cation at the cost of a lower recognition rate. Therefore, one of the most common problems in gesture interface development is how to set the trade-o between a high precision in recognition and the support for user's guidance. With this mind, the main goal of my Ph.D. is nding a way for lling the gap between these two approaches, bridging the gap between Machine Learning and declarative and compositional approaches.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        In recent years, gesture interaction gained importance in many interactive
settings, such as e.g., houses, o ces and hospitals. Nowadays, smartphones, tablets,
laptop and desktop computers are equipped with gesture input devices, which
complement more standard techniques like pointing and typing. In general, we
de ne an interactive gesture as a body movement that recognized through motion
sensing input device. The tracking capabilities in uence di erent aspect in the
interaction design. For instance di erent devices track di erent body parts:
ngertips (touchscreens), hands and ngers (Leap Motion [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], Intel RealSense [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]),
arms or legs (Kinect [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]).
      </p>
      <p>
        Despite their spread, window toolkits provide little support for managing
gestures in interface development: they allow either to de ne high-level events
or to track synchronously the raw data coming from the device. This increases
the development complexity and, in many cases, decreases the usability of the
resulting interface. A gesture is not an atomic event like a mouse-click: it
usually requires di erent movements and may last for many seconds, which is a
perceivable timespan for a user. During that time, s/he user needs guidance in
completing the movement and s/he should be informed about the system
interpretation of the input. For these reasons, developers need to provide users
with advices on gesture performance during the interaction: i) which movements
were recognized by the application (de ned as feedback) and ii) the possible
interface state(s) reachable after the user nishes the current gesture (de ned as
feedforward) [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        While introducing feedback and feedforward in gesture interfaces has positive
e ects on the user experience (UX) [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], developers have technical di culties in
building such systems. On the one hand, the interface needs a high recognition
accuracy. This problem was addressed through di erent machine learning methods,
like Hidden Markov Models (HMM), Dynamic Time Warping (DTW),
TimeDelay Neural Networks (TDNN) and Finite-State Machines (FTM) [
        <xref ref-type="bibr" rid="ref11 ref14 ref5">11,14,5</xref>
        ],
support vector machines [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. All methods demonstrated a high recognition
accuracy and they are suitable for recognizing complex gestures. However, they
do not support composition, they need many examples in the training phase
and they usually provide a single class label for the whole gesture, which is
represented as an atomic event even if the time dimension is internally taken
into account by the recognition approach. On the other hand, we can nd in
the literature di erent methods for describing the temporal evolution of a
gesture through composition, for instance GestIT [
        <xref ref-type="bibr" rid="ref18 ref19">18,19</xref>
        ] and Proton++ [
        <xref ref-type="bibr" rid="ref8 ref9">9,8</xref>
        ]. They
declaratively de ne a gesture through a set of composed sub-parts, allowing their
identi cation during the performance. The drawback in these methods is the
accuracy: the sub-part identi cation relies on heuristics which not are as precise as
the classi cation approaches. With this mind, we have proposed DEICTIC[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ],
which integrates the declarative and compositional gesture description model
GestIT [
        <xref ref-type="bibr" rid="ref18 ref19">18,19</xref>
        ] with HMMs and uses HMMs in order to recognize basic gesture
segments (or primitives) instead of whole gestures.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>In this section we will brie y summarise the state of the art for the approaches
we are trying to bridge.
2.1</p>
      <sec id="sec-2-1">
        <title>Machine Learning-based Approaches</title>
        <p>
          Computer Vision-based approaches have just tried to identify gesture sub-parts
in order to: i) reduce the size of the training dataset and ii) improve their
performances. Chen et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] de ne primitives in a context-grammar established in
advance using a top down approach, which is more suitable to UI designers;
however, grammars were not created taking into account the gesture meaning from
the user perspective. In [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ], Yang et al. identi ed primitives by using a
bottomup clustering approach, aimed at reducing the training set size and at improving
the organisation of unlabeled datasets for speeding up its processing. A gesture
is labelled and de ned with sequences of primitives. This representation at rst
may be useful for building UIs. Unfortunately, the automatic identi cation of
primitives usually lead to a set of basic gestures that is di cult to understand
for designers, and this decreases their applicability for creating feedback and
feed-forward systems.
        </p>
        <p>
          In [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], primitives are used together a three-level HMM classi er architecture
for recognizing i) the primitives, ii) their composition and iii) the pose or gesture.
However, also in this case unsupervised learning was used for de ning both
primitives and their composition, which is not suitable for building UIs.
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Compositional and Declarative-based Approaches</title>
        <p>
          As mentioned earlier, declarative approaches allow splitting a gesture into
several sub-components and there are di erent compositional approaches based on
heuristic gesture recognition. An example could be GestIT [
          <xref ref-type="bibr" rid="ref18 ref19">18,19</xref>
          ] and
Proton++ [
          <xref ref-type="bibr" rid="ref8 ref9">9,8</xref>
          ]. The latter, is a multi-touch framework, which allows describing
declaratively custom gestures, separating the temporal sequencing of the events
from the related code to the user interface behaviour. It also allows developers to
describe custom gestures in a declarative way, through regular expressions and
using the operators of concatenation, alternation and Kleene's star. Developers
could de ne a regular expression by a triplet which is composed by an event
type, a touch identi er and the interface item hit by the touch.
        </p>
        <p>
          In GestIT [
          <xref ref-type="bibr" rid="ref18 ref19">18,19</xref>
          ] gestures are modelled through expressions that de ne their
temporal evolution. The expressions are obtained by using two main elements: i)
ground term or primitive, which describes the smallest part of a gesture and in
general is associated to a fundamental movement which may considered atomic
by users, ii) and composite terms, which represent a set of operators allowing
to link ground terms between them or other composite terms, in order to de ne
more complex gesture. To describe the set of operators consider two gesture g
and h (either ground or composite terms): g is the continue iteration of g; g h
de nes the sequence that connects g with h, rstly the user performs g then h;
g k h de nes that g and h are performed in parallel simultaneously; g[ ]h is the
choice between either g or h; g[&gt; h disables the iteration of g by performing h;
g j=j h the connected gesture can be performed in any order (e.g. rst doing g
and then h or vice versa).
        </p>
        <p>
          Another example is Midas [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] that introduced a rule-based approach for
multitouch gestures. These rules work on di erent features, for example the 2D
positions, the nger tracking state or its speed and consists of two components:
a prerequisite part and an action part.
As we discussed in previous sections, feedback and feedforward systems were
developed to help user during the interaction, in particular for inexperienced
users who interface with specialised machines which have a low a ordance or
in situations where the users lack familiarity with the interface. These systems
help users in di erent ways, for example showing what will be the result of a
particular input or explaining how to complete a certain operation.
        </p>
        <p>
          In gesture interfaces feedback is used to show which portion of the gesture has
been recognized and what the system understood from the user's movements.
These suggestions can help the users to understand what movements they are
doing and, eventually, to repair both recognition or execution errors. Feedforward
helps users during the interaction di erently i.e., showing the e ects on the
UI of the possible ways for concluding the current gesture, thus guiding the
user towards the desired e ect. An e ective visualization of such guidance has
been described by Bau and Mackay [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], through a dynamic guide that combines
feedback and feedforward to help users to learn, execute and remember gesture
sets. Their approach shows a graphic representation of the gesture part that has
been correctly recognized and, at the same time, a representation for all possible
paths that may be followed for completing a correct interaction.
        </p>
        <p>
          LightGuide [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] proposes a similar approach in a 3D setting, showing
feedback and feedforward directly on the body parts which user should move to
perform a gesture. This approach using a set of projector and Kinects to show
the suggestions.
        </p>
        <p>
          Schwarz et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] proposed another model which is a general architecture
with the goal of proving support for continuous feedback about uncertainty.
Their work is based on prior work in modelling uncertainty using Monte Carlo
and tracks multiple interfaces; indeed, this architecture shows all these possible
new system states for each sequence of input that the user may have intended. It
works reducing the number of possible interfaces (according to the user input),
combining and showing these in a single interface. Summarising, the model shows
to user which will be the new program state basing on user input and the most
likely states whether there are some similar actions. In addition, when the user
completes the task, the architecture allows user to return to the previous state
or to select another state.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>DEICTIC</title>
      <p>
        As we said in previous section, de nitively DEICTIC we combined the
declarative and compositional gesture description model GestIT [
        <xref ref-type="bibr" rid="ref18 ref19">18,19</xref>
        ] with HMMs
in order to integrates the declarative advantages with the accuracy and the
robustness to input o ered by classi ers. DEICTIC, similarly to GestIT, de nes
a gesture by using ground and composite terms. In our approach, a primitive
movement is recognized by a single left-to-right trained HMM while the
composite operators are obtained by combining the HMMs of their operands. The result
of composition is a new HMM which provides information during the recognition
and recognizes the gesture according to the temporal semantics. The two follow
paragraphs will explain how DEICTC uses these components.
3.1
      </p>
      <sec id="sec-3-1">
        <title>Ground Terms.</title>
        <p>
          A basic movement in DEICTIC is described by a single HMM which uses a
leftto-right (or Bakis [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]) topology. For de ning a \basic" HMM, we must de ne his
number of states and train the HMM for learning the probability distributions
of both transitions and observations from a set of correct performance samples.
        </p>
        <p>In DEICTIC we de ned three 2D primitives:
1. Point which de nes the starting position of the stroke. It is de ned
specifying the x and y coordinates and its notations is P (x; y);
2. Line de nes a linear movement of a speci ed o set in the x and y axes
starting from the current position. It could be used to describe a vertical,
horizontal or diagonal movement and we represent a line with the notation
L( x; y);
3. Arc speci es a quarter of a circle starting from the current position and
nishing at the speci ed o set. The arc could follow a clockwise or
counterclockwise direction and we represent it with the notation A ( x; y)
(clockwise direction) or A ( x; y) (counter-clockwise direction).
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Composite Terms</title>
        <p>DEICTIC allows de ning complex gesture by using ve GestIT composite terms:
iteration, sequence, choice, disabling and parallel (we do not use order
independence because it could be derived from the other operators). These composite
terms are generated in DEICTIC starting from the HMM of their operands
automatically without requiring a new training step for the generated HMM.
We explain in the following list how the composition is accomplished for each
operator.</p>
        <p>{ Sequence. Consider a generic stroke g, its subcomponents and their HMMs.</p>
        <p>The HMM for recognizing g is obtained by connecting the end state of its
rst subcomponent with the start state of the second component and so on.
The number of states of the new HMM is made up of the sum of the numbers
of states of each model.
{ Iterative The HMM for recognizing the iteration of the generic stroke g is
obtained, starting from the basic HMM, by adding a transition from all states
that are connected with the ending state to all states that are connected
with the start state. In this way, we obtain a loop in the topology without
changes in the probability distributions. Obviously, the number of states of
the composed HMM will be the same of the basic HMM.
{ Choice. Consider now the generic strokes g. Now, suppose the possibility to
describe g through two di erent primitive sequences: g0 and g00. The choice
between the two descriptions is obtained putting the original HMMs in two
separate recognition lines, without any transition between their states. The
only contact points are the starting and the nal states of the new HMM. The
HMM obtained, once again, has a number of states which growns linearly.
{ Disabling. Consider the case of two generic strokes g and h and the
disabling operator. In DEICTIC the disabling operator for example g [&gt; h,
and its HMM are obtained in two steps: rst, we add a transition between
the starting states of each ground term in g stroke to the starting state of
h; second, a link is added from the ending state of g to the starting state of
h. This is necessary for supporting the case when the user performs the rst
gesture entirely before starting h. For maintaining the outgoing probability
property, we split the original transitions likelihood among all involved arcs.
{ Parallel. Finally, consider the possibility to do in parallel two di erent
strokes, g and h that are totally independent in their transitions. The
composite HMM for g k h will contain a state for each pair (stateg,stateh), that
are linked if only the transition is valid in both starting models. For example,
if we are in the state (g1; h1), we can: i) go forward on h, e.g. g1; h2; ii) go
forward on g, g2; h1 or iii) go forward on both, g2; h2). The hmm obtained
has a number of states equal to n m, where n is the number of states of Z
and m the number of states of . Since the two gestures are independent the
transition probabilities of the composite HMM are given from the product
of probabilities of each gesture.
In the test phase, our main goal is to prove the e ciency and the e cacy of our
approach recognizing gestures. First of all, we would prove if DEICTIC is usable
in a feedback and feed-forward systems. Second, we would done a comparison
between DEICTIC and machine learning-based approach about accuracy rate.
For doing it, we choose ad-hoc HMM that is a hmm trained with the whole
gesture.</p>
        <p>
          Taking this into consideration, we tested DEICTIC with a dataset created by
ourself and just used on our previous work about DEICTIC [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. It is made up
by 10 gestures strokes with 60 samples for each one. They are performed by 14
di erent people and recorded through Leap Motion device.
        </p>
        <p>In the preliminary test we used a subset of these gesture. The hmms created
with DEICTIC are trained only if they are used for recognizing basic movement;
conversely, the ad-hoc HMMs are trained with the whole gesture and we
evaluated the recognition performance using the leave-one-out technique. The 1 shows
their description model and the recognition rate results obtained with DEICTIC
and ad-hoc HMMs.</p>
        <p>Gesture</p>
        <p>Model
!
V
V
4</p>
        <p>P oint(0; 2)</p>
        <p>P oint(0; 0) Line( 1; 0)</p>
        <p>P oint(0; 0) Line(1; 0)
P oint(2; 2) Line(1; 2) Line(1; 2)</p>
        <p>P oint(0; 0) Line(1; 2) Line(1; 2)
P oint(0; 0) Line(2; 2) Line(0; 2)</p>
        <p>Line( 4; 2)
Line( 1; 2) Line(2; 0)</p>
        <p>Line( 1; 2)
P oint(0; 1) Line(0; 1) Line(2; 0)</p>
        <p>Line(0; 1) Line( 1; 0)</p>
        <p>DEICTIC
Summarising, DEICTIC is much accurate than the declarative methods and
allows automatic generation of composed HMM. In addition, taking advantage of
HMM, DEICTIC predicts the continuation of a gesture and permits sub-part
identi cation helping developers to create feedback and feed-forward systems.
Last but not least, the HMMs obtained from composition do not require
additional training with respect to the basic HMMs; it should be noted that, for
a particular gesture, the time asked to train and compose its basic HMMs is
much less than is required to de ne and train its ad-hoc HMM. However our
approach has some limits. First of all it does not support on-line recognition e
has not the same accuracy of ad-hoc HMM. Another problem is the number of
states of composed HMMs: it tends to increase or linearly in sequence, choice
and disabling case or quadratically in parallel case.</p>
        <p>In order to improve DEICTIC, we have thought to substitute HMMs trying other
machine learning-based method. We have looked for an approach which
manages sequential data, supports on-line recognition and preserves a high precision
rate. With this mind, we have started to study neural networks, in particular
the Time Delay Neural Network, or TDNN. Now we do not know whether this
method could be useful for us and we think this doctoral consortium is the best
opportunity to obtain advices and suggestions.</p>
        <p>At the same time, we are working on: i) de ning an extensive set of
information which reports all the point that gesture recognizer methods must expose in
order to build e ective feedback and feed-forward systems; ii) creating a set of
guidelines for helping developers to select the appropriate classi cation technique
according to the information needed by the feedback and feed-forward system
at hand; iii) using DEICTIC in a real-time scenarios and testing its features
building a feedback and feed-forward systems.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>University Doctoral Program Context</title>
      <p>I am Alessandro Carcangiu, a Phd student at the department of electrical
and electronic engineering of the University of Cagliari. My tutors are
Giorgio Fumera, Fabio Roli (from the same department) and Lucio Davide Spano
(from the Dep. of Mathematics and Computer Science). In October 2015 I have
started my Phd and I planed to defend my thesis in October 2018.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bau</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mackay</surname>
          </string-name>
          , W.E.:
          <article-title>Octopocus: a dynamic guide for learning gesture-based command sets</article-title>
          .
          <source>In: Proceedings of the 21st annual ACM symposium on User interface software and technology</source>
          . pp.
          <volume>37</volume>
          {
          <fpage>46</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Carcangiu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Gesture recognition through declarative and classi er approach</article-title>
          .
          <source>In: Proceedings of the 22nd International Conference on Intelligent User Interfaces Companion</source>
          . pp.
          <volume>185</volume>
          {
          <fpage>188</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Carcangiu</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spano</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fumera</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roli</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Gesture modelling and recognition by integrating declarative models and pattern recognition algorithms</article-title>
          .
          <source>In: Proceedings of the 22nd International Conference on Intelligent User Interfaces Companion</source>
          . pp.
          <volume>185</volume>
          {
          <fpage>188</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Georganas</surname>
            ,
            <given-names>N.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petriu</surname>
            ,
            <given-names>E.M.</given-names>
          </string-name>
          :
          <article-title>Real-time vision-based hand gesture recognition using haar-like features</article-title>
          .
          <source>In: Proceedings of IMTC 2007</source>
          . pp.
          <volume>1</volume>
          {
          <issue>6</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Cheng, H.,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Survey on 3d hand gesture recognition</article-title>
          .
          <source>IEEE Trans. Circuits Syst. Video Techn</source>
          .
          <volume>26</volume>
          (
          <issue>9</issue>
          ),
          <volume>1659</volume>
          {
          <fpage>1673</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Elliott</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aggoun</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            ,
            <given-names>J.B.</given-names>
          </string-name>
          :
          <article-title>Hidden Markov models: estimation and control</article-title>
          , vol.
          <volume>29</volume>
          . Springer Science &amp; Business
          <string-name>
            <surname>Media</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Intel:
          <article-title>Intel realse</article-title>
          . http://www.intel.com/content/www/us/en/ architecture-and
          <article-title>-technology/realsense-overview</article-title>
          .html
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hartmann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DeRose</surname>
          </string-name>
          , T.,
          <string-name>
            <surname>Agrawala</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Proton++ :
          <string-name>
            <given-names>A Customizable</given-names>
            <surname>Declarative</surname>
          </string-name>
          <article-title>Multitouch Framework</article-title>
          .
          <source>In: Proceedings of UIST 2012</source>
          . pp.
          <volume>477</volume>
          {
          <fpage>486</fpage>
          . ACM Press, Berkeley, California, USA (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kin</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hartmann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DeRose</surname>
          </string-name>
          , T.,
          <string-name>
            <surname>Agrawala</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Proton: multitouch gestures as regular expressions</article-title>
          .
          <source>In: Proceedings of CHI 2012</source>
          . pp.
          <volume>2885</volume>
          {
          <fpage>2894</fpage>
          . ACM Press, Austin, Texas, USA (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>10. Microsoft: Kinect. http://www.xbox.com/en-CA/en-EN/xbox-one/accessories/ kinect</mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mitra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Acharya</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Gesture recognition: A survey</article-title>
          .
          <source>IEEE Trans. Systems, Man, and Cybernetics</source>
          , Part C
          <volume>37</volume>
          (
          <issue>3</issue>
          ),
          <volume>311</volume>
          {
          <fpage>324</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Motion</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Leap motion</article-title>
          . https://www.leapmotion.com/
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Natarajan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nevatia</surname>
          </string-name>
          , R.:
          <article-title>Online, real-time tracking and recognition of human actions</article-title>
          .
          <source>In: Proceedings of WMVC 2008</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Rautaray</surname>
            ,
            <given-names>S.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Vision based hand gesture recognition for human computer interaction: a survey</article-title>
          .
          <source>Artif. Intell. Rev</source>
          .
          <volume>43</volume>
          (
          <issue>1</issue>
          ),
          <volume>1</volume>
          {
          <fpage>54</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Scholliers</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Signer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>De Meuter</surname>
          </string-name>
          , W.:
          <article-title>Midas: a declarative multitouch interaction framework</article-title>
          .
          <source>In: Proceedings of TEI 2011</source>
          . pp.
          <volume>49</volume>
          {
          <fpage>56</fpage>
          . TEI '11,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Schwarz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manko</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hudson</surname>
            ,
            <given-names>S.E.</given-names>
          </string-name>
          :
          <article-title>An architecture for generating interactive feedback in probabilistic user interfaces</article-title>
          .
          <source>In: Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems</source>
          . pp.
          <volume>2545</volume>
          {
          <fpage>2554</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Sodhi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benko</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Lightguide: projected visualizations for hand movement guidance</article-title>
          .
          <source>In: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems</source>
          . pp.
          <volume>179</volume>
          {
          <fpage>188</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Spano</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cisternino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paterno</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A Compositional Model for Gesture Definition</article-title>
          .
          <source>In: Proceedings of HCSE 2012</source>
          . pp.
          <volume>34</volume>
          {
          <fpage>52</fpage>
          . Springer (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Spano</surname>
            ,
            <given-names>L.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cisternino</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paterno</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fenu</surname>
          </string-name>
          , G.:
          <article-title>GestIT: a Declarative and Compositional Framework for Multiplatform Gesture De nition</article-title>
          .
          <source>In: Proceedings of EICS 2013</source>
          . pp.
          <volume>187</volume>
          {
          <fpage>196</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Vermeulen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luyten</surname>
          </string-name>
          , K., van den Hoven, E.,
          <string-name>
            <surname>Coninx</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Crossing the bridge over norman's gulf of execution: revealing feedforward's true identity</article-title>
          .
          <source>In: Proceedings CHI</source>
          <year>2013</year>
          . pp.
          <year>1931</year>
          {
          <year>1940</year>
          . ACM (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saleemi</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Discovering motion primitives for unsupervised grouping and one-shot learning of human actions, gestures, and expressions</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>35</volume>
          (
          <issue>7</issue>
          ),
          <volume>1635</volume>
          {
          <fpage>1648</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>