<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Multiparadigm Approach to Integrate Gestures and Sound in the Modeling Framework</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vasco Amaral</string-name>
          <email>vasco.amaral@fct.unl.pt</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonio Cicchetti</string-name>
          <email>antonio.cicchetti@mdh.se</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Romuald Deshayes</string-name>
          <email>romuald.deshayes@umons.ac.be</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Malardalen Research and Technology Centre (MRTC)</institution>
          ,
          <addr-line>Vasteras</addr-line>
          ,
          <country country="SE">Sweden</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Software Engineering Lab, Universit ́e de Mons-Hainaut</institution>
          ,
          <country country="BE">Belgium</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidade Nova de Lisboa</institution>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>57</fpage>
      <lpage>66</lpage>
      <abstract>
        <p>One of the essential means of supporting Human-Machine Interaction is a (software) language, exploited to input commands and receive corresponding outputs in a well-defined manner. In the past, language creation and customization used to be accessible to software developers only. But today, as software applications gain more ubiquity, these features tend to be more accessible to application users themselves. However, current language development techniques are still based on traditional concepts of human-machine interaction, i.e. manipulating text and/or diagrams by means of more or less sophisticated keypads (e.g. mouse and keyboard). In this paper we propose to enhance the typical approach for dealing with language intensive applications by widening available human-machine interactions to multiple modalities, including sounds, gestures, and their combination. In particular, we adopt a Multi-Paradigm Modelling approach in which the forms of interaction can be specified by means of appropriate modelling techniques. The aim is to provide a more advanced human-machine interaction support for language intensive applications.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The mean of supporting Human-Machine interaction are languages: a
welldefined set of concepts that can be exploited by the user to compose more or less
complex commands to be input to the computing device. Given the dramatic
growth of software applications and their utilization in more and more complex
scenarios, there has been a contemporary need to improve the form of interaction
in order to reduce users’ e↵ ort. Notably, in software development, it has been
introduced di↵ erent programming language generations, programming paradigms,
and modelling techniques aiming at raising the level of abstraction at which the
problem is faced. In other words, abstraction layers have been added to close
the gap between machine language and domain-specific concepts, keeping them
interconnected through automated mechanisms.</p>
      <p>While the level of abstraction of domain concepts has remarkably evolved,
the forms of interaction with the language itself indubitably did not. In
particular, language expressions are mainly text-based or a combination of text and
diagrams. The underlying motivation is that, historically, keyboard and mouse
have been exploited as standard input devices. In this respect, other forms of
interaction like gestures and sound have been scarcely considered. In this paper
we discuss the motivations underlying the need of enhanced forms of interaction
and propose a solution to integrate gestures and sound in modeling frameworks.
In particular, we define language intensive applicative domains the cases where
either the language evolves rapidly, or language customizations are part of the
core features of the application itself.</p>
      <p>In order to better grasp the previously mentioned problem, we consider a
sample case study in the Home Automation Domain, where a language has to be
provided as supporting di↵ erent automation facilities for a house, ranging from
everyday life operations to maintenance and security. This is a typical language
intensive scenario since there is a need to customize the language depending on
the customer’s building characteristics. Even more important, the language has
to provide setting features enabling a customer to create users’ profiles: notably,
children may command TV and lights but they shall not access kitchen
equipment. Likewise, the cleaning operator might have access to a limited amount of
rooms and/or shall not be able to deactivate the alarm.</p>
      <p>In the previous and other language intensive applicative domains it is
inconceivable to force users to exploit the “usual” forms of interaction for at least two
reasons: i) if they have to digit a command to switch on the lights they could
use the light switch instead and hence would not invest money in these type
of systems. Moreover, some users could be unable to exploit such interaction
techniques, notably disabled, children, and so forth; ii) the home automation
language should be easily customizable, without requiring programming skills.
It is worth noting that language customization becomes a user’s feature in our
application domain, rather than a pure developer’s facility. If we widen our
reasoning to the general case, the arguments mentioned so far can be referred to
as the need of facing accidental complexity. Whenever a new technology is
proposed, it is of paramount importance to ensure that it introduces new features
and/or enhances existing ones, without making it more complex to use, otherwise
it would not be worth to be exploited.</p>
      <p>Our solution is based on Multi-Paradigm Modelling (MPM) principles, i.e.
every aspect of the system has to be appropriately modeled and specified,
combining di↵ erent points of view of the system being then possible to derive the
concrete application. In this respect, we propose to precisely specify both the
actions and the forms of language interactions, in particular gestures and sound,
by means of models. In this way, flexible interaction modes with a language are
possible as well as languages accepting new input modalities such as sounds and
gestures.</p>
      <p>The remaining of the paper is organized as follows: sec. 2 discusses the
stateof-the-art; sec. 3 presents the communication between a human and a machine
through a case study related to home automation; sec. 4 introduces an interaction
metamodel and discusses how it can be used to generate advanced concrete
syntaxes relying on multiple modalities; finally, sec. 5 concludes.
2</p>
      <p>State of the Art
This section describes basic concepts and state-of-the-art techniques typically
exploited in sound/speech and gesture recognition together with their
combinations created to provide advanced forms of interaction. Our aim is to illustrate
the set of concepts usually faced in this domain and hence to elicit the
requirements for the interaction language we will introduce in Section 4.</p>
    </sec>
    <sec id="sec-2">
      <title>Sound and speech recognition</title>
      <p>
        Sound recognition is usually used for command-like actions; a word has to be
recognized before the corresponding action can be triggered. With the Vocal
Joystick [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] it is possible to use acoustic phonetic parameters to continuously
control tasks. For example, in a WIMP (Windows, Icons, Menus, Pointing device)
application, the type of vowel can be used to give a direction to the mouse cursor,
and the loudness can be used to control its velocity.
      </p>
      <p>
        In the context of environmental sounds, [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] proposed di↵ erent techniques
based on Support Vector Machines and Hidden Markov Models to detect and
classify acoustic events such as foot steps, a moving chair or human cough.
Detecting these di↵ erent events help to better understand the human and social
activities in smart-room environments. Moreover, an early detection of non-speech
sounds can help to improve the robustness of automatic speech recognition
algorithms.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Gesture recognition</title>
      <p>
        Typically, gesture recognition systems resort to various hardware devices such
as data glove or markers [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], but more recent hardware such as the Kinect or
other 3D sensors enable unconstrained gestural interaction [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. According to a
survey of gestural interaction [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], gestures can be of 3 types :
– hand and arm gestures: recognition of hand poses or signs (such as
recognition of sign language);
– head and face gestures: shaking head, direction of eye gaze, opening the
mouth to speak, happiness, fear, etc;
– body gestures: tracking movements of two people interacting, analyzing
movement of a dancer, or body poses for athletic training.
      </p>
      <p>
        The most widely used techniques for dynamic gestural recognition usually involve
hidden Markov models [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], particle filtering [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] or finite state machines [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>Multimodal systems</title>
      <p>
        The ”Put-that-there” [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] system, developed in the 80’s, is considered to be
the origin of human-computer interaction regarding the use of voice and sound.
Vocal commands such as ”delete this elements” while pointing at one object
displayed on the screen can be correctly processed by the system. According to
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], using multiple modalities, such as sound and gestures, helps to make the
system more robust and maintainable. They also proposed to split audio sounds
in two categories: human speech and environmental sounds.
      </p>
      <p>
        Multimodal interfaces involving speech and gestures have been widely used
for text input, where gestures are usually used to choose between multiple
possible utterances or correct recognition errors [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ]. Some other techniques
propose to use gestures on a touchscreen device in addition to speech recognition
to correct recognition errors [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Both modalities can also be used in an
asynchronous way to disambiguate between the possible utterances [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        More recently, the SpeeG system [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] has been proposed. It is a multimodal
interface for text input and is based on the Kinect sensor, a speech recognizer and
the Dasher [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] user interface. The contribution lies in the fact that, unlike the
aforementioned techniques, the user can perform speech correction in real-time,
while speaking, instead of doing it in a post processing fashion.
      </p>
      <p>
        Human-computer interaction modeling
In the literature, many modeling techniques have been used to represent
interaction with traditional WIMP user interfaces. For example, statecharts have been
dedicated to the specification and design of new interaction objects or widgets
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        Targetting virtual reality environment, Flownets [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is a modeling tool,
relying on high-level Petri nets, based on a combination of discrete and continuous
behavior to specify the interaction with virtual environments. Also based on
high-level Petri nets for dynamic aspects and an object oriented framework, the
Interactive Cooperative Objects (ICO) formalism [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] has been used to model
WIMP interfaces as well as multimodal interactions in virtual environments. In
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], a virtual chess game was developed in which the user can use a data glove to
manipulate virtual chess pieces. In [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], ICO has been used to create a framework
for describing gestural interaction with 3D objects has been proposed.
      </p>
      <p>
        Providing a UML based generic framework for modeling interaction
modalities such as speech or gestures enables software engineers to easily integrate
multimodal HCI in their applications. That’s the point defended by [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. In
their work, the authors propose a metamodel which focuses on the aspects of
an abstract modality. They distinguish between simple and complex modalities.
The first one represents a primitive form of interaction while the second
integrates other modalities and uses them simultaneously. With the reference point
being the computer, input and output modalities are defined as a specification of
simple modality. Input modalities can be event-based (e.g. performing a gesture
or sending a vocal command) or streaming based (e.g. drawing a circle or
inputting text using speech recognition) and output modalities are used to provide
static (e.g. a picture) or dynamic (e.g. speech) information to the user.
As discussed so far, di↵ erent techniques supporting more advanced forms of
interaction between humans and machines have already been proposed. Nonetheless,
their exploitation in current software languages has been noticeably limited.
This work proposes to widen the modalities of human-machine interaction as
depicted in Fig. 1. In general, a human could input information by means of
speech, gestures, texts, drawings, and so forth. The admitted ways of
interaction are defined in an interaction model, which also maps human inputs into
corresponding machine readable formats. Once the machine has completed its
work, it outputs the results to the user through sounds, diagrams, texts, etc.;
also in this case the interaction model prescribes how performed computations
should be rendered to a human comprehensible format.
      </p>
      <p>A software language supports the communication between humans and
machines by providing a set of well-defined concepts that typically abstract real-life
concepts. Hence, software language engineering involves the definition of three
main aspects: i) the internal representation of the selected concepts,
understandable by the machine and typically referred to as abstract syntax ; ii) how the
concepts are rendered to the users in order to close the gap with the
applicative domain, called concrete syntax ; iii) how the concepts can be interpreted to
get/provide domain-specific information, referred to as semantics. Practically, a
metamodel serves as a base for defining the structural arrangement of concepts
and their relationships (abstract syntax), the concrete syntax is “hooked” on
appropriate groups of its elements, and the semantics is generically defined as
computations over elements.</p>
      <p>In order to better understand the role of these three aspects, Fig. 2 and 1
illustrate excerpts of the abstract and concrete syntaxes, respectively, of a sample
Home Automation DSL. In particular, a home automation system manages a
House (see Fig. 2 right-hand side) that can have several rooms under domotic
control (i.e. Room and DomoticControl elements in the metamodel, respectively).
The control is composed by several devices and actions that can be performed
with them. Notably, a Light can be turned on and o↵ , while a Shutter can be
opened and closed.</p>
      <p>It is easy to notice that the abstract syntax representation would not be
user-friendly in general, hence a corresponding concrete syntax can be defined
in order to provide the user with easy ways of interaction. In particular, Prog.
1 shows an excerpt of the concrete syntax definition for the home automation
DSL using the Eugenia tool4. The script prescribes to depict a House element as
a graph node showing the rooms defined for the house taken into account.
Prog. 1 Concrete Syntax mapping of the DSL for Home Automation with
Eugenia
@gmf.node(label="HouseName", color="255,150,150", style="dash")
class House {
@gmf.compartment(foo="bar")
val Room[*] hasRooms;
attr String HouseName;
}</p>
      <p>Despite the remarkable improvements in language usability thanks to the
addition of a concrete syntax, the malleability of interaction modalities provided by
Eugenia and other tools (e.g. GMF5) is limited to the standard typing and/or
4 http://www.eclipse.org/epsilon/doc/eugenia/
5 http://www.eclipse.org/modeling/gmp/
drawing graphs. Such a limitation becomes evident when needing to provide
users with extended ways of interaction, notably defining concrete syntaxes as
sounds and/or gestures. For instance, a desirable concrete syntax for the home
automation metamodel depicted in Fig. 2 would define voice commands for
turning lights on and o↵ , or alternatively prescribe certain gestures to do the same
operations.</p>
      <p>By embracing the MPM vision, which prescribes to define any aspect of the
modeling activity as a model, next Section introduces a language for defining
advanced human-machine interactions. In turn, such a language can be combined
with abstract syntax specifications to provide DSLs with enhanced concrete
syntaxes.
4</p>
      <p>Interaction Modeling
The proposed language tailored to enhanced human-machine interaction
concrete syntax definition is shown in Fig. 3. In particular, it depicts the
metamodel to define advanced concrete syntaxes, encompassing sound and gestures,
while the usual texts writing and diagrams drawing are treated as particular
forms of gestures. Going deeper, elements of the abstract syntax can be linked
to (sequences of) activities (see Activity on the bottom-left part of Fig. 3). An
activity, in turn, can be classified as a Gesture or a Sound.
Regarding gestures, the language supports the definition of di↵ erent types
of primitive actions typically exploited in gesture recognition applications. A
Gesture is linked to the BodyPart that is expected to perform the gesture. This
way we can distinguish between performed actions, for example, by a hand or a
head. Move is the simplest action a body part can perform, it is triggered for each
displacement of the body part. Dragging (Drag) can only be triggered by a hand
a represents a displacement of a closed hand. ColinearDrag and NonColinearDrag
represent a movement of both hands going in either the same or opposite
directions while being colinear or not colinear, respectively. Open and Close are
triggered when the user opens or closes the hand.</p>
      <p>Sounds can be separated in two categories: i) Voice represents human speech,
which is composed of Sentences and/or Words. ii) Audio that relates to all non
speech sounds that can be encountered, such as knocking on a door, a guitar
chord or even someone screaming. As it is very generic, it is characterized by
fundamental aspects as tone, pitch, and intensity.</p>
      <p>Complex activities can be created by combining multiple utterances of sound
and gestures. For example one could define an activity to close a shutter by
closing the left hand and dragging it from top to bottom while pointing at the
shutter and saying “close shutters”.</p>
      <p>In order to better understand the usage of the proposed language, Fig. 4
shows a simple example defining the available concrete syntaxes for specifying
the turning the lights on command. In particular, it is possible to wave the right
hand, first left and then right to switch on light 1 (see the upper part of the
picture). Alternatively, it is possible to give a voice command made up of the
sound sequence “Turn Light 1 On”, as depicted in the bottom part of the figure.</p>
      <p>It is worth noting that, given the purpose of the provided interaction forms,
the system can be taught to recognize particular patterns as sounds, speeches,
gestures, or their combinations. In this respect, the technical problems related
to recognition can be alleviated. Even more important, the teaching process
discloses the possibility to extend the concrete syntax of the language itself,
since additional multimodal commands can be introduced as alternative ways of
interaction.
5</p>
      <p>Conclusions
The ubiquity of software applications is widening the modeling possibilities of
end users who may need to define their own languages. In this paper, we defined
these contexts as language-intensive applications given the evolutionary pressure
the languages are subject to. In this respect, we illustrated the needs of having
enhanced ways of supporting human-machine interactions and demonstrated them
by means of a small home automation example. We noticed that in general
concrete syntaxes usually refer to traditional texts writing and diagrams drawing,
while more complex forms of interaction are largely neglected. Therefore, we
proposed to extend the current concrete syntax definition approaches by adding
sounds and gestures, but also possibilities to compose them with traditional
interaction modalities. In this respect, by adhering to the MPM methodology we
defined an appropriate modeling language for illustrating concrete syntaxes that
can be exploited later on to generate corresponding support for implementing
the specified interaction modalities.</p>
      <p>As next steps we plan to extend the Eugenia concrete syntax engine in
order to be able to automatically generate the support for the extended
humanmachine interactions declared through the proposed language. This phase will
also help in the validation of the proposed concrete syntax metamodel shown in
Fig. 3. In particular, we aim at verifying the adequacy of the expressive power
provided by the language and extend it with additional interaction means.</p>
      <p>This work constitutes the base to build-up advanced modeling tools relying
on enhanced forms of interaction. Such improvements could be remarkably
important to widen tools accessibility to disabled developers as well as to reduce
the accidental complexity of dealing with big models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Bilmes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Malkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kilanski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kirchho↵</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Subramanya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Harada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Landay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dowden</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Chizeck</surname>
          </string-name>
          , “
          <article-title>The vocal joystick: A voice-based human-computer interface for individuals with motor impairments,” in HLT/EMNLP</article-title>
          . The Association for Computational Linguistics,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>A.</given-names>
            <surname>Temko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Malkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zieger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Macho</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Nadeu</surname>
          </string-name>
          , “
          <article-title>Acoustic event detection and classification in smart-room environments: Evaluation of chil project systems</article-title>
          ,” Cough, vol.
          <volume>65</volume>
          , p.
          <fpage>6</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>S.</given-names>
            <surname>Nakamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hiyane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Asano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nishiura</surname>
          </string-name>
          , and T. Yamada, “
          <article-title>Acoustical sound database in real environments for sound scene understanding and hands-free speech recognition,” in LREC</article-title>
          .
          <source>European Language Resources Association</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>D.</given-names>
            <surname>Navarre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Palanque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bastide</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Schyn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Winckler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Nedel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Freitas</surname>
          </string-name>
          , “
          <article-title>A formal description of multimodal interaction techniques for immersive virtual reality applications,” in INTERACT, ser</article-title>
          . Lecture Notes in Computer Science,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Costabile</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Paterno</surname>
          </string-name>
          `, Eds., vol.
          <volume>3585</volume>
          . Springer,
          <year>2005</year>
          , pp.
          <fpage>170</fpage>
          -
          <lpage>183</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yuan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , “
          <article-title>Robust hand gesture recognition with kinect sensor,”</article-title>
          <source>in ACM Multimedia</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>759</fpage>
          -
          <lpage>760</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>S.</given-names>
            <surname>Mitra</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Acharya</surname>
          </string-name>
          , “
          <article-title>Gesture recognition: A survey,”</article-title>
          <source>IEEE Trans. on Systems, Man and Cybernetics</source>
          - part C, vol.
          <volume>37</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>324</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ohya</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ishii</surname>
          </string-name>
          , “
          <article-title>Recognizing human action in time-sequential images using hidden Markov model,” in Proceed</article-title>
          . IEEE Conf.
          <source>Computer Vision and Pattern Recognition</source>
          ,
          <year>1992</year>
          , pp.
          <fpage>379</fpage>
          -
          <lpage>385</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>M.</given-names>
            <surname>Isard</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Blake</surname>
          </string-name>
          , “
          <article-title>Condensation - conditional density propagation for visual tracking</article-title>
          ,”
          <source>International Journal of Computer Vision</source>
          , vol.
          <volume>29</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>5</fpage>
          -
          <lpage>28</lpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>P.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Huang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Turk</surname>
          </string-name>
          , “
          <article-title>Gesture modeling and recognition using finite state machines,” in FG</article-title>
          .
          <source>IEEE Computer Society</source>
          ,
          <year>2000</year>
          , pp.
          <fpage>410</fpage>
          -
          <lpage>415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10. R. Bolt, “
          <article-title>Put-that-there: Voice and gesture at the graphics interface,” in Proceed. 7th annual conference on Computer graphics and interactive techniques, ser</article-title>
          .
          <source>SIGGRAPH '80</source>
          . New York, NY, USA: ACM,
          <year>1980</year>
          , pp.
          <fpage>262</fpage>
          -
          <lpage>270</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>B.</given-names>
            <surname>Demiroz</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Ar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ronzhin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Coban</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yalcn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karpov</surname>
          </string-name>
          , and L. Akarun, “
          <article-title>Multimodal assisted living environment</article-title>
          ,
          <source>” in eNTERFACE 2011, The Summer Workshop on Multimodal Interfaces</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>N.</given-names>
            <surname>Osawa</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y. Y.</given-names>
            <surname>Sugimoto</surname>
          </string-name>
          , “
          <article-title>Multimodal text input in an immersive environment</article-title>
          ,”
          <source>in ICAT 2002, 12th International Conference on Articial Reality and Telexistence</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>85</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13. K. Vertanen, “
          <article-title>E cient computer interfaces using continuous gestures, language models, and speech</article-title>
          ,” http://www.cl.cam.ac.uk/TechReports/UCAM-CLTR-
          <volume>627</volume>
          .pdf, Computer Laboratory, University of Cambridge,
          <source>Tech. Rep. UCAMCL-TR-627</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>D.</surname>
            Huggins-Daines and
            <given-names>A. I. Rudnicky</given-names>
          </string-name>
          , “
          <article-title>Interactive asr error correction for touchscreen devices,” in ACL (Demo Papers)</article-title>
          .
          <source>The Association for Computer Linguistics</source>
          ,
          <year>2008</year>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>P. O.</given-names>
            <surname>Kristensson</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Vertanen</surname>
          </string-name>
          , “
          <article-title>Asynchronous multimodal text entry using speech and gesture keyboards,” in INTERSPEECH</article-title>
          . ISCA,
          <year>2011</year>
          , pp.
          <fpage>581</fpage>
          -
          <lpage>584</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. L.
          <string-name>
            <surname>Hoste</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Dumas</surname>
            , and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Signer</surname>
          </string-name>
          , “
          <article-title>Speeg: a multimodal speech- and gesturebased text input solution</article-title>
          ,” in AVI, G. Tortora,
          <string-name>
            <given-names>S.</given-names>
            <surname>Levialdi</surname>
          </string-name>
          , and M. Tucci, Eds. ACM,
          <year>2012</year>
          , pp.
          <fpage>156</fpage>
          -
          <lpage>163</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>D. J. Ward</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <string-name>
            <surname>Blackwell</surname>
          </string-name>
          , and
          <string-name>
            <surname>D. J. C. MacKay</surname>
          </string-name>
          , “
          <article-title>Dasher - a data entry interface using continuous gestures and language models</article-title>
          ,” in UIST,
          <year>2000</year>
          , pp.
          <fpage>129</fpage>
          -
          <lpage>137</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>D</surname>
          </string-name>
          . A. Carr, “
          <article-title>Specification of interface interaction objects,” in Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, ser</article-title>
          .
          <source>CHI '94</source>
          . New York, NY, USA: ACM,
          <year>1994</year>
          , pp.
          <fpage>372</fpage>
          -
          <lpage>378</lpage>
          . [Online]. Available: http://doi.acm.
          <source>org/10</source>
          .1145/191666.191793
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>S.</given-names>
            <surname>Smith</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Duke</surname>
          </string-name>
          , “
          <article-title>Virtual environments as hybrid systems</article-title>
          ,” in Proceedings of Eurographics UK 17th Annual
          <string-name>
            <surname>Conference (EG-UK99)</surname>
          </string-name>
          , E. U. K. Chapter, Ed.,
          <string-name>
            <surname>United</surname>
            <given-names>Kingdom</given-names>
          </string-name>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Palanque</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Bastide</surname>
          </string-name>
          , “
          <article-title>Petri net based design of user-driven interfaces using the interactive cooperative objects formalism,” in DSV-</article-title>
          IS,
          <year>1994</year>
          , pp.
          <fpage>383</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>R.</given-names>
            <surname>Deshayes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mens</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Palanque</surname>
          </string-name>
          , “
          <article-title>A generic framework for executable gestural interaction models,”</article-title>
          <source>in Proc. VL/HCC</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Obrenovic</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Starcevic</surname>
          </string-name>
          , “
          <article-title>Modeling multimodal human-computer interaction</article-title>
          ,” IEEE Computer, vol.
          <volume>37</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>