<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshops, Los
Angeles, USA, March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Learn2Sign: Explainable AI for Sign Language Learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Prajwal Paudyal</string-name>
          <email>ppaudyal@asu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junghyo Lee</string-name>
          <email>jlee375@asu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Azamat Kamzin</string-name>
          <email>akamzin@asu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamad Soudki</string-name>
          <email>msoudki@asu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ayan Banerjee</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandeep K.S. Gupta</string-name>
          <email>sandeep.gupta@asu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Arizona State University Tempe</institution>
          ,
          <addr-line>Arizona</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Explainable AI; Sign Language Learning; Computer-aided learning</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>20</volume>
      <issue>2019</issue>
      <abstract>
        <p>Languages are best learned in immersive environments with rich feedback. This is specially true for signed languages due to their visual and poly-componential nature. Computer Aided Language Learning (CALL) solutions successfully incorporate feedback for spoken languages, but no such solution exists for signed languages. Current Sign Language Recognition (SLR) systems are not interpretable and hence not applicable to provide feedback to learners. In this work, we propose a modular and explainable machine learning system that is able to provide fine-grained feedback on location, movement and hand-shape to learners of ASL. In addition, we also propose a waterfall architecture for combining the sub-modules to prevent cognitive overload for learners and to reduce computation time for feedback. The system has an overall test accuracy of 87.9 % on real-world data consisting of 25 signs with 3 repetitions each from 100 learners.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Human-centered computing → Interaction design; •
Computing methodologies → Artificial intelligence ; • Applied
computing → Interactive learning environments.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        Signed languages are natural mediums of communication for the
estimated 466 million deaf or hard of hearing people worldwide [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
Families and friends of the deaf can also benefit from being able
to sign. The Modern Language Association [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] reports that the
enrollment in American Sign Language (ASL) courses in the U.S.
has increased nearly 6,000 percent since 1990 which shows that
interest to acquire sign languages is increasing. However, the lack
of resources for self-paced learning makes it dificult to acquire,
specially outside of the traditional classroom setting [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
      </p>
      <p>
        The ideal environment for language learning is immersion with
rich feedback [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] and this is specially true for sign languages [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
IUI Workshops’19, March 20, 2019, Los Angeles, USA
© 2019 Copyright for the individual papers by the papers’ authors. Copying permitted
for private and academic purposes. This volume is published and copyrighted by its
editors.
      </p>
      <p>
        Extended studies have shown that providing item-based feedback in
CALL systems is very important [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. Towards this goal, extensive
language learning softwares for spoken languages such as Rosetta
Stone or Duolingo support some form of assessments and automatic
feedback [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. Although, there are numerous instructive books [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ],
video tutorials or smartphone applications for learning popular
sign languages, there hasn’t yet been any work towards providing
automatic feedback as seen in Table 1. We conducted a survey [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]
of 52 first-time ASL users (29M, 21F) in 2018 and 96.2 % said that
reasonable feedback is important but lacking in solutions for sign
language learning (Table 2).
      </p>
      <p>Application
ASL Coach</p>
      <p>The ASL App
ASL Fingerspelling</p>
      <p>Marlee Signs
SL for Beginners</p>
      <p>WeSign</p>
      <p>
        Studies show that elaborated feedback such as providing
meaningful explanations and examples produce larger efect on learning
outcomes than just feedback regarding the correctness [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. The
simplest feedback that can be given to a learner is whether their
execution of a particular sign was correct. State-of-the-art SLR and
activity recognition systems can be easily trained to accomplish
this. However, to truly help a learner identify mistakes and learn
from them, the feedback and explanations generated must be more
ifne-grained.
      </p>
      <p>
        The various ways in which a signer can make mistakes during
the execution of a sign can be directly linked to how minimum pairs
are formed in the phonetics of that language. The work of Stokoe
postulates that the manual portion of an ASL sign is composed of
1) location, 2) movement and 3) hand-shape and orientation [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. A
black box recognition system cannot provide this level of feedback,
thus there is the need for an explainable AI system because feedback
from the system is analogous to explanations for its final decision.
Non-manual markers such as facial expressions and body gaits
also change the meaning of signs to some extent but they are less
important for beginner level language acquisition, so these will be
considered for future work.
      </p>
      <p>
        Studies have also shown that the efect of feedback is highest
if provided immediately [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ], thus feedback systems should be
real-time. The requirement for immediate feedback also restricts
the usage of complicated learning algorithms that require heavy
computing [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and extensive training. The usability and usefulness
of applications is enhanced if learning is self-paced, learners are
allowed to use their own devices, and the learning vocabulary can
be easily extended. However, current solutions for SLR require
retraining to support unseen words and large datasets initially. To
solve these challenges, we designed Learn2Sign(L2S), a smartphone
application that utilizes explainable AI to provide fine-grained
feedback on location, movement, orientation and hand-shape for ASL
learners. L2S is built using a waterfall combination of three
nonparametric models as seen in Figure1 to ensure extendibility to new
vocabulary. Learners can use L2S with any smartphone or
computer with a front-facing camera. L2S utilizes a bone localization
technique proposed by [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] for movement and location based
feedback and a light-weight pre-trained Convolutional Neural Network
(CNN) as a feature extractor for hand-shape feedback.
      </p>
      <p>
        The methodology and evaluations are provided in Sections 3
and 4. As part of the work, we collected video data from 100 users
executing 25 ASL signs three times each. The videos were recorded
by L2S users in real-world settings without restrictions on
devicetype, lighting conditions, distance to the camera or recording pose
(sitting or standing up). This was to ensure generalization to
realworld conditions, however, this makes the dataset more challenging.
More details about the resulting dataset of about 7500 instances
can be found in [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        There have been many works on providing meaningful feedback for
spoken language learners [
        <xref ref-type="bibr" rid="ref21 ref22 ref8">8, 21, 22</xref>
        ]. On the practical side, Rosetta
Stone provides both waveform and spectrograph feedback for
pronunciation mistakes by comparing acoustic waves of a learner to
that of a native speaker [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]. There has also been some recent work
on design principles for using Automatic Speech Recognition (ASR)
techniques to provide feedback for language learners [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. Sign
Language Recognition (SLR) is a research field that closely mirrors
ASR and can potentially be utilized by systems for sign language
learning. However, to the best of our knowledge, no such system
exists. This can be explained by the inherent dificulties in SLR as
well as the lack of detailed studies on design principles for such
systems. In this work, we propose some design principles and an
explainable smart system to meet this goal.
      </p>
      <p>
        Continuously translating a video recording of a signed language
to a spoken language is a very challenging problem and has been
(a) TIGER mostly in bucket (b) DECIDE in buckets 3 and
1 for left-hand. 6 for right-hand.
tackled recently by various researchers with some success [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. For
the purposes of this application, such complex measures are not
desirable, as they mandate extensive datasets for training and large
models for translation which decreases their usability. Isolated
Sign Language Recognition has the goal of classifying various sign
tokens into classes that represent some spoken language words [
        <xref ref-type="bibr" rid="ref11 ref12 ref18 ref29">11,
12, 18, 29</xref>
        ]. Some researchers have utilized videos [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] while some
others have attempted to use wearable sensors [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ] with varying
performances. In this work, we utilize the insights and advances
from such systems to help a new learner acquire the sign language
words. To our knowledge, this work is the first attempt at such a
practical and much needed application.
      </p>
      <p>
        For this work, we require an estimation of human pose,
specifically the estimates on the location of various joints throughout a
video, known as keypoints. There have been several works towards
this goal [
        <xref ref-type="bibr" rid="ref25 ref3 ref32 ref33 ref4 ref5">3–5, 25, 32, 33</xref>
        ]. Some of these works first detect the
keypoints in 2D and then attempt to ‘lift’ that set to 3D space while
others return the 2D coordinates of the various keypoints relative
to the image. In order to fulfill the requirement to use pervasive
cameras, we did not focus on the approaches that utilize depth
information such as Microsoft Kinect [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Thus, we utilized the
pose estimates from a Tensorflow JS implementation of a model
proposed by Papandreou et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] which can run on devices with
or without GPUs (Graphical Processing Units).
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>METHODOLOGY</title>
      <p>
        Stokoe proposed that a sign in ASL consists of three parts which
combine simultaneously: the tab (location of the sign), the dez
(handshape) and the sig (movement) [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. Signs like ‘HEADACHE’ and
‘STOMACH ACHE’ that are similar in hand-shape and movement
may difer only by the signing location. Similarly, there will be
other minimal pairs of signs that difer only by the movement or
hand-shape. Following this understanding, L2S is composed of three
corresponding recognition and feedback modules.
3.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>User Interface</title>
      <p>
        For initial data collection and for testing the UI, we developed an
android application called L2S. We preloaded the application with
25 tutorial videos from Signing Savvy corresponding to 25 ASL
signs [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The application has three main components: a) Learning
Module b) Practice Module, and c) Extension.
3.1.1 Learning Module. The learning module of the L2S application
is where all the tutorial videos are accessible. A learner selects an
ASL word/phrase to learn and can then view the tutorial videos.
The learner can pause, play, and repeat the tutorials as many times
as needed. In this module, the learner can also record executions of
their signs for self-assessment.
3.1.2 Practice Module. The practice module is designed to give
automatic feedback to the learners. A learner selects a sign to practice
and sets up their device to record their execution. After this, L2S
determines if the learner performed the sign correctly. The result
is correct if the sign meets the thresholds for movement, location,
and hand-shape and a ‘correct’ feedback is given. If, the system
determines that the learner did not execute the sign correctly, an
appropriate feedback is provided as seen in Figure 1. Details about the
recognition and feedback mechanisms is discussed in Section 3.4.
3.1.3 Extension Module. To extend the supported vocabulary of
L2S, a learner can upload one or more tutorial videos from a source
of their choosing. The application processes them for usability
before they appear in the Learning Module as new tutorial sign(s).
3.2
      </p>
    </sec>
    <sec id="sec-6">
      <title>Data Collection</title>
      <p>We collected signing videos from 100 learners, for 25 ASL signs
with three repetitions each in real-world settings using L2S app.
Learners used their own devices, with no restrictions on lighting
conditions, distance to the camera or recording pose (sitting or
standing up). After reviewing a tutorial video, a learner was given
a 5 s setup time before recording a 3 s video using a front-facing
camera. Both the tutorial and the newly recorded video were then
displayed in the same screen for the user to accept or reject. This
self-assessment served not only as a review but it also helped prune
incorrect data due to device or timing errors as suggested by the
new learner survey in Table 2.
3.3</p>
    </sec>
    <sec id="sec-7">
      <title>Preprocessing</title>
      <p>
        Determining joint locations: Since, diferent devices record in
diferent resolutions, all videos for learning, practice or extension
are first converted to a 320*240 resolution. Then, PoseNet Javascript
API for single pose estimation [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] was used to compute the
estimated locations and confidence levels for the various keypoints as
seen in Table 3. Figure 2 shows the estimated eyes, shoulder and
wrist locations for the signs TIGER and DECIDE for all the frames
in one video.
      </p>
      <p>
        Normalization: There is a diference in scale of the bodies relative
to the frame-size corresponding to the distance between the learner
and the camera. This scaling factor can negatively impact
recognition since the relative location, movement and hand-shape will
vary with distance. We perform min-max normalization and
zeroing based on the distance between the average estimated locations
for the right and left shoulders throughout the video frame as
suggested by [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Normalization was found to be specially important
for correct movement recognition.
3.4
      </p>
    </sec>
    <sec id="sec-8">
      <title>Recognition and Feedback</title>
      <p>L2S is designed to give incremental feedback to learners for the
various modalities in sign language: a) Location b) Movement and
c) Hand-shape. The various models are arranged in a waterfall
architecture as seen in Figure 1. If the location of signing was not
correct, then immediate feedback is provided and the learner is
prompted to try again. Similarly, if the movement of the elbows or
the wrists for either hand was incorrect, the learner is prompted
to try again. Finally, if the shape and orientation of either of the
hands does not appear to be correct, a hand-shape based feedback
is provided. Consequently, the learner can move on to a practice
a new sign, only if all these modalities were suficiently correct.
A waterfall architecture was chosen in the final application over
a linear weighted combination to make learning progressive and
to decrease the cognitive load on the learner due to the potential
of mistakes in multiple modalities. This architecture also helps
to reduce the time taken for recognition and feedback since the
models are stacked in an increasing order of execution time. Each
of the feedback screens shown to the user also has a link to the
tutorial video. Users can also manually tune the amount of feedback
by altering the value of ‘feedback sensitivity’ in the application
settings. Increasing this value alters the thresholds for each of the
sub-modules so that the overall rate of feedback is increased. This
involves a trade-of in performance which is summarized in Figure 3.
3.5</p>
    </sec>
    <sec id="sec-9">
      <title>Location</title>
      <p>To correctly and eficiently determine the location for signing, we
ifrst assume the shoulders stay fairly stationary throughout the
execution of a sign. This is a fair assumption for ASL since there are
no minimal pairs exclusively associated with a signer’s shoulders.
Then we divide the video canvas into 6 diferent sub-sections called
buckets as seen in Figure 2. Then, as the learner executes any given
sign, the location of both the wrist joints is tracked for each bucket
resulting in a vector of length 6.</p>
      <p>This same procedure is followed for the tutorials, and a
cosinebased comparison between is done between the two vectors. A
heuristic threshold that is determined during training is utilized
as a cut-of point. If the resulting cosine similarity is lower than a
threshold, some feedback is shown to the learner as seen in Figure 4.
For each hand, the user’s own video is replayed in Graphics
Interchange Format (GIF) with a red highlight on the location section
that was incorrect and a green highlight on the section of the frame
where the sign should have been executed. A text feedback with
details and a link to the tutorial is also provided and the learner is
prompted to try again.
3.6</p>
    </sec>
    <sec id="sec-10">
      <title>Movement</title>
      <p>
        Determination of correct movement is perhaps the single most
important feedback we can provide to a learner. We compute a
segmental DTW distance between a learner and the tutorial
using keypoints for the wrists, elbows and shoulders as suggested
in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Normalization as discussed in Section 3.3 was found to be
very important. Experimental results showed that segmental DTW
outperformed DTW or Global Alignment Kernel (GAK).
      </p>
      <p>The dataset had a wide variation in the number of frames per
video. It was found that this afected the distance scores adversely.
Thus, as an additional step of preprocessing, the video with the
higher number of frames was down-sampled before comparison and
the segmental DTW is utilized to find the best sub-sample matching.
Thresholds for the signs were determined experimentally using 10
training videos for each sign. If segmental DTW distance between
a learner’s recording and a tutorial was higher than the threshold
for each arm section, then a movement-based feedback is provided
as seen in Figure 1. A GIF is replayed to the user with the section(s)
of the arm for which the movement was incorrect in red as seen in
Figure 5b. A textual feedback is also generated with an explanation
after which the user is prompted to watch the tutorial and try again.
3.7</p>
    </sec>
    <sec id="sec-11">
      <title>Hand Shape and Orientation</title>
      <p>ASL signs which are otherwise similar, may difer only by the
shape or orientation of the hands. Since, CNNs have
state-of-theart image recognition results, we utilized Inception v3 or Mobilenet
CNN depending on the device being used. A model that was
pretrained on ImageNet is retrained using hand-shape images from the
training users. The wrist location obtained during pre-processing
was used as a guide to auto-crop these hand-shape images. During
recognition time, hand-shape images from each hand are extracted
automatically in a similar way from a learner’s recording. Then 6
images for each hand are passed separately through the CNN and
the softmax layer is obtained and are concatenated together as seen
in Figure 1. Similar processing is done on the tutorial video to obtain
a vector of the same length. Then a cosine similarity is calculated
on the resultant vector. If the similarity between a learner’s sign
and that of a tutorial is above a set threshold for a sign, then the
execution is determined to be correct, otherwise the hand-shape
based feedback as seen in Figure 5 is provided.</p>
      <p>Although the retrained CNN could theoretically be used as a
classifier, we use it only as a feature extractor for cosine similarity
to ensure that the system can extend to unseen classes. A new
tutorial can then be efectively added to the system without the need
for retraining. An analysis of the efectiveness of hand-shape and
orientation recognizer is provided in Section 4.Similar to location
and movement, feedback for hand shape and orientation is also
provided in the form of a replay GIF and text. A zoomed in image
of the incorrect hand shape is shown side by side with the correct
image from a tutorial as seen in Figure 5(a).
4</p>
    </sec>
    <sec id="sec-12">
      <title>RESULTS AND EVALUATION</title>
      <p>An ideal system should give feedback to a learner only if their
execution is incorrect. Giving unnecessary feedback for correct
executions will hinder the learning process and decrease the usability.
Conversely, providing sound and timely explanations for incorrect
executions helps to improve utility and user trust. Smart systems
such as L2S that use explainable machine learning tend to have a
trade-of between explainability and performance which should be
minimized.</p>
      <p>
        The overall performance of the system was tested for 10 test
users for a total of 750 signs. The training of the CNN for
handshape feature extraction and optimal threshold determination was
done using the remaining users. For each sign, 30 executions from
the test dataset were taken as true class while 30 randomly
selected executions from the pool of remaining signs was taken as
(a) Hand-shape feedback for AFTER. (b) Movement Feedback for ABOUT.
incorrect class to avoid class imbalance. A pre-trained model from
C3D [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] was retrained with the data we collected and was used as
the baseline for comparison. This model has an accuracy of 82.3 %
on UCF101 [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] and 87.7 % in YUPENN-Scene [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] datasets. The final
recognition accuracy of C3D on L2S dataset using the same
traintest split was 45.38 %. Our approach achieves a higher accuracy of
87.9 % while still ofering explanations about its decisions in the
form of learner feedback.
      </p>
      <p>To obtain the results, data collected from one learner was selected
at random and served as the tutorial dataset. Then each sign for each
user in the test dataset was compared against the corresponding
tutorial sign. The location module had an overall recall of 96.4 % and
precision of 24.3 %. The lower precision is due to the fact that many
signs in the test dataset had similar locations. We performed a test
comparing only the signs ‘LARGE’ to the sign ‘FATHER’ and both
the precision and recall were 100 %. The movement module had an
overall recall of 93.2 % and a precision of 52.4 %. The hand-shape
module had a recall of 89 % and a precision of 74 %. The overall
model is constructed as a waterfall combination of all three models
such that the movement model is executed only when the location
was found to be correct, and the hand-shape model is executed only
when both the location and movement were correct. The overall
precision, recall, f-1 score and accuracies is summarized in Table 4.
5</p>
    </sec>
    <sec id="sec-13">
      <title>DISCUSSION AND FUTURE WORK</title>
      <p>
        We demonstrated the need for a feedback based technological
solution for sign language learning and provided an implementation
with a modular feedback mechanism. The user preference for the
desired amount of feedback can be changed by altering the value for
‘Feedback Sensitivity’. The trade-of between ‘Feedback Sensitivity’
and amount of feedback received as well as other performance
metrics is summarized in Figure 3. Although, we designed our feedback
mechanism based on principles from linguistics and user survey,
only a large scale usage of such an application will provide
definitive best practices for the most efective feedback. In such future
studies, issues such as the extent of user control for determining
types of feedback and the possibility of peer-to-peer feedback for
on-line learning has to be evaluated as suggested by works such
as [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. This work provides the foundations and feasibility for
interactive and intelligent sign language learning to pave the path
for such future work.
      </p>
      <p>We collected usage and interaction data from 100 new learners
as part of this work, which will be foundational to assist future
researchers. Although, the focus of this work was on the
manual portion of sign languages, the preprocessing includes location
estimates for the eyes, ears and the nose. This can be utilized for
including facial expression recognition and feedback in future works.
We evaluated only 25 isolated words for ASL, but in the future,
this work can be extended to more words and phrases and to
include other sign languages since the general principles will remain
the same. In this work, we used sign language as a test
application, however, the insights from this work can be easily applied to
other gesture domains such as combat sign training for military or
industrial operator signs.
6</p>
    </sec>
    <sec id="sec-14">
      <title>CONCLUSION</title>
      <p>
        There is an increasing need and demand for learning sign language.
Feedback is very important for language learning and intelligent
language learning softwares must provide efective and meaningful
feedback. There has also been significant advances in research for
recognizing sign languages, however technological solutions that
leverage them to provide intelligent learning environments do not
exist. In this work, we identify diferent types of potential feedback
we can provide to learners of sign language and address some
challenges in doing so. We propose a pipeline of three non-parametric
recognition modules and an incremental feedback mechanism to
facilitate learning. We tested our system on real-world data from a
variety of devices and settings to achieve a final recognition
accuracy of 87.9 %. This demonstrates that using explainable machine
learning for gesture learning is desirable and efective. We also
provided diferent types of feedback mechanisms based on results
of a user survey and best practices in implementing them. Finally,
we collected data from 100 users of L2S with 3 repetitions for each
of the 25 signs for a total of 7500 instances [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
7
      </p>
    </sec>
    <sec id="sec-15">
      <title>ACKNOWLEDGMENTS</title>
      <p>
        We thank SigningSavvy[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] for letting us use their tutorial videos
in the application.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Anguera</surname>
          </string-name>
          , Robert Macrae, and
          <string-name>
            <given-names>Nuria</given-names>
            <surname>Oliver</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Partial sequence matching using an unbounded dynamic time warping algorithm</article-title>
          .
          <source>In Acoustics Speech and Signal Processing (ICASSP)</source>
          ,
          <source>2010 IEEE International Conference on. IEEE</source>
          ,
          <fpage>3582</fpage>
          -
          <lpage>3585</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Modern</given-names>
            <surname>Language Association</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Language Enrollment Database</article-title>
          . https: //apps.mla.org/flsurvey_search. [Online; accessed 24-September-2018].
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Federica</given-names>
            <surname>Bogo</surname>
          </string-name>
          , Angjoo Kanazawa, Christoph Lassner,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Gehler</surname>
          </string-name>
          , Javier Romero, and
          <string-name>
            <surname>Michael</surname>
          </string-name>
          J Black.
          <year>2016</year>
          .
          <article-title>Keep it SMPL: Automatic estimation of 3D human pose and shape from a single image</article-title>
          .
          <source>In European Conference on Computer Vision</source>
          . Springer,
          <fpage>561</fpage>
          -
          <lpage>578</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Ching-Hang Chen</surname>
            and
            <given-names>Deva</given-names>
          </string-name>
          <string-name>
            <surname>Ramanan</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>3d human pose estimation= 2d pose estimation+ matching</article-title>
          .
          <source>In CVPR</source>
          , Vol.
          <volume>2</volume>
          . 6.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Xianjie</given-names>
            <surname>Chen and Alan L Yuille</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Articulated pose estimation by a graphical model with image dependent pairwise relations</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>1736</volume>
          -
          <fpage>1744</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Necati</given-names>
            <surname>Cihan</surname>
          </string-name>
          <string-name>
            <surname>Camgoz</surname>
          </string-name>
          , Simon Hadfield, Oscar Koller, Hermann Ney, and Richard Bowden.
          <year>2018</year>
          .
          <article-title>Neural Sign Language Translation</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>7784</fpage>
          -
          <lpage>7793</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Konstantinos</surname>
            <given-names>G Derpanis</given-names>
          </string-name>
          ,
          <article-title>Matthieu Lecce</article-title>
          , Kostas Daniilidis, and Richard P Wildes.
          <year>2012</year>
          .
          <article-title>Dynamic scene understanding: The role of orientation features in space and time in scene classification</article-title>
          .
          <source>In Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <source>2012 IEEE Conference on. IEEE</source>
          ,
          <fpage>1306</fpage>
          -
          <lpage>1313</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Farzad</given-names>
            <surname>Ehsani</surname>
          </string-name>
          and
          <string-name>
            <given-names>Eva</given-names>
            <surname>Knodt</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Speech technology in computer-aided language learning: Strengths and limitations of a new CALL paradigm</article-title>
          . (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Karen</given-names>
            <surname>Emmorey</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Language, cognition, and the brain: Insights from sign language research</article-title>
          . Psychology Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Rebecca</surname>
            <given-names>Fiebrink</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perry R Cook</surname>
            , and
            <given-names>Dan</given-names>
          </string-name>
          <string-name>
            <surname>Trueman</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Human model evaluation in interactive supervised learning</article-title>
          .
          <source>In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>147</volume>
          -
          <fpage>156</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Kirsti</given-names>
            <surname>Grobel</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marcell</given-names>
            <surname>Assan</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Isolated sign language recognition using hidden Markov models</article-title>
          .
          <source>In Systems, Man, and Cybernetics</source>
          ,
          <year>1997</year>
          .
          <string-name>
            <given-names>Computational</given-names>
            <surname>Cybernetics</surname>
          </string-name>
          and Simulation.,
          <source>1997 IEEE International Conference on</source>
          , Vol.
          <volume>1</volume>
          . IEEE,
          <fpage>162</fpage>
          -
          <lpage>167</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Pradeep</surname>
            <given-names>Kumar</given-names>
          </string-name>
          , Himaanshu Gauba, Partha Pratim Roy, and Debi Prosad Dogra.
          <year>2017</year>
          .
          <article-title>Coupled HMM-based multi-sensor data fusion for sign language recognition</article-title>
          .
          <source>Pattern Recognition Letters</source>
          <volume>86</volume>
          (
          <year>2017</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Impact</given-names>
            <surname>Lab</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Learn2Sign Details Page</article-title>
          . https://impact.asu.edu/projects/ sign
          <article-title>-language-recognition/learn2sign</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Kian</given-names>
            <surname>Ming</surname>
          </string-name>
          <string-name>
            <surname>Lim</surname>
          </string-name>
          ,
          <source>Alan WC Tan, and Shing Chiang Tan</source>
          .
          <year>2016</year>
          .
          <article-title>A feature covariance matrix with serial particle filter for isolated sign language recognition</article-title>
          .
          <source>Expert Systems with Applications</source>
          <volume>54</volume>
          (
          <year>2016</year>
          ),
          <fpage>208</fpage>
          -
          <lpage>218</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Malek</surname>
            <given-names>Nadil</given-names>
          </string-name>
          , Feryel Souami, Abdenour Labed, and
          <string-name>
            <given-names>Hichem</given-names>
            <surname>Sahbi</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>KCCAbased technique for profile face identification</article-title>
          .
          <source>EURASIP Journal on Image and Video Processing</source>
          <year>2017</year>
          ,
          <volume>1</volume>
          (
          <year>2016</year>
          ),
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16] World Health Organization.
          <year>2018</year>
          .
          <article-title>Deafness and hearing loss</article-title>
          . http://www.who. int/news-room/fact-sheets/detail/deafness-and
          <string-name>
            <surname>-</surname>
          </string-name>
          hearing-loss. [Online; accessed 24-September-2018].
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>George</surname>
            <given-names>Papandreou</given-names>
          </string-name>
          , Tyler Zhu, Nori Kanazawa, Alexander Toshev, Jonathan Tompson, Chris Bregler, and
          <string-name>
            <given-names>Kevin</given-names>
            <surname>Murphy</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards accurate multi-person pose estimation in the wild</article-title>
          .
          <source>In CVPR</source>
          , Vol.
          <volume>3</volume>
          . 6.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Prajwal</surname>
            <given-names>Paudyal</given-names>
          </string-name>
          ,
          <source>Ayan Banerjee, and Sandeep KS Gupta</source>
          .
          <year>2016</year>
          .
          <article-title>Sceptre: a pervasive, non-invasive, and programmable gesture recognition technology</article-title>
          .
          <source>In Proceedings of the 21st International Conference on Intelligent User Interfaces. ACM</source>
          ,
          <volume>282</volume>
          -
          <fpage>293</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Prajwal</surname>
            <given-names>Paudyal</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Junghyo</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <source>Ayan Banerjee, and Sandeep KS Gupta</source>
          .
          <year>2017</year>
          .
          <article-title>Dyfav: Dynamic feature selection and voting for real-time recognition of fingerspelled alphabet using wearables</article-title>
          .
          <source>In Proceedings of the 22nd International Conference on Intelligent User Interfaces. ACM</source>
          ,
          <volume>457</volume>
          -
          <fpage>467</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Fabrizio</surname>
            <given-names>Pedersoli</given-names>
          </string-name>
          , Sergio Benini, Nicola Adami, and
          <string-name>
            <given-names>Riccardo</given-names>
            <surname>Leonardi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>XKin: an open source framework for hand pose and gesture recognition using kinect</article-title>
          .
          <source>The Visual Computer</source>
          <volume>30</volume>
          ,
          <issue>10</issue>
          (
          <year>2014</year>
          ),
          <fpage>1107</fpage>
          -
          <lpage>1122</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Martha</surname>
            <given-names>C</given-names>
          </string-name>
          <string-name>
            <surname>Pennington and Pamela</surname>
          </string-name>
          Rogerson-Revell.
          <year>2019</year>
          .
          <article-title>Using Technology for Pronunciation Teaching, Learning, and Assessment</article-title>
          .
          <source>In English Pronunciation Teaching and Research</source>
          . Springer,
          <fpage>235</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Sean</surname>
            <given-names>Robertson</given-names>
          </string-name>
          , Cosmin Munteanu, and
          <string-name>
            <given-names>Gerald</given-names>
            <surname>Penn</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Designing Pronunciation Learning Tools: The Case for Interactivity against Over-Engineering</article-title>
          .
          <source>In Proceedings of the 2018 CHI Conference on Human Factors in Computing Systems. ACM</source>
          ,
          <volume>356</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Russell</surname>
            <given-names>S</given-names>
          </string-name>
          <string-name>
            <surname>Rosen</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>American sign language curricula: A review</article-title>
          .
          <source>Sign Language Studies</source>
          <volume>10</volume>
          ,
          <issue>3</issue>
          (
          <year>2010</year>
          ),
          <fpage>348</fpage>
          -
          <lpage>381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Signing</given-names>
            <surname>Saavy</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Signing Saavy: Your Sign Language Resouce</article-title>
          . https://www. signingsavvy.com/. [Online; accessed 28-September-2018].
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Nikolaos</surname>
            <given-names>Sarafianos</given-names>
          </string-name>
          , Bogdan Boteanu,
          <source>Bogdan Ionescu, and Ioannis A Kakadiaris</source>
          .
          <year>2016</year>
          .
          <article-title>3d human pose estimation: A review of the literature and analysis of covariates</article-title>
          .
          <source>Computer Vision and Image Understanding</source>
          <volume>152</volume>
          (
          <year>2016</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>YoungHee</given-names>
            <surname>Sheen</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Corrective feedback and learner uptake in communicative classrooms across instructional settings</article-title>
          .
          <source>Language teaching research 8</source>
          ,
          <issue>3</issue>
          (
          <year>2004</year>
          ),
          <fpage>263</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Skehan</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>A cognitive approach to language learning</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Khurram</surname>
            <given-names>Soomro</given-names>
          </string-name>
          , Amir Roshan Zamir, and
          <string-name>
            <given-names>Mubarak</given-names>
            <surname>Shah</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>UCF101: A dataset of 101 human actions classes from videos in the wild</article-title>
          .
          <source>arXiv preprint arXiv:1212.0402</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Thad</surname>
            <given-names>Starner</given-names>
          </string-name>
          , Joshua Weaver, and
          <string-name>
            <given-names>Alex</given-names>
            <surname>Pentland</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>Real-time american sign language recognition using desk and wearable computer based video</article-title>
          .
          <source>IEEE Transactions on pattern analysis and machine intelligence</source>
          <volume>20</volume>
          , 12 (
          <year>1998</year>
          ),
          <fpage>1371</fpage>
          -
          <lpage>1375</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>William</surname>
            <given-names>C Stokoe</given-names>
          </string-name>
          <string-name>
            <surname>Jr</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Sign language structure: An outline of the visual communication systems of the American deaf</article-title>
          .
          <source>Journal of deaf studies and deaf education 10</source>
          ,
          <issue>1</issue>
          (
          <year>2005</year>
          ),
          <fpage>3</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Rosetta</given-names>
            <surname>Stone</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Talking back required</article-title>
          . https://www.rosettastone.com/ speech-recognition.
          <source>[Online; accessed 28-September-2018].</source>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Denis</surname>
            <given-names>Tome</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Russell</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Lourdes</given-names>
            <surname>Agapito</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Lifting from the deep: Convolutional 3d pose estimation from a single image</article-title>
          .
          <source>CVPR 2017 Proceedings</source>
          (
          <year>2017</year>
          ),
          <fpage>2500</fpage>
          -
          <lpage>2509</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Jonathan J Tompson</surname>
            , Arjun Jain, Yann LeCun, and
            <given-names>Christoph</given-names>
          </string-name>
          <string-name>
            <surname>Bregler</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Joint training of a convolutional network and a graphical model for human pose estimation</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          . 1799-
          <fpage>1807</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Du</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Lubomir Bourdev, Rob Fergus, Lorenzo Torresani, and
          <string-name>
            <given-names>Manohar</given-names>
            <surname>Paluri</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Learning spatiotemporal features with 3d convolutional networks</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          . 4489-
          <fpage>4497</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <surname>Fabienne M Van der Kleij</surname>
          </string-name>
          ,
          <source>Remco CW Feskens, and Theo JHM Eggen</source>
          .
          <year>2015</year>
          .
          <article-title>Efects of feedback in a computer-based learning environment on students' learning outcomes: A meta-analysis</article-title>
          .
          <source>Review of educational research 85</source>
          ,
          <issue>4</issue>
          (
          <year>2015</year>
          ),
          <fpage>475</fpage>
          -
          <lpage>511</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>Ping</surname>
            <given-names>Yu</given-names>
          </string-name>
          , Yingxin Pan,
          <string-name>
            <given-names>Chen</given-names>
            <surname>Li</surname>
          </string-name>
          , Zengxiu Zhang, Qin Shi, Wenpei Chu, Mingzhuo Liu, and
          <string-name>
            <given-names>Zhiting</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>User-centred design for Chinese-oriented spoken english learning system</article-title>
          .
          <source>Computer Assisted Language Learning</source>
          <volume>29</volume>
          ,
          <issue>5</issue>
          (
          <year>2016</year>
          ),
          <fpage>984</fpage>
          -
          <lpage>1000</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>