<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Alexandria poral scale convolutional neural network for micro-
Engineering Journal 59 (2020). doi:10.1016/j.aej. expression recognition</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1109/TPAMI.2014.2329301</article-id>
      <title-group>
        <article-title>unsupervised learning method for micro gesture recognition based on skeleton modality</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wenxuan Yuan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shanchuan He</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jianwen Dou</string-name>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Multi-Scale TCN, Unsupervised Network, Temporal Deconvolution, VAE structure</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>2</volume>
      <fpage>2466</fpage>
      <lpage>2482</lpage>
      <abstract>
        <p>We propose a novel unsupervised model for micro-gesture classification, called MSTCN-VAE, which follows the VAE structure by adding the Multi-scale TCN and hidden feature extraction block to the encoder, and the decoder is embedded with the Temporal Deconvolution block. The MSTCN-VAE model collects more temporal information from the input action sequences due to the advanced time series information integration method and thus exhibits better classification performance. By evaluation of the iMiGUE dataset, our approach outperforms the current state-of-the-art unsupervised methods in micro-gesture classification and is comparable to the accuracy of slightly earlier supervised models. Also, we validate the efectiveness of our model on the SMG dataset.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The recognition of human gestures and actions plays a
computer interaction [1][2][3] to video surveillance
[4][5][6] and robotics [7][8][9][10]. Over the years, there
has been significant progress in the field of
skeletonresentation of human body movements is utilized for
based action recognition [11][12], where the skeletal rep- raw data.
crucial role in various domains, ranging from human- and local temporal features [21]. Unsupervised methods
analyzing and understanding human gestures. Skeleton- are predominantly supervised [22][23], relying on labeled
representations or discover patterns from unlabeled or
weakly labeled data. Some articles explore techniques
such as hidden Markov models [19], sparse coding [20] ,
play a crucial role in scenarios where annotated training
data is scarce or unavailable, allowing for the discovery of
meaningful micro-gesture representations directly from</p>
      <p>Micro-gesture recognition networks in the mainstream
data for training. However, collecting micro-gesture
datasets presents challenges due to the dificulty of
capturing and annotating subtle hand movements. This
process often results in multiple labels for the same sample,
introducing ambiguity. To address the limitations of
supervised methods, researchers have explored
unsupervised approaches for micro-gesture datasets. One notable
method is Predict &amp; Cluster framework [24], which
provides a way to automatically recognize actions from
skeleing results on multiple benchmark datasets. Another one
is unsupervised S-VAE (U-S-VAE) [25], which indicates
the efectiveness of using multi-layer BLSTM to extract
recognition. However, both GRU (Gated Recurrent Unit)
[26] and LSTM (Long Short-Term Memory) [27] have
certain limitations because of the computational complexity
and limited time information integration when it comes
to efectively modeling long-term dependencies and
integrating time information. These limitations have led
to the development of alternative architectures like TCN
(Temporal Convolutional Network) [28], which benefits
based approaches [13][14][15] ofer a compact and
informative representation that captures the spatial and
temporal dynamics of human actions, enabling the
eficient processing of gestures and facilitating the
extraction of relevant features for recognition tasks. While
skeleton-based action recognition has achieved
remarkable success, there is a growing interest in exploring
micro-gesture recognition, which focuses on recognizing
subtle and fine-grained hand movements. Micro-gestures
poral variations, making them challenging to capture and
understand. To address this research frontier, many
methods have been proposed. In supervised micro-gesture</p>
      <sec id="sec-1-1">
        <title>On the other hand, unsupervised methods aim to learn</title>
        <p>China.
†These authors contributed equally.
nEvelop-O
and the ability to increase the receptive field
exponensupervised neural network, we can only use data that we
tially with depth to learn patterns over extended time
have labeled, and more unmarked data will be wasted.
horizons and the stability of gradient. These qualities</p>
      </sec>
      <sec id="sec-1-2">
        <title>The unsupervised method based on skeleton data has</title>
        <p>make TCN a promising alternative for tasks involving
come into view to conquer the aforementioned
probtime series data.
lems. Graph convolutional neural networks are widely
To address these issues, an innovative unsupervised
used in graph correlation recognition [34][35], but action
network model for gesture classification is proposed in
recognition depends on long-term information. Most of
this paper. The network is based on a VAE structure [29]
the frameworks are based on recurrent neural networks
using a temporal convolutional network (TCN) [28] and
(RNNs), convolutional neural networks (CNNs), or
grapha multiscale temporal convolutional network (MSTCN)
based CNNs. Employing methods directly tends to ignore
as the encoder and a TDCN (Temporal Deconvolutional
the most important information in action recognition,
Network) as the decoder. By conducting experiments
which includes the interrelationship and timing of the
on the original skeleton data as well as on the data after
movements. Diferent from the above framework, a novel
extraction of angular information, we demonstrate the
model-aware gesture-to-gesture translation method is
advantages of the model, such as label dependence on the
proposed, which presents novel approaches, called
Selfdataset and capturing the hidden feature vectors that are</p>
      </sec>
      <sec id="sec-1-3">
        <title>Attention Network (SAN) [36]. Furthermore, a Focal</title>
        <p>crucial for micro-gesture classification. Diferent variants
and Global Spatial-Temporal Transformer network
(FGof our network are evaluated on the now popular micro- STFormer)[37].
gesture dataset iMiGUE [25] and compared with
stateDiferent from the previous form of network
optiof-the-art supervised unsupervised methods. The wide
mization, a new unsupervised model [24] is based on
applicability of the network is validated on the SMG
an encoder-decoder system. The encoder is responsible
dataset [30]. The potential of our approach in advancing
for feature extraction from the original data to obtain a
skeleton-based micro-gesture recognition and further
feature vector that can be separated. The decoder needs
improving human-computer interaction is highlighted.
to restore the extracted features to the original action</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related</title>
    </sec>
    <sec id="sec-3">
      <title>Work</title>
      <p>Skeleton-based action recognition has been a growing
area of focus in computer vision research. The related
works broadly cover methodologies from hand-crafted
features to deep learning models.</p>
      <p>An early approach [31], in which a skeleton-based
representation called Actionlet Ensemble is used for action
recognition. It identifies and groups related parts of the
skeletons that form meaningful sub-actions, termed
Actionlets. The advent of deep learning has significantly
improved the performance of skeleton-based action
recognition. Additionally, a hierarchical RNN [32] was proposed
for skeleton-based recognition. The model hierarchically
constructs five parts of the body and then connects them
in a temporal recurrent layer. More recent works leverage
attention mechanisms to focus on discriminative joints
or frames. Moreover, an attention mechanism in Long</p>
      <sec id="sec-3-1">
        <title>Short-Term Memory (LSTM) networks [33] is proposed,</title>
        <p>which can selectively focus on informative joints in the
skeleton.</p>
        <p>Although there are already many excellent supervised
skeleton-based methods to recognize, these methods rely
on labels that we have made. Manual annotation not only
requires a lot of manpower and financial resources, and
accuracy cannot be guaranteed. If we take a supervised
approach, we must classify this set of actions into the
types of actions we already know, and there may be
some kinds we can’t discern. Besides, when we train a
sequence. They set up an evaluation system to measure
the diference between the original data and the restored
data. The cluster used the intermediate feature vectors
generated by the encoder. Specifically, the encoder is a
multi-layered bidirectional Gated Recurrent Unit (GRU)
and the decoder is a uni-directional GRU. Afterward,
another structure U-S-VAE [25], which is diferent from
[5] in that BLSTM is used instead of BI-GRU. Our
structure is also similar to several approaches [24][25]. The
encoder-decoder system is also a vital part of our network
structure. We adopt the hidden representation from the
encoder as our classification feature vector. Furthermore,
we incorporated TCN and multi-scale TCN in the
encoder to integrate temporal information. In addition, due
to the excellent performance of deconvolutional neural
networks in GAN networks for data generation [38], we
embedded TDCN (temporal deconvolutional network) in
the decoder.
3.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Methods</title>
      <sec id="sec-4-1">
        <title>3.1. Preliminary</title>
        <p>data is a sequence  
joints node:</p>
        <p>Data angle information extraction: The skeleton
3 of T frames, and each frame is the</p>
        <sec id="sec-4-1-1">
          <title>2D location information and confidence about the K-th</title>
          <p>3 = { 1,  2, ...  , ...,   }
  = { 1,   1,  1,  2,   2,  2, … ,  
 ,  

,   }
of the character in the view, we choose to use Angle in- paper, we will refer to the data extracted from the angle
Where   is the confidence of the 2D location information
( 1,   1), which is about the k-th joints node in the t frame.</p>
          <p>To overcome the diference in the position information
formation instead of the original 2D Cartesian coordinate
data.The Angel data is a sequence  
each frame is the angular formation of three nodes in</p>
          <p>2 of T frames, and
sequence and confidence:</p>
          <p>2 = { 1,  2, ...  , ...,   }
  = { 1, 
1, 
2, 
2, ..., 


,   }</p>
          <p>is the order of
Where   is the confidence of 

 . ℎ
the three adjacent nodes (e.g., right shoulder, right elbow,
right hand). To deal with the unity of left and right angles,
we take counterclockwise or clockwise angles on both
sides (for example, the angles of the right shoulder, right
elbow, and right hand is counterclockwise, and the angles
of the left hand, left elbow and left hand is clockwise).</p>
          <p>While we use the angular information,we should note
that we can no longer convert Angle data to coordination
data.Therefore, when this transformation happens, we
lose some information that we can’t be sure of useful.So
we propose two data supplement solutions.Both methods
add distance information to the original Angle data.The
ifrst is the distance from the center of the Angle to the
center of the body, which is the shoulder center, and
the other is the length of the second side formed by the
Angle.</p>
          <p>3 = { 1,  2, …   , … ,   }
  = { 1, 
1,  1, 
2, 
2,  2, … ,   ,</p>
          <p>,   }</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>For the sake of convenience in the later part of this information as AE data, while the data in the dataset that has not been changed in any way, i.e. original skeleton data, will be referred to as OS data.</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Model Architecture</title>
        <p>MSTCN-VAE network structure: The advantage of
unsupervised methods over fully supervised methods is
that they do not require manually labeled data. In this
paper, referring to the VAE unsupervised model proposed
by previous researchers [25][24], an encoder-decoder
model is introduced to learn unlabeled micro-gesture
sequence data (key points-based or
angle-informationextracted). However, compared to existing unsupervised
models our network has the following key diferences:</p>
        <sec id="sec-4-2-1">
          <title>1) We put TCN and Multi-scale TCN (MSTCN) into the</title>
          <p>encoder for integration of temporal information,
respectively. This is because TCN has been shown to have
better integration of temporal information compared to
RNNs, LSTMs, and GRUs [39]. Also, MSTCN will collect
more information than TCN due to the joint efect of
diferent size receptive fields [</p>
          <p>40]. 2) We embed a
temporal deconvolutional network module in the decoder
to generate an initial sequence of gesture actions based
on hidden features. On the one hand, deconvolutional
neural networks are widely used in Generative
adversarial networks to generate data [41],[42], and on the other
hand to use operations in the decoder similar to those MSTCN-VAE Motion Prediction: Our proposed
in the encoder as a way to better generate the original MSTCN-VAE network framework is depicted in detail in
micro-gesture sequence data. In terms of the loss func- Figure 1 and Figure 2. A four-dimensional data  of the
tion, similar to U-S-VAE [25] we use a linear combination shape ( ⋅  , ,  ,  ) , where  is equal to the size of the
of   and   as the plausible loss.   computes the MSE batch size,  represents the number of people in each
loss between the decoder-generated vector and the input frame,  represents the number of channels,  represents
vector, and this term aims to make the decoder-generated the frame length of each action sample, and  represents
result as similar as possible to the input action sequence the number of human features in each frame. It is
impordata.   computes the Kullback-Leibler (KL) divergence, tant to clarify that  = 3 in OS data and consists of the
and the KL divergence norm term is to ensure a closer x-coordinate and y-coordinate of the joint point and the
approximation to the joint distribution and the product of confidence level of that point, and  = 3 in AE data and
the marginals, i.e. makes the encoder-generated hidden consists of the angle value of the pinch angle, the length
variables conform to the standard normal distribution as of the line segment of the corresponding joint point, and
much as possible. the corresponding confidence level.  is first extracted by</p>
          <p>Hidden feature vector clustering: A vital feature the MSTCN module of the encoder with diferent scales
in our network architecture is that we use two fully con- of convolutional kernels for temporal information, and
nected layers after multi-scale convolution in the tempo- then the data is stitched according to the  dimension
ral dimension and use this to form feature clusters. In for data stitching and then shaped into ( ⋅  ,  ′ ⋅  ⋅  )
other words, the feature clusters consist of hidden fea- data ( ′ represents the size of  dimension after
stitchtures integrated by temporal convolution [43]. Such a ing). Subsequently, a HFE block consisting of two fully
strategy is efective and promising when unsupervised connected layers performs dimensionality reduction on
methods are used for clustering multidimensional se- this data, reducing the computational efort for clustering
quences, such as in body and gesture junction sequences while not losing information as much as possible. For
[44][25]. It has been experimented with and displayed the dimensionality reduction, the data is then shaped
that fully connected layers are extensively applicable to into ( ⋅  , ,  ⋅  ) three-dimensional data  ̃ after the
RNN architectures [24], and in our demonstration, it can deconvolution module and the initial time series data 
be found that fully connected layers under VAE struc- which reshaped as ( ⋅  , ,  ⋅  ) is used to calculate
tures will also help temporal convolution extract hidden the Loss value by  =   +  ⋅   , where   = ‖ −  ‖ ̃ 2, 
features to some extent. Therefore, we put a hidden fea- is used to describe the weight of the kl-divergence loss.
ture extraction (HFE) block which consists of two fully
connected layers into the end of the encoder to extract 3.3. Classification methods
the multi-nodal temporal information after MSTCN
integration. In this way, we implement a codec system,
called Multi-Scale Temporal Convolutional variational
autoencoder (MSTCN-VAE), in which the original time
series is input to the encoder and the encoder passes the
low-dimensional hidden feature vectors to the decoder.</p>
          <p>Unsupervised K-nearest neighbors classifier: In
order to evaluate our action classification efect more
explicitly, for the hidden feature vectors generated by
the encoder, we use the K-nearest neighbors classifier
(KNN). In other words, all the sequence data in the
training set are forward propagated in the current training</p>
          <p>Supervised
Unsupervised
iMiGUE dataset</p>
          <p>Methods
S-VAE
ST-GCN
Shift-GCN</p>
          <p>MS_G3D
TCN_VAE(with RFC)</p>
          <p>(OS data)(Our )
MSTCN_VAE(with RFC)</p>
          <p>(OS data)(Our )
MSTCN_VAE(with RFC)</p>
          <p>(AE data)(Our )
MSTCN_VAE(with RFC)
(AE data + OS data)(Our )</p>
          <p>P&amp;C</p>
          <p>U-S-VAE
TCN_VAE (with out HFE)
(OS data) (Our )</p>
          <p>TCN_VAE
(OS data)(Our )
MSTCN_VAE
(OS data)(Our )</p>
          <p>MSTCN_VAE
(AE data + OS data)(Our )
39.11
41.23
network to obtain the hidden feature vectors of all the
training data when calculating the accuracy, and this is
used to form the KNN classification space. After the same
forward propagation for each sample in the test set, the
KD-tree algorithm is used to quickly search for the
neighboring samples in the just-formed classification space.</p>
          <p>It is worth noting that although the composition of the
KNN classification space uses the labels of the training
set, the labels are only used to assign categories and are
not involved in model training.</p>
          <p>Supervised random forest classifier: The
supervised classification method was used to evaluate the
performance of our model from multiple perspectives.</p>
          <p>Specifically, for the hidden feature vector generated by
the encoder, we put it into a Random Forest classifier (RF
classifier). All training sets are also forward propagated
under the current network to obtain the hidden feature
vectors. The RF classifier is used to fit these vectors and
the accuracy is calculated on the test data after forward
propagation in the evaluation phase. It is important to
clarify that since the RF classifier uses the label
information to form the classification space, the model at this
point belongs to the supervised network.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Experiment</title>
      <sec id="sec-5-1">
        <title>4.1. Dataset</title>
        <p>iMiGUE: iMiGUE dataset [25] focuses on unconscious
micro-gesture movements without identity information.</p>
        <p>The dataset uses the OpenPose video dataset [45] ges- SMG dataset
ture estimation toolbox to extract 18499 action samples Methods Top1 Top5
from 359 post-race press conference videos, which are ST-GCN 41.48 86.07
wttcnhhaoeertdleelesgekaso,dseralioiemnztndeoeedntnhnsimenoiocntanoop-smo3:cr1oidtcwnirmnsooaii-s-ctgdterseisomso-otgfeufeVnerseast=ciucohar2netp2aeolaugicsponptprtiayoect.noriaEnbclsaoaiccdstohteyogofjrofrodadriimnianettasesataeaoinsssf USnvvusiisspueeepddre-r- MMSSTT(COCNSNSMh__dVi(SVft-OaAG_AtuEGaCEr()3(w)(NOODitSuhrd)RaFtaC)) 55634.4023...1570956 87944.1593...4542488
and prediction confidence scores. In MiGA Workshop
&amp; Challenge 2023, the entire dataset was divided into a
training dataset consisting of 13670 samples and a test VTaalbidleat2ion on the SMG dataset and Comparison with currently
dataset consisting of 4562 samples. known methods(best supervised method: Black with bold,</p>
        <p>
          SMG: SMG dataset [30] is a novel spontaneous micro- best unsupervised method: Blue with bold).RFC denotes
gesture dataset. From 414 long video instances con- random forest classifier.
taining 40 participants, the SMG dataset extracted 3712
micro-gesture action clips, where the average length of
these clips was 51.3 frames, and labeled them with 16
micro-gesture action categories as well as one non-micro- 4.2. Implementation Details
gesture category. Based on the authors’ suggestion, we
evaluated our proposed model with 610 test samples in
the body skeleton data model of this dataset. Following
the convention of some articles [46] [
          <xref ref-type="bibr" rid="ref1">47</xref>
          ], we show the
accuracy of Top1 and Top5 on this dataset.
        </p>
        <p>
          To train the network, each action sample was
downsampled by up to 100 frames. The joint point data in each
skeleton map were also normalized. For the
optimization hyperparameters, unless otherwise stated, all models
were optimizer: SGD, batch size: 32, an initial learning
rate: 0.0001, epoch: 200, LR decay rate: 0.1, LR decay step:
(100, 150). By random hyperparametric grid search, 1) for comparatively in table 1 and table 2.
the encoder that only uses TCN blocks, setting the follow- In table 1, first, compared between diferent
MSTCNing network structure: Encoder: using one TCN block, VAE variants. We find that the hidden feature extraction
each of which convolves T-dimension into 75, 50, 25, and block has about 3% improvement in the accuracy of the
1, Gradually; Decoder: by one TDCN block, which de- model, and Multi-scale has a 7% positive impact on the
convolutes T-dimension of sizes 50 and 100, Gradually. model. AE data has a negative impact on the model
com2) for the encoder that uses TCN block and HFE block, pared to OS data, but when AE data and OS data are
the following network structure is set: Encoder: consists judged together it brings a 4% improvement. In addition,
of one TCN block and one HFE block, TCN block con- the application of supervised classification in the VAE
volves T-dimension into 75, 50, 25, 1, Gradually; Decoder: network structure brings a significant 10% increase in the
one TDCN block, deconvolutes T-dimension into 50, 100, model. Second, compared to supervised algorithms that
Gradually. 3) For the encoder using one MSTCN block are also based on skeleton recognition, our supervised
and HFE, the following network structure is set: Encoder: model is at a considerable disadvantage since the graph
similar to the setting for MSTCN blocks [
          <xref ref-type="bibr" rid="ref2">48</xref>
          ], utilizing connectivity property in the skeleton data is not taken
one MSTCN block and one HFE block; Decoder: consists into account. It is worth mentioning that for the
superof one TDCN block, which deconvolutes the T-dimension vised algorithm S-VAE, which also does not consider this
into 50, 100, Gradually. property, our model has a considerable improvement in
        </p>
        <p>For the above 1) and 2) models, the hidden feature prediction. Finally, compared to similar unsupervised
vectors used for classification are all 66-dimensional, algorithms, the use of temporal convolution gives better
and for the 3) model, the hidden feature vectors are 128- classification results for skeleton-based data under the
dimensional. In order to avoid gradient explosion dur- encoder-decoder system. However, since the process of
ing the training process, gradient truncation will be per- calculating the Top5 of P&amp;C and U-S-VAE in article [25]
formed when the maximum norm is greater than 10. In is ambiguous, this leads to the accuracy of our top5 and
calculating the loss and performing backpropagation, we the top5 of the two models mentioned above not being
found that the overall training efect of the model was directly comparable.
best when the value of λ in the loss function was 0.8 after Of course, by validating the results on the SMG dataset
several experiments. For the iMiGUE dataset, both the in table 2, it can be found that the supervised and
unsuoriginal skeleton data and the pre-processed data with pervised MSTCN-VAE models are equally efective for
the above angular information were input to the model other micro-gesture datasets. It is worth noting that our
to demonstrate the improvement of the model accuracy approach is the first to use a completely unsupervised
with the new data. However, for the SMG dataset, we temporal convolution method in skeleton-based
recogonly validated the efectiveness of our model on the raw nition, and the results validate the efectiveness of our
skeleton data. approach.</p>
        <p>In the model evaluation session, for our diferent
MSTCN-VAE variants (a combination of the methods
described in Section 3.2), Top1 accuracy and Top5 accu- 5. Conclusion
racy are calculated uniformly using k=1 under the KNN
classifier and k=1, 2, 3, 4, 5 combined, and random_state
=1 under the Random forest classifier Top1 accuracy and
random_state =1, 2, 3, 4, 5 are used to calculate Top5
accuracy. In the top5 calculation, for the classification
results under five diferent parameters of the classifier,
the prediction is considered correct as long as it contains
the correct category. All experiments are based on an
RTX 3080 (10GB) GPU and a 12 vCPU Intel(R) Xeon(R)
Platinum 8255C CPU for training and evaluation.</p>
        <p>In this paper, we propose a novel skeleton-based
microgesture recognition method. Our model connects a
multiscale temporal convolutional network with a hidden
feature extraction block as an encoder to aggregate out
hidden feature vectors and uses a temporal deconvolutional
network in the decoder to generate action sequences from
the hidden feature vectors. Through experiments on the
iMiGUE dataset, we continuously improve and
demonstrate the improvement of the MSTCN-VAE model over
previous unsupervised methods, in addition to validation
on the SMG dataset to further illustrate the efectiveness
of our model.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.3. Evaluation and Comparison</title>
        <sec id="sec-5-2-1">
          <title>State-of-the-art supervised and unsupervised action</title>
          <p>
            recognition methods based on skeleton data have been Acknowledgments
applied to the iMiGUE dataset and the SMG dataset, e.g.
[
            <xref ref-type="bibr" rid="ref2">48</xref>
            ],[
            <xref ref-type="bibr" rid="ref3">49</xref>
            ],[
            <xref ref-type="bibr" rid="ref4">50</xref>
            ],[
            <xref ref-type="bibr" rid="ref5">51</xref>
            ],[24],[25]. To highlight the advantages Thanks to the developers of MS_G3D https://github.
of our model, the above methods and the detailed ac- com/kenziyuliu/ms-g3d, P&amp;C https://github.com/shlizee/
curacy of our method on these two data are presented Predict-Cluster and TCN https://github.com/locuslab/
[23] R. Zhi, J. Hu, F. Wan, Micro-expression based action recognition, Proceedings of the
recognition with supervised contrastive learn- AAAI Conference on Artificial Intelligence 34 (2020)
ing, Pattern Recognition Letters 163 (2022) 11045–11052. doi:1 0 . 1 6 0 9 / a a a i . v 3 4 i 0 7 . 6 7 5 9 .
25–31. URL: https://www.sciencedirect.com/ [36] S. Cho, M. H. Maqbool, F. Liu, H. Foroosh,
Selfscience/article/pii/S0167865522002690. doi:h t t p s : attention network for skeleton-based human action
/ / d o i . o r g / 1 0 . 1 0 1 6 / j . p a t r e c . 2 0 2 2 . 0 9 . 0 0 6 . recognition, in: 2020 IEEE Winter Conference on
[24] K. Su, X. Liu, E. Shlizerman, Predict &amp; cluster: Unsu- Applications of Computer Vision (WACV), 2020, pp.
pervised skeleton based action recognition, in: 2020 624–633. doi:1 0 . 1 1 0 9 / W A C V 4 5 5 7 2 . 2 0 2 0 . 9 0 9 3 6 3 9 .
IEEE/CVF Conference on Computer Vision and [37] Z. Gao, P. Wang, P. Lv, X. Jiang, Q. Liu, P. Wang,
Pattern Recognition (CVPR), 2020, pp. 9628–9637. M. Xu, W. Li, Focal and global spatial-temporal
doi:1 0 . 1 1 0 9 / C V P R 4 2 6 0 0 . 2 0 2 0 . 0 0 9 6 5 . transformer for skeleton-based action recognition,
[25] X. Liu, H. Shi, H. Chen, Z. Yu, X. Li, G. Zhao, imigue: in: L. Wang, J. Gall, T.-J. Chin, I. Sato, R. Chellappa
An identity-free video dataset for micro-gesture (Eds.), Computer Vision – ACCV 2022, Springer
understanding and emotion analysis, arXiv preprint Nature Switzerland, Cham, 2023, pp. 155–171.
arXiv:2107.00285, 2021. URL: https://arxiv.org/abs/ [38] S. Addepalli, G. Nayak, A. Chakraborty, R. Babu,
2107.00285, [cs.CV]. Degan: Data-enriching gan for retrieving
represen[26] K. Cho, B. van Merrienboer, C. Gulcehre, D. Bah- tative samples from a trained classifier, Proceedings
danau, F. Bougares, H. Schwenk, Y. Bengio, Learn- of the AAAI Conference on Artificial Intelligence 34
ing phrase representations using rnn encoder- (2020) 3130–3137. doi:1 0 . 1 6 0 9 / a a a i . v 3 4 i 0 4 . 5 7 0 9 .
decoder for statistical machine translation, 2014. [39] S. Bai, J. Z. Kolter, V. Koltun, An empirical
a r X i v : 1 4 0 6 . 1 0 7 8 . evaluation of generic convolutional and
recur[27] S. Hochreiter, J. Schmidhuber, Long short-term rent networks for sequence modeling, ArXiv
memory, Neural computation 9 (1997) 1735–1780. abs/1803.01271 (2018).
[28] S. Bai, J. Z. Kolter, V. Koltun, An empirical evalua- [40] J. Zhang, Y. Wang, J. Tang, J. Zou, S. Fan, Ms-tcn: A
tion of generic convolutional and recurrent net- multiscale temporal convolutional network for fault
works for sequence modeling, arXiv preprint diagnosis in industrial processes, in: 2021 American
arXiv:1803.01271 (2018). Control Conference (ACC), 2021, pp. 1601–1606.
[29] D. P. Kingma, M. Welling, Auto-encoding varia- doi:1 0 . 2 3 9 1 9 / A C C 5 0 5 1 1 . 2 0 2 1 . 9 4 8 2 7 2 8 .
          </p>
          <p>tional bayes, arXiv preprint arXiv:1312.6114 (2013). [41] A. Radford, L. Metz, S. Chintala, Unsupervised
rep[30] H. Chen, H. Shi, X. Liu, X. Li, G. Zhao, Smg: A resentation learning with deep convolutional
genmicro-gesture dataset towards spontaneous body erative adversarial networks, CoRR abs/1511.06434
gestures for emotional stress state analysis, Interna- (2015).
tional Journal of Computer Vision 131 (2023) 1–21. [42] M. Arjovsky, S. Chintala, L. Bottou,
Wasserdoi:1 0 . 1 0 0 7 / s 1 1 2 6 3 - 0 2 3 - 0 1 7 6 1 - 6 . stein generative adversarial networks, ICML’17,
[31] S. Maji, L. Bourdev, J. Malik, Action recognition JMLR.org, 2017, p. 214–223.</p>
          <p>from a distributed representation of pose and ap- [43] M. Farrell, S. Recanatesi, G. Lajoie, E. Shea-Brown,
pearance, in: CVPR 2011, 2011, pp. 3177–3184. Recurrent neural networks learn robust
representadoi:1 0 . 1 1 0 9 / C V P R . 2 0 1 1 . 5 9 9 5 6 3 1 . tions by dynamically balancing compression and
ex[32] Y. Du, W. Wang, L. Wang, Hierarchical recurrent pansion, in: Real Neurons &amp; Hidden Units: Future
neural network for skeleton based action recog- directions at the intersection of neuroscience and
nition, in: 2015 IEEE Conference on Computer artificial intelligence @ NeurIPS 2019, 2019. URL:
Vision and Pattern Recognition (CVPR), 2015, pp. https://openreview.net/forum?id=BylmV7tI8S.
1110–1118. doi:1 0 . 1 1 0 9 / C V P R . 2 0 1 5 . 7 2 9 8 7 1 4 . [44] K. Su, E. Shlizerman, Clustering and recognition of
[33] J. Liu, G. Wang, P. Hu, L.-Y. Duan, A. C. Kot, Global spatiotemporal features through interpretable
emcontext-aware attention lstm networks for 3d action bedding of sequence to sequence recurrent neural
recognition, in: 2017 IEEE Conference on Computer networks, 2020. a r X i v : 1 9 0 5 . 1 2 1 7 6 .
Vision and Pattern Recognition (CVPR), 2017, pp. [45] Z. Cao, G. Hidalgo, T. Simon, S.-E. Wei, Y. Sheikh,
3671–3680. doi:1 0 . 1 1 0 9 / C V P R . 2 0 1 7 . 3 9 1 . Openpose: Realtime multi-person 2d pose
estima[34] D. Miki, S. Chen, K. Demachi, Weakly supervised tion using part afinity fields, IEEE Transactions
graph convolutional neural network for human ac- on Pattern Analysis and Machine Intelligence 43
tion localization, in: 2020 IEEE Winter Conference (2021) 172–186. doi:1 0 . 1 1 0 9 / T P A M I . 2 0 1 9 . 2 9 2 9 2 5 7 .
on Applications of Computer Vision (WACV), 2020, [46] W. Kay, J. Carreira, K. Simonyan, B. Zhang,
pp. 642–650. doi:1 0 . 1 1 0 9 / W A C V 4 5 5 7 2 . 2 0 2 0 . 9 0 9 3 5 5 1 . C. Hillier, S. Vijayanarasimhan, F. Viola, T. Green,
[35] L. Huang, Y. Huang, W. Ouyang, L. Wang, Part- T. Back, P. Natsev, M. Suleyman, A. Zisserman, The
level graph convolutional network for skeleton- kinetics human action video dataset (2017).</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>A. Online Resources</title>
      <sec id="sec-6-1">
        <title>The sources code for the MSTCN-VAE model are avail</title>
        <p>able via
• MSTCN-VAE
• iMiGUE
• SMG</p>
      </sec>
      <sec id="sec-6-2">
        <title>The data set used in this article is available at</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>S.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Spatial temporal graph convolutional networks for skeleton-based action recognition</article-title>
          ,
          <source>AAAI'18/IAAI'18/EAAI'18</source>
          , AAAI Press,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [48]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <article-title>Disentangling and unifying graph convolutions for skeleton-based action recognition</article-title>
          ,
          <source>2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2020</year>
          )
          <fpage>140</fpage>
          -
          <lpage>149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [49]
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zhao, Bidirectional long short-term memory variational autoencoder</article-title>
          ,
          <source>in: British Machine Vision Conference</source>
          <year>2018</year>
          ,
          <article-title>BMVC 2018, Newcastle</article-title>
          , UK, September 3-
          <issue>6</issue>
          ,
          <year>2018</year>
          , BMVA Press,
          <year>2018</year>
          , p.
          <fpage>165</fpage>
          . URL: http://bmvc2018.org/ contents/papers/0963.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [50]
          <string-name>
            <given-names>S.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Spatial temporal graph convolutional networks for skeleton-based action recognition</article-title>
          ,
          <source>Proceedings of the AAAI Conference on Artificial Intelligence</source>
          <volume>32</volume>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1609/ aaai.v32i1.
          <fpage>12328</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [51] K. Cheng,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , J. Cheng, H. Lu,
          <article-title>Skeleton-based action recognition with shift graph convolutional network</article-title>
          ,
          <source>in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>180</fpage>
          -
          <lpage>189</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR42600.
          <year>2020</year>
          .
          <volume>00026</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>