<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>X (J. She);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Micro-gesture Recognition Using Data Augmentation and Spatial-Temporal Attention</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pengyu Liu</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kun Li</string-name>
          <email>kunli.hfut@gmail.com</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fei Wang</string-name>
          <email>jiafei127@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yanyan Wei</string-name>
          <email>weiyy@hfut.edu.cn</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Junhui She</string-name>
          <email>shejunhui@mail.ustc.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dan Guo</string-name>
          <email>guodan@hfut.edu.cn</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Guangzhou, China.</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Advanced Technology, University of Science and Technology of China</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Artificial Intelligence, Hefei Comprehensive National Science Center</institution>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Key Laboratory of Knowledge Engineering with Big Data (HFUT), Ministry of Education</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>ReLER, CCAI, Zhejiang University</institution>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>School of Computer Science and Information Engineering, School of Artificial Intelligence, Hefei University of Technology</institution>
          ,
          <addr-line>HFUT</addr-line>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Xinsight Lab, Research Institute, Hefei Zhongjuyuan Intelligent Technology Co., Ltd.</institution>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>In this paper, we introduce the latest solution developed by our team, HFUT-VUT, for the Micro-gesture Online Recognition track of the IJCAI 2025 MiGA Challenge. The Micro-gesture Online Recognition task is a highly challenging problem that aims to locate the temporal positions and recognize the categories of multiple microgesture instances in untrimmed videos. Compared to traditional temporal action detection, this task places greater emphasis on distinguishing between micro-gesture categories and precisely identifying the start and end times of each instance. Moreover, micro-gestures are typically spontaneous human actions, with greater diferences than those found in other human actions. To address these challenges, we propose hand-crafted data augmentation and spatial-temporal attention to enhance the model's ability to classify and localize micro-gestures more accurately. Our solution achieved an F1 score of 38.03, outperforming the previous state-of-the-art by 37.9%. As a result, our method ranked first in the Micro-gesture Online Recognition track.</p>
      </abstract>
      <kwd-group>
        <kwd>Micro-gesture online recognition</kwd>
        <kwd>micro action</kwd>
        <kwd>video understanding</kwd>
        <kwd>spatio-temporal attention</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        When humans express emotions or interact with the world, various non-verbal forms of communication
play a crucial role in the transmission of emotion and information [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1, 2, 3, 4, 5, 6, 7, 8, 9</xref>
        ]. During such
interactions, the human body often displays numerous spontaneous actions and gestures. Understanding
these subtle behaviors [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref13">10, 11, 12, 13</xref>
        ] is essential for gaining deeper insight into human behavior patterns
and emotional states [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. Examples include gestures such as “folding arms”, “playing or adjusting
hair”, and “crossing legs”. In many scenarios, individuals may consciously suppress or conceal their
genuine emotions due to social etiquette or contextual considerations. However, because micro-gestures
are often spontaneous, they can serve as an indicator of a person’s true emotional state. Although there
have been great successes in conventional action understanding [
        <xref ref-type="bibr" rid="ref14 ref15">14, 15</xref>
        ], the research on micro-gesture
analysis [
        <xref ref-type="bibr" rid="ref16 ref17 ref18">16, 17, 18</xref>
        ] is still in its infancy.
      </p>
      <p>
        Due to the significant imbalance in the category distribution of the SMG [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] dataset, we introduced
data augmentation to expand the number of samples of the rare category. At the same time, to address
the issue of the model’s insuficient focus on key temporal and spatial information for micro-gesture
recognition, we designed a spatial-temporal attention module to strengthen the ability to recognize key
areas. In summary, the main contributions of this paper are as follows:
      </p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073
• We introduce a spatial-temporal attention to enhance the baseline’s localization head,
encouraging the model to focus on more informative areas of the feature output. Compared with the
baseline model, our approach achieves improved action classification and more precise boundary
localization.
• To mitigate the severe category imbalance caused by the spontaneous nature of micro-gestures
in real-world scenarios, we augment the dataset to improve the model’s sensitivity to gesture
categories with fewer samples. This improves the classification performance for gesture categories
with fewer samples.
• In the Micro-gesture Online Recognition challenge, our proposed method achieved an F1 score of
38.03 in the test set, securing first place in the competition. Experimental results demonstrate
that our model is capable of efectively distinguishing and localizing micro-gestures 1.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. Micro-Gesture Analysis Datasets</title>
        <p>
          The SMG [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] dataset is designed for studying naturally occurring micro-gestures under stress. It
contains micro-gestures collected from 40 participants of diverse ages, genders, and ethnic backgrounds.
The dataset categorizes micro-gestures into 16 categories and has been widely used in micro-gesture
recognition and emotion analysis tasks, demonstrating its practicality and efectiveness in these
research areas. The iMiGUE [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] dataset is the first publicly available dataset aimed at recognizing and
understanding suppressed or hidden emotions through micro-gestures. It includes 359 videos with a
total duration of 2,092 minutes, collected from 72 subjects from 28 countries. Some studies suggest that
micro-actions [
          <xref ref-type="bibr" rid="ref10 ref19 ref20 ref7">7, 10, 19, 20</xref>
          ], which focus on spontaneous actions of the whole body, can better reflect
subtle emotional changes in humans. The MA-52 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] dataset consists of 52 micro-action categories and 7
body part labels, covering a wide range of natural micro-actions. It comprises 22,422 instances collected
from 205 participants during psychological interview sessions. In addition, Li et al. [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] introduced the
Multi-label Micro-Action 52 (MMA-52) dataset and proposed a Multi-label Micro-Action Detection task,
which aims to recognize all micro-actions within a video sequence for fine-grained understanding. The
MMA-52 dataset comprises 6,528 videos and 19,782 action instances collected from 203 subjects.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Micro-gesture Online Recognition</title>
        <p>
          Guo et al. [22] proposed a novel deep network that integrates Graph Convolutional Networks (GCNs)
and Transformers to extract motion features from 2D skeleton sequences. This hybrid design leverages
the strengths of both GCNs and Transformers, efectively capturing spatial relationships and long-range
temporal dependencies. Their method achieved first place in the Micro-gesture Online Recognition
track of the MiGA 2023 Challenge. Wang et al. [23] developed a deep network with dual-stream
input for micro-gesture online recognition. They first used a sequential action recognition model to
extract gesture features from RGB and skeleton sequences, respectively, and then used a multi-scale
Transformer encoder to process these features as a detection model. Their approach secured first place
in the Micro-gesture Online Recognition track of the MiGA 2024 Challenge. Additionally, Liu et al. [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]
proposed a model using learnable query points and Mamba blocks. They used learnable points to learn
the positions of frames that are more important for gesture recognition and leveraged Mamba’s ability
to eficiently capture complex relationships in sequence data, significantly improving the model’s ability
to recognize micro-gestures, and achieved second place in the MiGA 2024 challenge using only RGB
data.
Raw Data
Augmented Data
1
2
3
4
5
6
7
11
12
13
14
15
16
        </p>
        <p>17</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Temporal Action Detection</title>
        <p>Temporal Action Detection (TAD) [24, 25, 26, 27] aims to locate and classify all actions in untrimmed
videos. Existing methods can generally be divided into two categories: feature-based approaches and
end-to-end approaches. Feature-based methods typically rely on pre-trained feature extractors to
obtain video representations, which are then used for subsequent processing. In contrast, end-to-end
approaches jointly optimize video encoders and decoders to achieve better task-specific feature
representations. End-to-end approaches allow more seamless modeling through the simultaneous optimization of
both encoding and decoding stages. For example, Tan et al. [24] proposed an end-to-end action detection
model, PointTAD, which utilizes learnable query points to accurately localize and diferentiate actions
in videos. Liu et al. [25] introduced the concept of fine-tuning large language models into the TAD task
by employing VideoMAE [28, 29] as the backbone and fine-tuning it for action localization, achieving
precise classification and localization. On the other hand, feature-based approaches are favored for
their eficiency and lower computational cost. Tirupattur
et al. [30] incorporated an attention-based
multi-label dependency layer into their model, significantly improving the modeling of co-occurrence
and temporal dependencies between actions. Dai et al. [31] proposed a novel ConvtransFormer network
that efectively integrates global and local temporal relations. Zhang et al. [26] employed a Transformer
encoder to capture long-range dependencies, while Shi et al. [32] introduced a ternary point modeling
approach for more accurate boundary localization. Yang et al. [27] dynamically aggregated multi-scale
features to handle actions of varying temporal lengths. Together, these approaches provide valuable
insights into accurately localizing and distinguishing complex temporal actions in untrimmed videos.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <sec id="sec-3-1">
        <title>3.1. Problem Definition</title>
        <p>Given an untrimmed video  , represented as a sequence of feature vectors  = { 1,  2, … ,   }, where 
denotes the temporal length, each   ∈  is typically extracted using pre-trained video encodes, such as
I3D [33] or VideoMAE [28, 29]. The objective of Micro-gesture Online Recognition is to predict a set of
action instances Ψ = { 1,  2, … ,   }, where  = {1, 2, … , }
Each instance   = {


,   ,   } is defined by its start time 
 , end time   , and category label   , where
denotes the number of predicted instances.
  ∈  , and  is the set of all predefined action categories.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Data Augmentation</title>
        <p>The SMG dataset captures micro-gestures that naturally occur in daily life. However, the frequency of
occurrence varies between diferent gesture categories. For instance, gestures like “Moving legs” appear
far more frequently than rarer gestures such as “Touching or covering suprasternal notch.” To address
the severe category imbalance in the training set, we designed a category-frequency-based adaptive
label augmentation strategy.</p>
        <p>+DynE layer
  +DynElayer
  +DynElayer
  +DynElayer
e
c
n
e
u
q
e
o
e
d
i
S</p>
        <p>B
V</p>
        <p>V
e
n
o
b
k
c
a
o
e
d
i
DS</p>
        <p>LN</p>
        <p>DynEM
⊕</p>
        <p>GN</p>
        <p>FFN</p>
        <p>⊕
Squeeze</p>
        <p>D Conv
D Conv
D Conv
Shift
⊕
⊗
⊗</p>
        <p>D Conv
G Conv</p>
        <p>⊕</p>
        <sec id="sec-3-2-1">
          <title>Feature Pyramid Multi-scale Spatial-Temporal Attention TAD Head</title>
          <p>LN
LN
LN
LN
 
 
 
 
  +
 
  −
US Path
DSPath
US Path</p>
          <p>DSPath
Spatial-Temporal
Atention
Spatial-Temporal
Atention
Path Spatial-Temporal
Atention</p>
          <p>LN
LN
LN
Path ⊕ ⊕
  +</p>
          <p>Up
Sample
Down
Sample
 
 
 2
 
 
 
 
 
 
Input
feature
 
  ′</p>
          <p>Temporal Atention</p>
          <p>Module
MaxPool
AvgPool
⊗
Shared MLP</p>
          <p>Spatial
Atention
Module
MaxPool &amp; AvgPool</p>
          <p>Temporal Atention Module</p>
          <p>Conv
Spatial Atention Module
ClassHead
Localization Head
ClassHead
Localization Head
ClassHead
Localization Head
ClassHead
Localization Head
⊗   ′′</p>
          <p>Refined
feature
⊕   ′
  ′′</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>DynE Layer</title>
        </sec>
        <sec id="sec-3-2-3">
          <title>Feature Fusion</title>
        </sec>
        <sec id="sec-3-2-4">
          <title>Spatial-Temporal Attention</title>
          <p>where  is the minimum instance threshold, and   is the number of instances of category  in the training
set. Specifically, for each instance, if its   &lt;  , we consider it a rare category. We then replicate its
annotations in the training data according to the calculated   , efectively increasing the representation
of that category. This augmentation is performed at the annotation level rather than duplicating raw
video data, preserving both the structure and diversity of the dataset. It avoids redundancy caused by
naive duplication and enhances the model’s ability to learn rare categories. Figure 1 demonstrates the
distribution of instances across the gesture category before and after applying our data augmentation
strategy. It clearly shows a significant reduction in category imbalance, thus enabling the model to
better learn rare gestures during training.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Overall Architecture</title>
        <p>Considering the unique characteristics of micro-gestures, we enhance the DyFADet [27] to build a
micro-gesture online recognition model that extracts discriminative representations from video encoder
features and dynamically adjusts the detection head for actions of varying lengths. Following the
DyFADet model architecture, our approach is composed of three main components: a feature extraction
backbone, an encoder, and a multi-scale spatial-temporal attention TAD head. Specifically, we first
use a pretrained video encoder to extract video features. These features are then passed through an
encoder to generate a feature pyramid. Within this pyramid, the features are downsampled using a
stride of 2 via the Dynamic Encoder (DynE) layer to obtain representations at diferent temporal scale
features F  , where  = 1, 2, … ,  , and  is the total number of pyramid levels. Finally, the multi-scale
spatial-temporal attention TAD head is used to detect action categories and their temporal boundaries.
The overall architecture is illustrated in Figure 2.</p>
        <p>The feature encoder in DyFADet introduces the DynE to enhance both global and local modeling
capabilities during action detection. DynE is designed based on the Transformer, replacing standard
self-attention with DynE. It contains two parallel branches: the instance-dynamic branch and the
multi-kernel branch. These branches collaboratively generate dynamic weighted masks to improve the
discriminative power of the feature representations. In the instance-dynamic branch, a DFA convolution
with kernel size 1 is applied to model features and generate a global temporal attention mask that
captures overall action information. In contrast, the multi-kernel branch applies DFA convolutions
with multiple kernel sizes to generate attention masks with various receptive fields, better adapting to
local structural diversity. The two branches are defined as:</p>
        <p>F instance-dynamic = DFA_Conv1 (Squeeze(LN(DS(F ))) ,</p>
        <p>F multi-kernel = DFA_Conv, (LN(DS(F ))),
where DS denotes the down-sample the feature with a scale of 2 to generate the representations with
diferent temporal resolutions. LN is Layer Normalization, Squeeze represents average pooling along
the channel dimension, and  is a parameter used to expand the convolution window size for better
temporal modeling. The outputs of both branches are then added to the original input features, forming
the final representation of the DynE. In the complete feature encoding process, each DynE performs
downsampling on the input features, constructing a multi-scale temporal feature representation. This
dynamic feature selection mechanism efectively mitigates the lack of feature discriminability in previous
models and significantly improves the overall action detection performance.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Spatial-Temporal Attention</title>
        <p>To address the inconsistency in detecting short-duration and long-duration instances caused by
traditional static methods, shared detection heads, DyFADet [27] adopts a Multi-Scale TAD Head architecture.
This architecture dynamically fuses features across multiple scales to guide the detection head in
adaptively adjusting its parameters based on the input context. However, micro-gestures exhibit strong
spatial-temporal dependencies, and the current Multi-Scale TAD Head design lacks suficient attention
to spatial-temporal features, which may lead to inaccurate boundary localization. Therefore, we
introduce a Spatial-Temporal Attention to enhance the model’s sensitivity to spatial-temporal information.
As illustrated in the figure 2, our detection head takes input features from the current level of the
pyramid along with its neighboring upper and lower levels. Through the Spatial-Temporal Attention
module, along with upsampling (US) and downsampling (DS) operations, we construct three parallel
paths for subsequent detection tasks.</p>
        <p>For the Spatial-Temporal Attention module, in the Temporal Attention branch, we first compress
the input feature map along the spatial dimensions to obtain a one-dimensional vector. During spatial
compression, both average pooling and max pooling are considered. These operations aggregate the
spatial information of the feature maps and are fed into a shared MLP network to generate a temporal
attention map. The spatially compressed features are summed element-wise to yield the final temporal
attention, which is defined as:</p>
        <p>(F ) =  ( MLP(AvgPool(F )) +MLP(MaxPool(F ))),
where  denotes the sigmoid function.</p>
        <p>In the Spatial Attention branch, the input features are compressed along the channel dimension using
both average pooling and max pooling. The resulting pooled features are concatenated and passed
through a convolutional layer to produce the spatial attention map. This can be formulated as:
 
 (F ) =  (</p>
        <p>7×7([AvgPool(F );MaxPool(F )])),
where  is the sigmoid function and 7 × 7 denotes the convolution kernel size.</p>
        <p>Finally, our classification and boundary regression modules operate on the aggregated features from
all pyramid levels. The class head uses a 1D convolution followed by a sigmoid function to predict the
(2)
(3)
(4)
(5)
probability of each action category at each time. The localization head applies a ReLU-activated 1D
convolution to estimate the temporal ofsets from the current time to the start and end times of the
action instance. This unified structure enables robust detection of actions with varying durations while
incorporating both temporal and spatial information, thereby achieving more generalized and accurate
temporal action localization.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <sec id="sec-4-1">
        <title>4.1. Dataset and Evaluation Metric</title>
        <p>
          The SMG [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] dataset consists of 3,692 samples covering 17 micro-gesture categories. A cross-subject
evaluation protocol is adopted, wherein 40 subjects are divided into two groups: a training group
comprising long video sequences from 35 subjects, and a testing group consisting of sequences from
the remaining 5 subjects. The dataset provides both RGB data and skeleton data. However, we achieve
state-of-the-art performance using only RGB data as input. We jointly evaluate the detection and
classification performance of algorithms using the F1 score:
 1 =
2 ⋅ Precision ⋅ Recall
Precision + Recall .
(6)
        </p>
        <p>Given a long video sequence for evaluation, Precision is the ratio of correctly classified
microgestures to the total number of gestures retrieved by the algorithm in the sequence. Recall is the ratio
of correctly retrieved micro-gestures to the total number of annotated micro-gestures in the ground
truth. This metric comprehensively reflects the algorithm’s ability to both detect and correctly classify
micro-gestures.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Implementation Details</title>
        <p>We use VideoMAEv2-g [28] as the video backbone to extract features from video sequences. The videos
are processed at the original frame rate of 28 fps, and a sliding window mechanism is adopted for
feature extraction, where each window contains 16 frames with a stride of 4 frames. To standardize
the input size of the model, all frames are resized to 224×224. The batch size is set to 128, the initial
learning rate is 1e-4, and the training is conducted for a total of 400 epochs.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Experimental Results</title>
        <p>Main Comparison. The experimental results comparing the performance of our method with other
models are shown in Table 1. On the test set of the SMG dataset, our method using VideoMAEv2-g
features surpassed the previous state-of-the-art performance, and our team’s model ranked first. Our
solution achieved an F1 score of 38.03, outperforming the previous state-of-the-art by 37.9%. The
proposed data augmentation and spatial-temporal attention demonstrate high performance, proving
their efectiveness in micro-gesture online recognition and indicating their ability to capture richer
semantic features.</p>
        <p>Ablation Studies. In Table 2, we report the results of our experiments conducted on several
baselines, comparing our proposed method with existing approaches to demonstrate its efectiveness.
We adopted VideoMAEv2-g as the backbone and applied the method proposed by Liu et al. [25],
achieving the best result with an F1 score of 38.03. However, due to time constraints, we did not
further optimize or fine-tune the model, and we report this result solely as a reference for future
research. Furthermore, we illustrate the efectiveness of the proposed data augmented strategy and
the Spatial-Temporal Attention. We conducted the following experiments: a baseline experiment, an
experiment incorporating data augmentation, one incorporating Spatial-Temporal Attention, and one
combining both techniques. The performance of all experiments outperformed the baseline. Since data
augmentation efectively addressed the category imbalance of micro-gestures, enabling the model to
focus more on and distinguish rare categories, performance improved by 15.33%. For the detection head,
we find that using spatial-temporal attention in the detection head can better leverage information,
improving performance by 17.69%. However, using data augmentation or spatial-temporal attention
alone only provides limited improvements. When combining data augmentation and spatial-temporal
attention, our method achieves a performance of 33.44, significantly outperforming the baseline.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Error Analysis</title>
        <p>In addition, to evaluate the performance of our proposed model, we followed the standard practice in
action detection by using the action detection evaluation toolkit proposed by Alwassel et al. [34].</p>
        <p>False Negative Analysis. Figures 3(a) and 3(b) illustrate the false negative analysis for the baseline
and our improved methods, focusing on the true action instances that were not detected by the models.
By analyzing the missed detection rates at a tIoU threshold of 0.5 under diferent Coverage, Length, and
Instance conditions, we can evaluate the performance of both models in various scenarios. To align
with the characteristics of the SMG dataset, we define the Length intervals as [0, 2, 5, 7, 9.75, INF] and
the Instance intervals as [-1, 15, 100, 200, INF], which are labeled on the axes as [XS, S, M, L, XL]. It
can be observed that our model significantly reduces the false negative rate for short-duration actions,
which account for the majority of the data. Additionally, in low-density videos, our model shows a
lower rate of missed detections compared to the baseline. In summary, the false negative analysis
indicates that our model demonstrates stronger detection capability on samples with a higher number
of action instances.</p>
        <p>False Positive Analysis. Figures 3(c) and 3(d) show the false positive analysis of the baseline and
our improved method, focusing on five common types of false detection errors. We present the false
positive analysis at tIoU = 0.5, where the x-axis represents the top NG predictions, with G denoting the
number of ground truth instances. From the comparison, it can be observed that false positives are
mainly concentrated in confusion errors and background errors. Compared to the baseline, our method
G1 G2 G3 G4 G5 G6 G7 G8 G9 G01</p>
        <p>Top Predictions</p>
        <p>G1 G2 G3 G4 G5 G6 G7 G8 G9 G01</p>
        <p>Top Predictions
(c) False Positive Analysis on Baseline</p>
        <p>(d) False Positive Analysis on Our Method
Double Detection Err</p>
        <p>True Positive
R1e.0moving Error Impact</p>
        <p>0.8
)
PAN (t%0.8
-em en0.6 0.50.5
ag vom0.4 0.3
rveA Irpm0.2 0.1
0.0</p>
        <p>Error Type
significantly reduces confusion errors. Moreover, for nearly all categories of removing error impact, our
model achieves a clear reduction, indicating improvements not only in boundary localization accuracy
but also in action classification precision. In summary, the false positive analysis indicates that our
proposed approach achieves improvements across nearly all top predictions. At the 10G level, our model
shows the most significant reduction in confusion errors, demonstrating a substantial enhancement in
the model’s ability to distinguish between diferent actions.</p>
        <p>Sensitivity Analysis. Finally, we analyzed the model’s sensitivity to variations in action
characteristics. As shown in Figures 3(e) and 3(f), compared to the baseline, our method exhibits a noticeably
reduced sensitivity to changes in coverage, length, and instance count. In particular, the sensitivity drop
is more significant for actions with length category L and instance count category XS. This indicates
that our model achieves stronger robustness and generalization in micro-gesture online detection.</p>
        <p>Figure 4 shows typical examples of incorrect micro-gesture classification by the model. The left figure
shows the model incorrectly predicting “Rubbing hands and crossing fingers” as “Folding arms,” and
the right figure shows the model incorrectly predicting “Scratching or touching facial parts other than
eyes” as “Playing or adjusting hair.” These two incorrect predictions demonstrate the model’s lack of
sensitivity to finger movements and its dificulty in classifying micro-gestures involving multiple body
parts, such as limbs and the head.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, we presented our solution for the Micro-gesture Online Recognition track of the IJCAI
2025 MiGA Challenge. Our approach is based on the DyFADet model, with the introduction of a
data augmentation strategy to alleviate the severe category imbalance commonly observed in
realworld micro-gestures. Additionally, we enhance the model’s ability to capture the spatial-temporal
dependencies of micro-gestures by incorporating a Spatial-Temporal Attention into the detection head.
Our final model achieved a score of 38.03 on the test set of the SMG dataset. Notably, our model relies
solely on RGB data for recognition. In future research, we plan to explore a wider variety of data
augmentation methods to address the issue of overfitting that may be caused by data augmentation
based on category frequency, thereby improving the model’s generalization. At the same time, we plan
to incorporate skeleton data and explore joint modeling using both RGB and skeleton modalities to
further improve online micro-gesture recognition performance.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work is supported by National Key R&amp;D Program of China (NO.2024YFB3311602), Natural Science
Foundation of China (62272144), the Anhui Provincial Natural Science Foundation (2408085J040), and the
Major Project of Anhui Provincial Science and Technology Breakthrough Program (202423k09020001),
and the Fundamental Research Funds for the Central Universities (JZ2024HGTG0309, JZ2024AHST0337).</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the manuscript preparation, ChatGPT was used exclusively to assist with grammar correction,
spelling checks, and minor language polishing. All outputs were carefully reviewed and revised by the
authors, who take full responsibility for the final content of the publication.
[22] X. Guo, W. Peng, H. Huang, Z. Xia, Micro-gesture online recognition with graph-convolution and
multiscale transformers for long sequence., in: MiGA@ IJCAI, 2023.
[23] Y. Wang, L. Kerui, H. Huang, Z. Xia, Micro-gesture online recognition with dual-stream multi-scale
transformer in long videos, MiGA@ IJCAI (2024).
[24] J. Tan, X. Zhao, X. Shi, B. Kang, L. Wang, Pointtad: Multi-label temporal action detection with
learnable query points, Advances in Neural Information Processing Systems 35 (2022) 15268–15280.
[25] S. Liu, C.-L. Zhang, C. Zhao, B. Ghanem, End-to-end temporal action detection with 1b parameters
across 1000 frames, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 2024, pp. 18591–18601.
[26] C.-L. Zhang, J. Wu, Y. Li, Actionformer: Localizing moments of actions with transformers, in:</p>
      <p>Proceedings of the European Conference on Computer Vision, Springer, 2022, pp. 492–510.
[27] L. Yang, Z. Zheng, Y. Han, H. Cheng, S. Song, G. Huang, F. Li, Dyfadet: Dynamic feature aggregation
for temporal action detection, in: Proceedings of the European Conference on Computer Vision,
Springer, 2024, pp. 305–322.
[28] Z. Tong, Y. Song, J. Wang, L. Wang, Videomae: Masked autoencoders are data-eficient learners
for self-supervised video pre-training, Advances in Neural Information Processing Systems 35
(2022) 10078–10093.
[29] L. Wang, B. Huang, Z. Zhao, Z. Tong, Y. He, Y. Wang, Y. Wang, Y. Qiao, Videomae v2: Scaling
video masked autoencoders with dual masking, in: Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition, 2023, pp. 14549–14560.
[30] P. Tirupattur, K. Duarte, Y. S. Rawat, M. Shah, Modeling multi-label action dependencies for
temporal action localization, in: Proceedings of the IEEE/CVF Conference on Computer Vision
and Pattern Recognition, 2021, pp. 1460–1470.
[31] R. Dai, S. Das, K. Kahatapitiya, M. S. Ryoo, F. Brémond, Ms-tct: Multi-scale temporal
convtransformer for action detection, in: Proceedings of the IEEE/CVF Conference on Computer Vision and
Pattern Recognition, 2022, pp. 20041–20051.
[32] D. Shi, Y. Zhong, Q. Cao, L. Ma, J. Li, D. Tao, Tridet: Temporal action detection with relative
boundary modeling, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 2023, pp. 18857–18866.
[33] J. Carreira, A. Zisserman, Quo vadis, action recognition? a new model and the kinetics dataset,
in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp.
6299–6308.
[34] H. Alwassel, F. C. Heilbron, V. Escorcia, B. Ghanem, Diagnosing error in temporal action detectors,
in: Proceedings of the European Conference on Computer Vision, 2018, pp. 256–272.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zhao, imigue: An identity-free video dataset for microgesture understanding and emotion analysis</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10631</fpage>
          -
          <lpage>10642</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zhao, Smg: A micro-gesture dataset towards spontaneous body gestures for emotional stress state analysis</article-title>
          ,
          <source>International Journal of Computer Vision</source>
          <volume>131</volume>
          (
          <year>2023</year>
          )
          <fpage>1346</fpage>
          -
          <lpage>1366</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . Zhao,
          <article-title>Analyze spontaneous gestures for emotional stress state recognition: A micro-gesture dataset and analysis with deep learning</article-title>
          ,
          <source>in: 2019 14th IEEE International Conference on Automatic Face &amp; Gesture Recognition</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Ying</surname>
          </string-name>
          ,
          <article-title>A survey on fmri-based brain decoding for reconstructing multimodal stimuli</article-title>
          ,
          <source>arXiv preprint arXiv:2503.15978</source>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Balasubramanian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hoai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Samaras</surname>
          </string-name>
          ,
          <article-title>Learning visual emotion representations from web data</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>13106</fpage>
          -
          <lpage>13115</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. W.</given-names>
            <surname>Schuller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kälviäinen</surname>
          </string-name>
          ,
          <article-title>Identity-free artificial emotional intelligence via micro-gesture understanding</article-title>
          ,
          <source>arXiv preprint arXiv:2405.13206</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Benchmarking micro-action recognition: Dataset, methods, and applications</article-title>
          ,
          <source>IEEE Transactions on Circuits and Systems for Video Technology</source>
          <volume>34</volume>
          (
          <year>2024</year>
          )
          <fpage>6238</fpage>
          -
          <lpage>6252</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <year>Mac 2024</year>
          :
          <article-title>Micro-action analysis grand challenge</article-title>
          ,
          <source>in: Proceedings of the 32nd ACM International Conference on Multimedia</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>11304</fpage>
          -
          <lpage>11305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Temporal-frequency state space duality: An eficient paradigm for speech emotion recognition</article-title>
          ,
          <source>in: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2025</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          , G. Chen,
          <string-name>
            <given-names>C.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Prototypical calibrating ambiguous samples for micro-action recognition</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>39</volume>
          ,
          <year>2025</year>
          , pp.
          <fpage>4815</fpage>
          -
          <lpage>4823</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Data augmentation for human behavior analysis in multiperson conversations</article-title>
          ,
          <source>in: Proceedings of the 31st ACM International Conference on Multimedia</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>9516</fpage>
          -
          <lpage>9520</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Frequency decoupling for motion magnification via multi-level isomorphic architecture</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>18984</fpage>
          -
          <lpage>18994</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          , Eulermormer:
          <article-title>Robust eulerian motion magnification via dynamic ifltering within transformer</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>38</volume>
          ,
          <year>2024</year>
          , pp.
          <fpage>5345</fpage>
          -
          <lpage>5353</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Repetitive action counting with hybrid temporal relation modeling</article-title>
          ,
          <source>IEEE Transactions on Multimedia</source>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <article-title>Exploiting ensemble learning for cross-view isolated sign language recognition</article-title>
          ,
          <source>arXiv preprint arXiv:2502.02196</source>
          (
          <year>2025</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Joint skeletal and semantic embedding loss for micro-gesture classification</article-title>
          ,
          <source>arXiv preprint arXiv:2307.10624</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Prototype learning for micro-gesture classification</article-title>
          ,
          <source>arXiv preprint arXiv:2408.03097</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Micro-gesture online recognition using learnable query points</article-title>
          ,
          <source>arXiv preprint arXiv:2407.04490</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Motion matters: Motion-guided modulation network for skeleton-based micro-action recognition</article-title>
          ,
          <source>in: Proceedings of the 33rd ACM International Conference on Multimedia</source>
          ,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Qu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Unified multi-modal unsupervised representation learning for skeleton-based action understanding</article-title>
          ,
          <source>in: Proceedings of the 31st ACM International Conference on Multimedia</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>2973</fpage>
          -
          <lpage>2984</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Mmad: Multi-label micro-action detection in videos</article-title>
          ,
          <source>arXiv preprint arXiv:2407.05311</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>