<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Challenge on Micro-gesture Analysis for Hidden Emotion Understanding (MiGA) 2024: Dataset and Results</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haoyu Chen</string-name>
          <email>chen.haoyu@oulu.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Björn W. Schuller</string-name>
          <email>bjoern.schuller@imperial.ac.uk</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ehsan Adeli</string-name>
          <email>eadeli@stanford.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guoying Zhao</string-name>
          <email>guoying.zhao@oulu.fi</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CMVS, University of Oulu</institution>
          ,
          <country country="FI">Finland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>GLAM, Imperial College London</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Stanford University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper summarizes the 2nd Challenge of Micro-gesture Analysis for Hidden Emotion Understanding (MiGA) 2024. The competition was split into two independent tracks: micro-gesture classification from pre-segmented data clips, and micro-gesture online recognition in sequences of continuous data. In this edition of the MiGA challenge, both tracks use multi-modal data (RGB and skeleton as modalities). For evaluation, accuracy for classification and F1 score for online recognition are used as the evaluation measure. Two large micro-gesture datasets (iMiGUE and SMG) were made publicly available and the Kaggle platform was used to manage the competition. Results achieved a classification accuracy of 70.25% for micro-gesture classification, showing a significant improvement compared to last year's competition, meanwhile, an F1 score for online recognition is about 0.2757 was achieved for multi-modal gesture recognition, showing the task is still challenging and leaves considerable margin for improvement. Afective computing, behavior analysis, multi-modal gesture recognition, micro-gestures, emotion Understanding emotions is fundamental to human intelligence and should hold a similar significance in artificial intelligence [ 1]. In the domains of emotion analysis and recognition, prior research has largely concentrated on facial expressions, vocal intonations, and physiological indicators such as heart rate. Relatively few studies, however, have explored the interpretation of emotions through gestural behaviors. Psychological research indicates that body language plays a crucial role in understanding emotions [2]. Notably, when individuals attempt to conceal their feelings, they often adjust their facial expressions but struggle to completely suppress micro-expressions. Moreover, only a few people actively manage their body movements in these ∗Corresponding author.</p>
      </abstract>
      <kwd-group>
        <kwd>understanding</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        CEUR
Workshop
Proceedings
situations. This suggests that gestures may ofer valuable insights into hidden or suppressed
emotions [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ].
      </p>
      <p>
        With the above observations, we initiated the workshop series focusing on the micro-gesture
(MG) analysis for hidden emotion understanding, which is a novel research direction for
the computer vision community. Micro-gestures (MGs) are subtle, involuntary body movements
that can reveal suppressed or hidden emotions, often used in psychology to interpret inner
feelings. MGs encompass various gestures – such as scratching the head, touching the nose,
or fidgeting with clothing – and difer from typical gestures by lacking any communicative
or illustrative purpose. Instead, they are spontaneous responses to certain stimuli, especially
negative ones. Existing researchers mainly work on ordinary gestures/actions [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ], which
are often used to convey semantic meanings or express attitudes. Meanwhile, MGs may occur
when individuals try to mask true emotions, like stress or nervousness, in high-stakes situations.
These micro-expressions can provide insight into hidden emotional states and may also correlate
with neurological or mental disorders, making them valuable for diagnostic support. Automatic
MG recognition has promising applications in areas such as human-computer interaction, social
media, public safety, and healthcare.
      </p>
      <p>
        In 2023, MiGA organized the first challenge on MG recognition with only skeleton modality
data recorded with Kinect V21. In the 2023 challenge, 54 entrants participated in the MiGA
challenge which was devoted to skeleton-based MG classification and online recognition. With
more research interests gained in the recognition of micro-gestures [
        <xref ref-type="bibr" rid="ref10 ref11 ref8 ref9">8, 9, 10, 11</xref>
        ], we chose
to continuously host the MiGA competition this year. In the edition of 2024 this year 2, we
have organized a second round of the same two tasks (classification and online recognition)
including more modalities (both RGB and skeleton). Two public MG datasets (iMiGUE and
SMG) [
        <xref ref-type="bibr" rid="ref12 ref13 ref3">12, 13, 3</xref>
        ] are used.
      </p>
      <p>In this paper, we detail how the MiGA-IJCAI 2014 challenge was organized, the datasets, the
results achieved by 72 entrants who joined the competition, and the implementing schemes of
the winning methods.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Competition Tracks and Datasets</title>
      <p>In this section, we introduce the two challenge tracks and their corresponding characteristic, as
well as the datasets used in each track.</p>
      <sec id="sec-3-1">
        <title>2.1. Track 1: Multi-modal MG classification</title>
        <p>In this track, we focus on classifying micro-gestures (MGs) based on both skeleton data and RGB
data from pre-segmented short video clips. This track utilizes an in-the-wild MG dataset, the
iMiGUE dataset, which includes video footage of tennis players during post-match interviews,
featuring detailed ground-truth annotations for MGs. Unlike typical action or gesture datasets,
MGs capture finer, more subtle body movements that occur naturally in real-world interactions.
Key challenges in this classification task include learning these intricate movement patterns,</p>
        <sec id="sec-3-1-1">
          <title>1https://cv-ac.github.io/MiGA2023/</title>
          <p>2https://cv-ac.github.io/MiGA2/
managing the imbalanced distribution of MG samples, and distinguishing the high variability
in MGs across classes.</p>
          <p>For this year’s challenge, we also provide RGB data alongside the skeleton data and encourage
participants to explore multimodal approaches that integrate both modalities. The training and
testing sets are drawn from the iMiGUE dataset, following a cross-subject evaluation protocol:
72 subjects are split, with 37 subjects allocated for training and 35 for testing. Specifically, for
MG classification, 13,936 clips are designated for training, 3,692 clips for validation, and an
additional 4,563 clips are used for testing without annotations.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Track 2: Multi-modal MG online recognition</title>
        <p>In this track, we tackle online micro-gesture (MG) recognition by utilizing both skeleton data
and RGB data from extended video sequences. Unlike traditional online action or gesture
recognition datasets, where actions are typically well-structured and sequentially organized,
MG samples appear spontaneously and in varied sequences, resembling natural
communicative behaviors. Consequently, the task of online MG recognition requires handling complex
transitions between body movements, including the simultaneous occurrence of multiple MGs,
partial or incomplete MGs, and intricate transitions. Additionally, detecting subtle MGs amidst
other, less relevant body movements adds another layer of complexity, presenting challenges
not commonly addressed in previous gesture recognition research.</p>
        <p>The same as track 1, we encourage participants to adopt multimodal approaches, while
the SMG dataset serves as the foundation for this track. A cross-subject evaluation protocol
is applied, with sequences from 35 subjects allocated for training and sequences from the
remaining 5 subjects reserved for testing.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Competition Itinerary</title>
      <sec id="sec-4-1">
        <title>3.1. Competition agenda</title>
        <p>This section encompasses both the competition schedule and relevant participant details.
The challenge was managed using the Kaggle competition framework. The schedule of the
competition was as follows:
March 29, 2024. Call for Challenge online. Registration starts.</p>
        <p>April 9, 2024. Release of training data, development toolkit, and sample codes.
May 2, 2024. Release of testing data.</p>
        <p>May 12, 2024. Final testing data and result submission. Registration ends.</p>
        <p>May 17, 2024. Release of challenge results.</p>
        <p>May 30, 2024. Paper submission deadline (workshop).</p>
        <p>June 04, 2024. June 07, 2024. Notification to authors.</p>
        <p>June 04, 2024. June 12, 2024. Camera-ready deadline.</p>
        <p>August 03 – 09, 2024. MiGA IJCAI 2024 Workshop, Jeju, Korea.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Participants</title>
        <p>As stated, the competition has been conducted using Kaggle, a well-known challenge
opensource platform. We created a diferent competition for each track, having separate information
and leaderboard 3 4. A total of 72 users have been registered in the Kaggle platform, 43 for
track 1 and 29 for track 2 (note that some users might have been registered for more than one
track but we count each track severately). All these users were able to access the data for the
developing stage and submit their predictions for this stage. For the final evaluation stage, team
registration was mandatory, and a total of 16 teams were successfully registered: 12 for track 1,
and 4 for track 2. During the challenge period, in total 323 submissions were made with 235 for
track 1 and 88 for track 2.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Protocol and Evaluation</title>
      <p>In this section, we introduce the evaluation metrics used to evaluate the participants for the
two tracks.</p>
      <sec id="sec-5-1">
        <title>4.1. Multi-modal MG classification</title>
        <p>We evaluate participants’ methods based on Top-1 accuracy in the challenge, but we encourage
participants to report their results when submitting papers to our MiGA workshop on the
following subsets of the test set: 1) Overall: All segments in the test split; 2) Tail Classes: Due
to the long-tailed nature of the datasets, among a total of 33 classes in the iMiGUE dataset, 28
classes are tail classes (approx. 57.8 % of the data).</p>
        <p>As to the submission format for classification track, participants must submit their predictions
in a single .csv file. Instructions and sample submission files is released with the data. For each
‘Id’ in the validation set, they must predict a probability for the ‘Target’ variable. The file should
contain a header (Id, Target).</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. Multi-modal MG online recognition</title>
        <p>As for the MG online recognition track, we jointly evaluate the detection and classification
performances of algorithms by using the F1 score measurement defined below: F1
=2Precision*Recall/(Precision+Recall), given a long video sequence that needs to be evaluated, Precision
is the fraction of correctly classified MGs among all gestures retrieved in the sequence by
algorithms, while Recall (or sensitivity) is the fraction of MGs that have been correctly retrieved
over the total amount of annotated MGs.</p>
        <p>Considering the submission format for online recognition track, participants must submit
their predictions in a single .csv file. The submission .csv file should consist of the following
columns with headers: ID: incremental index, class: prediction label, start_frame: staring frame,
end_frame: ending frame, sample_id: represents the subject, i.e., the sample folder name is
Sample0005, the sample_id is 5.</p>
        <sec id="sec-5-2-1">
          <title>3https://www.kaggle.com/competitions/2nd-miga-ijcai-challenge-track1/ 4https://www.kaggle.com/competitions/2nd-miga-ijcai-challenge-track2/</title>
          <p>For both of the two tracks, the results are evaluated on the server and displayed on the
ranking list in real time. The organization team has the right to examine the participants’ source
code to ensure the reproducibility of the algorithms. The final results and ranking are confirmed
and announced by the organizers after verifying the reproducibility of the source code.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Challenge Results and Methods</title>
      <p>In this section, we report the winning methods proposed by the participants. For the two tracks,
we asked the top three teams to submit their source code and predictions for the test sets. Below,
we introduce the implementation details of each method.</p>
      <sec id="sec-6-1">
        <title>5.1. Track 1: Multi-modal MG classification</title>
        <p>Table 1 summarizes the methods of the three teams who ranked the top three on the test set of
track 1. One can see that all the methods use both RGB and skeleton modalities and tend to
convert skeleton modality data into heat map presentations with a PoseConv3D [14] backbone.
Next, we describe the main characteristics of the three winning methods.</p>
        <p>First place: The HFUT-VUT team proposes the first place scheme [ 15] with the core structure
of the proposed approach as two separate branches for RGB and Pose data. Initially, it employs
the PoseConv3D network [14] as its backbone, optimizing the learning of spatio-temporal
features and improving resilience against noise. Specifically, the method is built on a
twostream 3D CNN backbone, with the top pathway dedicated to RGB data processing and the
bottom pathway handling skeleton data. A cross-attention fusion module is then introduced
to capture interactions between the RGB and Pose modalities. Finally, drawing inspiration
from [16], a prototypical refinement module is added. This module establishes prototype
representations for each fine-grained micro-gesture class during training, prompting the model
to refine ambiguous samples across diferent gesture categories. By jointly leveraging the
PoseConv3D backbone for both the RGB and skeleton modalities, the model achieves 67.91%
accuracy on the skeleton-only modality and 70.25% accuracy when combining RGB and skeleton
data.</p>
        <p>Second place: The NPU-MUCIS team introduces a framework named M2HEN [17], which
constructs a heterogeneous ensemble network by combining two fundamentally diferent
deep learning models: a 3D convolution-based model and a Transformer-based model. This
heterogeneous ensemble approach enhances feature diversity and strengthens the model’s
representational power. For the 3D convolutional model, they present the MiG-enhanced
Multi-modal and Multi-scale 3D Convolutional sub-Network (M3CN), while the Ensemble
Hypergraph-Convolution Transformer (EHCT) [18] is employed as the Transformer model.
With either RGB or skeleton data alone, the framework achieves 61.14% and 61.11% accuracy,
respectively. When fusing both modalities, accuracy increases to 66.57% (baseline) and further
improves to 70.19% with the ensemble models and group training.</p>
        <p>Third place: The ywww11 team presents a method rooted in the CLIP framework. Building
on Froster CLIP [19], they introduce a token attenuation strategy within the video encoding
module, which incrementally filters out less significant tokens at each layer. For the skeleton
modality, they align it with the video modality’s CLIP model by applying text embeddings from
the video modality to the skeleton network. The PoseConv-3D model is specifically enhanced
with CLIP text embeddings [14], allowing it to work in conjunction with the CLIP text encoder.
This integration fosters collaborative processing. By fully leveraging feature extraction from
both the skeleton and video modalities, this approach boosts performance in micro-gesture
recognition tasks by using both skeleton sequences and video frames. By combining three
modalities (RGB, Skeleton Joint, Skeleton Limb) with optimized weights, they reach an accuracy
of 68.90%.</p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Track 2: Multi-modal MG online recognition</title>
        <p>Second place: The approach introduced by the team HFUT-VUT [24] is composed of a
video encoder and an action decoder. For each video sequence, the model learns a series of
trainable query points that help identify action boundary positions, along with query vectors
that interpret action semantics and locations from the input features. The action decoder,
incorporating a Mamba-MHSA module and a multi-level interaction module, then maps features
to linear projection layers, which decode action labels from the query vectors and convert the
query points into detection outputs. This method achieves F1 scores of 0.1835 for RGB-only,
0.2269 for skeleton-only, and 0.2757 when combining both RGB and skeleton data. With a single
RGB modality, it achieves an F1 score of 0.1434.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Discussion</title>
      <p>This paper has described the main characteristics of the MiGA 2024 Challenge hosted at IJCAI
2024 which included tracks on (i) Multi-modal MG classification, and (ii) Multi-modal MG online
recognition. Two large datasets (the SMG and iMiGUE datasets) were introduced and made
publicly available with corresponding toolkits to the participants for a fair comparison of the
performance results.</p>
      <p>Analyzing the methods introduced by the above participants, several conclusions can be
drawn. For multi-modal MG classification (track 1), although a large improvement has been
made in the performances from this year’s method compared to last year’s models [25, 18], there
is still considerable room for improvement (the current highest accuracy is 70.25%). On the other
hand, there are still many ways to improve in the MG online recognition task from multi-modal
data. For instance, the detection performances are still quite low with the highest F1 score
as 0.2757, showing that spotting MG is a challenging task to perform even by humans. Aside
from those winning schemes proposed for the MiGA competition 2024, some other interesting
research related to MG is also included in the MiGA 2024 workshop. For instance, Xia et al.
propose to use event data to recognize micro-gestures and micro-expressions which is a novel
research entry that meets the nature of those short and rapid movements of human behaviors
[26].</p>
      <p>Future trends in MiGA may include hidden emotion understanding via MGs with the analysis
of social signals, and face expression analysis as relevant information cues. Besides, extending
the datasets to larger scales and diverse materials toward more real-world scenarios can also be
a promising direction.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>We wish to acknowledge the teams who participated in MiGA 2024. We especially thank
Associate Professor Xiaobai Li for assisting the organizing the event, and Marko Savic, Atif
Shah, and Abdelrahman Mostafa for the technical support. We also wish to thank the platforms
Codalab and Kaggle for providing free sources for us to organize the challenges of MiGA 2023
and 2024.</p>
      <p>This MiGA challenge and workshop was supported by the Academy of Finland for Academy
Professor project EmotionAI (grants 336116, 345122), the University of Oulu &amp; The Academy of
Finland Profi 7 Hybrid Intelligence (grant 352788), by the Ministry of Education and Culture of
Finland for AI forum project. We also wish to acknowledge the CSC – IT Center for Science,
Finland, for computational resources.
IEEE International Conference on Automatic Face &amp; Gesture Recognition (FG 2019), IEEE,
2019, pp. 1–8.
[14] H. Duan, Y. Zhao, K. Chen, D. Lin, B. Dai, Revisiting skeleton-based action recognition, in:
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
(CVPR), 2022, pp. 2969–2978.
[15] G. Chen, F. Wang, K. Li, Z. Wu, H. Fan, Y. Yang, M. Wang, D. Guo, Prototype learning for
micro-gesture classification, in: MiGA@ IJCAI, 2024.
[16] H. Zhou, Q. Liu, Y. Wang, Learning discriminative representations for skeleton based
action recognition, in: Proceedings of the IEEE/CVF Conference on Computer Vision and
Pattern Recognition, 2023, pp. 10608–10617.
[17] H. Huang, Y. Wang, L. Kerui, Z. Xia, Multi-modal micro-gesture classification via
multiscale heterogeneous ensemble network, in: MiGA@ IJCAI, 2024.
[18] H. Huang, X. Guo, W. Peng, Z. Xia, Micro-gesture classification based on ensemble
hypergraph-convolution transformer., in: MiGA@ IJCAI, 2023.
[19] X. Huang, H. Zhou, K. Yao, K. Han, Froster: Frozen clip is a strong teacher for
openvocabulary action recognition, ICLR2024 (2024).
[20] H. Chen, X. Liu, J. Shi, G. Zhao, Temporal hierarchical dictionary guided decoding for
online gesture segmentation and recognition, IEEE Transactions on Image Processing 29
(2020) 9689–9702.
[21] H. Chen, X. Liu, G. Zhao, Temporal hierarchical dictionary with hmm for fast gesture
recognition, in: 2018 24th international conference on pattern recognition (ICPR), IEEE,
2018, pp. 3378–3383.
[22] Y. Wang, L. Kerui, H. Huang, Z. Xia, Micro-gesture online recognition with dual-stream
multi-scale transformer in long videos, in: MiGA@ IJCAI, 2024.
[23] X. Guo, X. Zhang, L. Li, Z. Xia, Micro-expression spotting with multi-scale local transformer
in long videos, Pattern Recognition Letters 168 (2023) 146–152.
[24] P. Liu, F. Wang, K. Li, G. Chen, Y. Wei, S. Tang, Z. Wu, D. Guo, Micro-gesture online
recognition using learnable query points, in: MiGA@ IJCAI, 2024.
[25] K. Li, D. Guo, G. Chen, X. Peng, M. Wang, Joint skeletal and semantic embedding loss for
micro-gesture classification, MiGA@ IJCAI (2023).
[26] K. Xia, L. Wei, L. Yu, A spatio-temporal event transformer on versatile tasks for human
behavior analysis, in: MiGA@ IJCAI, 2024.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>From emotion ai to cognitive ai</article-title>
          ,
          <source>International Journal of Network Dynamics and Intelligence</source>
          (
          <year>2022</year>
          )
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. B.</given-names>
            <surname>Adams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Shafir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tsachor</surname>
          </string-name>
          ,
          <article-title>Unlocking the emotional world of visual media: An overview of the science, research, and impact of understanding emotion</article-title>
          ,
          <source>Proceedings of the IEEE</source>
          <volume>111</volume>
          (
          <year>2023</year>
          )
          <fpage>1236</fpage>
          -
          <lpage>1286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zhao, Smg: A micro-gesture dataset towards spontaneous body gestures for emotional stress state analysis</article-title>
          ,
          <source>International Journal of Computer Vision</source>
          <volume>131</volume>
          (
          <year>2023</year>
          )
          <fpage>1346</fpage>
          -
          <lpage>1366</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Aviezer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Trope</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Todorov</surname>
          </string-name>
          ,
          <article-title>Body cues, not facial expressions, discriminate between intense positive and negative emotions</article-title>
          ,
          <source>Science</source>
          <volume>338</volume>
          (
          <year>2012</year>
          )
          <fpage>1225</fpage>
          -
          <lpage>1229</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ekman</surname>
          </string-name>
          , Darwin, deception, and
          <article-title>facial expression</article-title>
          ,
          <source>Annals of the new York Academy of sciences 1000</source>
          (
          <year>2003</year>
          )
          <fpage>205</fpage>
          -
          <lpage>221</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zhao, Searching multi-rate and multi-modal temporal enhanced networks for gesture recognition</article-title>
          ,
          <source>IEEE Transactions on Image Processing</source>
          <volume>30</volume>
          (
          <year>2021</year>
          )
          <fpage>5626</fpage>
          -
          <lpage>5640</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>W.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Learning graph convolutional network for skeletonbased human action recognition by neural searching</article-title>
          ,
          <source>in: Proceedings of the AAAI conference on artificial intelligence</source>
          , volume
          <volume>34</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>2669</fpage>
          -
          <lpage>2676</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Benchmarking micro-action recognition: Dataset, method, and application</article-title>
          ,
          <source>IEEE Transactions on Circuits and Systems for Video Technology</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <year>Mac 2024</year>
          :
          <article-title>Micro-action analysis grand challenge</article-title>
          ,
          <source>in: Proceedings of the 32nd ACM International Conference on Multimedia</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>11304</fpage>
          -
          <lpage>11305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . Zhao,
          <article-title>Representation learning for topology-adaptive micro-gesture recognition and analysis</article-title>
          ,
          <source>in: IJCAI-MIGA Workshop &amp; Challenge on Micro-gesture Analysis for Hidden Emotion Understanding (MiGA) July</source>
          <volume>21</volume>
          ,
          <year>2023</year>
          Macao, China, Redaktion
          <string-name>
            <surname>Sun</surname>
            <given-names>SITE</given-names>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . Zhao,
          <article-title>Naive data augmentation might be toxic: Data-prior guided self-supervised representation learning for micro-gesture recognition</article-title>
          ,
          <source>in: 2024 IEEE 18th International Conference on Automatic Face and Gesture Recognition (FG)</source>
          , IEEE,
          <year>2024</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Zhao, imigue: An identity-free video dataset for micro-gesture understanding and emotion analysis</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10631</fpage>
          -
          <lpage>10642</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <surname>G</surname>
          </string-name>
          . Zhao,
          <article-title>Analyze spontaneous gestures for emotional stress state recognition: A micro-gesture dataset and analysis with deep learning</article-title>
          ,
          <source>in: 2019 14th</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>