arXiv · 2606.07355
Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition
Abstract
Micro-gesture online recognition aims to temporally localize and classify subtle gestures in untrimmed videos. Owing to their extremely short duration, low motion amplitude, and ambiguous visual cues, capturing discriminative spatiotemporal representations remains highly challenging. Existing parameter-efficient adapters typically employ a single branch to model spatial and temporal cues jointly, which may fail to capture the fine-grained patterns of micro-gestures. To address this limitation, we propose a Spatial-Temporal Decoupled Adapter that decomposes video adaptation into independent temporal and spatial branches via lightweight depthwise convolutions. In addition, to alleviate the long-tailed class distribution inherent in the benchmark dataset, we introduce an Adaptive Soft Balanced Augmentation method, which dynamically allocates augmentation intensity based on class rarity and learning difficulty, without manual thresholds. Our method achieves an F1 score of 0.43808, ranking 1st in Track 2 of the 4th EI-MiGA-IJCAI Challenge.
Explore related subjects
Keep this discovery
Xucheng Shen, Kun Li, Fei Wang, Wei Qian, Jin Jiang, Dan Guo. 2026-08-29. Spatial-Temporal Decoupled Adapter for Micro-gesture Online Recognition. https://arxiv.org/abs/2606.07355
Cite the original work for its findings. Save a collection to share your selection of sources.
Discover connections
Connections use source metadata and explicit phrase matches, not verified experimental comparisons.