arXiv · 2608.00235
Attention-Steered Vision-Language Models for Sign Language Translation
Abstract
Vision-language models (VLMs) have emerged as a powerful framework for multimodal video understanding. However, they remain limited in the sign language translation task, where we identify a key failure mode of existing VLMbased translators: poor spatial-temporal visual grounding. In particular, we find that standard next-token cross-entropy does not directly provide signal for where and when the model should attend, causing models to overlook sign-relevant regions and frames. To address this challenge, we propose AttnSign, a VLM-based spatial-temporal attention steering framework for sign language translation. AttnSign first introduces spatial attention supervision for sign-relevant regions, such as face and hands, in each frame; then develops an RL-based motion-cadence steering method that encourages the model to explore and focus on sign-level keyframes. Experimental results on How2Sign and OpenASL benchmarks show that our proposed AttnSign consistently outperforms existing methods.
Explore related subjects
Keep this discovery
Meibo Hu, Guohao Sun, Annemarie D. Ross, Sheng Li, Zhiqiang Tao. 2026-07-31. Attention-Steered Vision-Language Models for Sign Language Translation. https://arxiv.org/abs/2608.00235
Cite the original work for its findings. Save a collection to share your selection of sources.