arXiv · 2609.35225
SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale
Abstract
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Zhaoyi An, Sihan Tan, Youngbae Hwang, Kazuhiro Nakadai, Rei Kawakami. 2026-09-28. SignFLIP: A Unified Model for Sign Language Translation and Generation via Stage-wise Alignment at Scale. https://arxiv.org/abs/2609.35225
Cite the original work for its findings. Save a collection to share your selection of sources.