π€ AI Summary
This study addresses the insufficient cross-modal alignment in sign language translation and generation tasks, which stems from their isolated treatment and limited data availability. To this end, we propose a unified framework based on large language models (LLMs) that employs a symmetric network architecture. Specifically, we introduce a novel staged progressive alignment strategy: leveraging large-scale pretraining to achieve preliminary textβsign alignment for optimizing translation, and subsequently enhancing generation through translation-derived representations, thereby effectively breaking down task isolation barriers. Experimental results demonstrate that the proposed approach achieves performance comparable to task-specific models across multiple benchmarks while exhibiting superior transferability to sign language recognition.
π Abstract
Sign language translation and generation share the goal of bidirectional alignment between text and sign representations. However, existing approaches either treat them as isolated tasks or are only verified on limited datasets, limiting effective modeling between modalities. In this paper, we propose SignFLIP, a unified LLM-centered framework for translation and generation. To enable bidirectional mapping between text and sign, SignFLIP adopts a symmetric architecture together with a stage-wise training strategy built on large-scale data. The shared sign--text representation is progressively refined: pre-alignment facilitates subsequent SLT, while the SLT-adapted representation further benefits SLG. Extensive experiments on multiple benchmarks show that SignFLIP shows competitive performance compared with task-specific models on both translation and generation tasks, as well as strong transferability to sign language recognition.