🤖 AI Summary
This study addresses the challenge of adapting existing gloss-free sign language translation paradigms to decoder-only large language models by proposing a unified translation framework. Methodologically, it achieves cross-modal alignment through hierarchical pretraining combined with target-domain retrieval bank construction. Furthermore, it innovatively introduces retrieval-utility-guided reinforcement fine-tuning, which effectively overcomes cross-modal optimization imbalance and mitigates harmful retrieval dependencies. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance across multiple benchmarks and, for the first time, comprehensively surpasses gloss-supervised methods on the CSL-Daily dataset, validating its superiority in efficient gloss-free sign language translation.
📝 Abstract
Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{https://github.com/shahelaojieraozhi/SignRAG}{GitHub}, together with models of different sizes to support future academic research.