SignRAG: Unified Retrieval-Augmented Gloss-Free Sign Language Translation

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of adapting existing gloss-free sign language translation paradigms to decoder-only large language models by proposing a unified translation framework. Methodologically, it achieves cross-modal alignment through hierarchical pretraining combined with target-domain retrieval bank construction. Furthermore, it innovatively introduces retrieval-utility-guided reinforcement fine-tuning, which effectively overcomes cross-modal optimization imbalance and mitigates harmful retrieval dependencies. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance across multiple benchmarks and, for the first time, comprehensively surpasses gloss-supervised methods on the CSL-Daily dataset, validating its superiority in efficient gloss-free sign language translation.
📝 Abstract
Contemporary decoder-only large language models (LLMs) have demonstrated strong capabilities across a wide range of domains. However, existing pretraining paradigms for gloss-free sign language translation (SLT) are largely designed around conventional encoder-decoder pretrained language models, which limits their direct applicability to decoder-only LLMs. To address this limitation, we propose SignRAG, a unified framework combining hierarchical pretraining, target-domain retrieval augmentation, and retrieval-aware reinforcement fine-tuning. Hierarchical pretraining first learns linguistically grounded sign representations and then jointly aligns the sign encoder with an LLM, mitigating cross-modal optimization imbalance. For downstream adaptation, SignRAG complements parameter-based fine-tuning with a target-domain retrieval gallery that provides instance-specific translation cues. To ensure that retrieved contexts are used appropriately, we further introduce Retrieval Utility-Guided Reinforcement Fine-Tuning (RUG-RFT), which combines translation-quality and retrieval-utility rewards to encourage beneficial retrieval use while suppressing harmful reliance. Experiments on multiple SLT benchmarks establish new state-of-the-art performance. In particular, to the best of our knowledge, SignRAG is the first gloss-free approach to outperform gloss-supervised methods across all reported metrics on CSL-Daily. Our code has been released at \href{https://github.com/shahelaojieraozhi/SignRAG}{GitHub}, together with models of different sizes to support future academic research.
Problem

Research questions and friction points this paper is trying to address.

Sign Language Translation
Gloss-Free
Decoder-Only LLMs
Retrieval-Augmented Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Sign Language Translation
Retrieval-Augmented Generation
Decoder-only LLMs
Reinforcement Fine-Tuning
Gloss-Free
Z
Zhi Rao
Faculty of Innovation Engineering, Macau University of Science and Technology, Macau, China
Yucheng Zhou
Yucheng Zhou
University of Macau | Fudan
Machine LearningLarge Language ModelsDeep Generative ModelsMultimodal LearningAI Healthcare
Q
Qianran Sun
MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China
Y
Yiqing Huang
Faculty of Innovation Engineering, Macau University of Science and Technology, Macau, China
L
Longcan Yuan
Faculty of Innovation Engineering, Macau University of Science and Technology, Macau, China
J
Jiayi Hou
Yale University, New Haven, CT, USA
C
Chengwen Yao
VIVO AI Lab, China
L
Lin Cheng
VIVO AI Lab, China
D
Donghui Sun
VIVO AI Lab, China
Xiaoxin Chen
Xiaoxin Chen
Coriell Institute for Medical Research
J
Jun Wan
MAIS, Institute of Automation, Chinese Academy of Sciences, Beijing, China