ReSCUE: Re-translation with Sentence Commitment for Unsegmented Long-Form Simultaneous Sign Language Translation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of existing sign language translation systems on pre-segmentation and their difficulty in processing unsegmented long video streams. To overcome these limitations, this work proposes a unified simultaneous translation framework tailored for real-world streaming scenarios. The framework integrates inference-aware training, a stable retranslation strategy, and an online sentence commitment mechanism to achieve precise online segmentation and low-latency output correction. Experimental results demonstrate that the proposed method attains state-of-the-art translation quality under strict latency constraints, with performance approaching the upper bound of offline systems. These findings validate the feasibility of deploying the framework for real-time sign language translation applications.
📝 Abstract
Simultaneous Sign Language Translation (SLT) is critical for real-time communication, yet existing methods remain largely confined to sentence-level, offline settings that assume pre-segmented inputs. These assumptions hinder deployment in realistic scenarios involving continuous, unsegmented video streams. We present ReSCUE, a unified framework for simultaneous SLT on unsegmented long-form sign language videos that aligns training and inference with realistic streaming conditions. ReSCUE combines inference-aware training to handle partial inputs, non-signing pauses, and multi-sentence contexts, stabilized re-translation to enable low-latency yet revisable predictions with reduced output flicker, and a sentence commitment mechanism for online segmentation and memory management. Experiments on standard sentence-level benchmarks show that ReSCUE achieves lower latency and the best translation quality under low-latency settings. On long-form unsegmented datasets, ReSCUE approaches the translation quality of oracle offline systems that use ground-truth sentence boundaries, while operating at substantially lower latency, demonstrating its practicality for real-world streaming scenarios.
Problem

Research questions and friction points this paper is trying to address.

Simultaneous Sign Language Translation
Unsegmented Long-Form Video
Real-time Streaming
Low Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simultaneous Sign Language Translation
Unsegmented Long-Form Video
Re-translation
Sentence Commitment
Inference-Aware Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
S
Sihan Ren
ShanghaiTech University
G
Gaozheng Li
ShanghaiTech University
Y
Yuanshang Quan
ShanghaiTech University
Yiming Qin
Yiming Qin
EPFL, LTS4
generative modelsgraph machine learningdrug discovery
F
Fuyi Yang
ShanghaiTech University
C
Chang Liu
ShanghaiTech University
L
Lan Xu
ShanghaiTech University
Minye Wu
Minye Wu
KU Leuven
computer vision