🤖 AI Summary
This study addresses the high latency inherent in existing sign language translation systems that require complete video inputs, which precludes real-time cross-lingual communication. To overcome this limitation, this work proposes the first simultaneous streaming sign language-to-sign language translation system. Methodologically, two Wait-k strategies and a computation-aware latency metric, ca-Stream-AL, are designed to accommodate cross-lingual word order discrepancies. Furthermore, random multi-path supervised training is integrated with Wait-k inference at test time to enable concurrent input processing and output generation. Experimental results demonstrate that the proposed system reduces the average ca-Stream-AL by 38% across six translation directions while maintaining high translation quality and low pose estimation error. These findings confirm its effectiveness in supporting real-time interactive scenarios, including broadcasting and bidirectional conversational applications.
📝 Abstract
Deaf and hard-of-hearing (DHH) signers cannot converse in real time across different sign languages today: existing sign-to-sign translation systems run offline, requiring the full source clip before any target sign is emitted. Live use cases - e.g. broadcast interpretation and two-way video calls - instead demand simultaneous output, while the source signer is still signing. We present, to our knowledge, the first simultaneous sign-to-sign (S2S) translation system, with two wait-k regimes: test-time wait-k inference applied directly to a full-sentence model, and a trained wait-k model via stochastic multi-path supervision. We further introduce ca-Stream-AL, a computation-aware latency metric for streaming output. Averaged across six S2S directions on both a smaller human-verified test set and a larger synthetic S2S corpus, our streaming system achieves a 38% ca-Stream-AL reduction while staying within a 9% DTW-PA-MPJPE increase and a 2.1 BLEU-4 drop compared to the full-sentence baseline. A word-order case study probes how the streaming model handles word order mismatch between different sign languages - a consequence of simultaneous translation.