🤖 AI Summary
This study addresses the latency-accuracy trade-off in streaming speech translation by proposing cross-attention-based adaptive strategies, namely RFAP and DCAP. This work pioneers the use of attention mechanisms to dynamically regulate information flow, enabling existing offline encoder-decoder models to be directly adapted to streaming scenarios without retraining. Experimental results demonstrate that the proposed approach maintains high-quality outputs under low-latency conditions, achieving a 4.0 BLEU score improvement while reducing latency by nearly one second. Consequently, this method realizes a significant synergistic optimization of both performance and efficiency with zero additional training cost.
📝 Abstract
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.