Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the latency-accuracy trade-off in streaming speech translation by proposing cross-attention-based adaptive strategies, namely RFAP and DCAP. This work pioneers the use of attention mechanisms to dynamically regulate information flow, enabling existing offline encoder-decoder models to be directly adapted to streaming scenarios without retraining. Experimental results demonstrate that the proposed approach maintains high-quality outputs under low-latency conditions, achieving a 4.0 BLEU score improvement while reducing latency by nearly one second. Consequently, this method realizes a significant synergistic optimization of both performance and efficiency with zero additional training cost.
📝 Abstract
Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional training. Results on three different language translation pairs over the CVSS-C corpus show that the RFAP is able to surpass other policies with gains of up to 4.0 BLEU while reducing the translation delay by almost 1 second. Moreover, the DCAP is able to preserve a high translation quality when the latency is very low.
Problem

Research questions and friction points this paper is trying to address.

Simultaneous Speech-to-Text Translation
Streaming
Translation Delay
Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simultaneous Speech-to-Text Translation
Attention-Based Policy
Cross-Attention Mechanism
Streaming Inference
Zero Additional Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
F
Filip Tăşădan
University of Copenhagen
E
Ema Tomanová
University of Copenhagen
O
Ondrej Lopuch
University of Copenhagen
P
Paweł Bilko
University of Copenhagen
Anders Søgaard
Anders Søgaard
Full Professor in NLP and Machine Learning, University of Copenhagen
Natural Language ProcessingMachine Learning.