Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge in simultaneous speech translation of generating target text before the source utterance concludes while preserving translation quality. To this end, we propose an end-to-end approach that fine-tunes a full-sentence speech model using prefix supervision, enabling streaming translation without manual annotation. Furthermore, we design a multi-round incremental decoding strategy coupled with a confidence thresholding mechanism. By integrating forced prefix decoding, confidence calibration, and synthetic margin adjustment, our method substantially improves the quality-latency trade-off. Experimental results demonstrate that the proposed approach effectively enhances performance across multilingual benchmarks, with multi-round training reducing early commitment calibration errors by 80%.
📝 Abstract
Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts nor human translations. We compare single-turn forced-prefix and multi-turn append-only decoding, use a confidence threshold to control the inference-time quality--latency trade-off, and vary the density of training prefixes with a separate synthesis margin. On FLEURS and CoVoST2 in three language directions, prefix training improves quality--latency frontiers over the unadapted model, and confidence provides the broadest consistently competitive operating range. Multi-turn decoding is generally stronger at low latency; under multi-turn training, commit-calibration error falls by 63--68% overall and 68--80% at early prefixes, whereas single-turn training provides only modest overall calibration gains and no early-prefix improvement. A small synthesis margin sometimes extends the frontier to lower latency, particularly on shorter utterances, while a larger margin degrades translation quality and calibration. Prefix adaptation therefore improves simultaneous speech translation, especially under multi-turn append-only decoding, while synthesis density introduces a non-monotonic quality--latency trade-off.
Problem

Research questions and friction points this paper is trying to address.

Simultaneous Speech Translation
Partial Speech
Commit Strategy
Quality-Latency Trade-off
End-to-End
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simultaneous Speech Translation
Prefix Supervision
Multi-turn Decoding
Confidence Threshold
Commit Calibration
🔎 Similar Papers
No similar papers found.