All In Good Time: Causality-Aware Framework for LLM-Based Simultaneous Speech-to-Speech Translation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of causal alignment difficulties and high latency induced by fixed policies in large language model-based simultaneous speech translation. To this end, it proposes a causality-aware framework comprising three key components: a novel data pipeline that generates high-fidelity segments to mitigate data scarcity, a factorized FAST architecture that decouples translation from synthesis, and CAP, an adaptive policy that dynamically optimizes read/write decisions. Furthermore, this work introduces a causality-aware latency metric, significantly reducing training data dependence while enhancing speaker fidelity. Multilingual experiments demonstrate that the proposed approach achieves state-of-the-art performance, yielding a 1.2 BLEU improvement and a 26%–38.8% reduction in relative latency, thereby realizing an optimal trade-off between translation quality and latency.
📝 Abstract
Large Language Models (LLMs) have shown strong performance in low-resource offline translation; however, extending them to simultaneous speech-to-speech translation (Simul-S2ST) remains challenging due to the scarcity of causally aligned training data with high cross-lingual speaker fidelity. In addition, existing approaches rely on fixed translation policy or confidence heuristics, leading to suboptimal quality and higher latency. We propose a causality-aware Simul-S2ST framework with a novel data pipeline that generates high-fidelity, causally aligned segments with improved voice transfer. The framework introduces (i) a factorized S2ST architecture (FAST), (ii) a causality-aware adaptive policy (CAP), and (iii) causality-aware latency metric. Experiments on CVSS Spanish, German, and French show that FAST-CAP consistently improves the quality-latency trade-off, achieving up to +1.2 BLEU and a 26% relative latency reduction over a fixed policy. Despite using substantially less training data than existing systems, FAST-CAP achieves state-of-the-art results in speech translation quality and speaker fidelity while yielding up to a 38.8% relative reduction in latency.
Problem

Research questions and friction points this paper is trying to address.

Simultaneous Speech-to-Speech Translation
Large Language Models
Causal Alignment
Speaker Fidelity
Translation Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Simultaneous Speech-to-Speech Translation
Causality-Aware Framework
Factorized Architecture
Adaptive Policy
Large Language Models
🔎 Similar Papers
2024-02-16International Workshop on Spoken Language TranslationCitations: 4