DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the audio-visual misalignment caused by duration inconsistency in speech-to-speech translation by proposing a duration-aligned inference framework. The method leverages chain-of-thought reasoning to jointly plan target phrasing and speech duration prior to synthesizing speech tokens. To support this approach, we construct DuraSet-440K, a large-scale corpus, and introduce multimodal reinforcement learning coupled with a modality-aware reward attribution mechanism for optimization. Experimental results demonstrate that the proposed framework achieves an optimal balance between translation quality and duration alignment on the CVSS-T benchmark, significantly outperforming existing open-source and commercial baseline systems.
📝 Abstract
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.
Problem

Research questions and friction points this paper is trying to address.

Speech-to-Speech Translation
Duration Alignment
Video Dubbing
Temporal Planning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Speech-to-Speech Translation
Chain-of-Thought
Reinforcement Learning
Duration Alignment
Reward Attribution
🔎 Similar Papers
No similar papers found.