🤖 AI Summary
This study addresses the audio-visual misalignment caused by duration inconsistency in speech-to-speech translation by proposing a duration-aligned inference framework. The method leverages chain-of-thought reasoning to jointly plan target phrasing and speech duration prior to synthesizing speech tokens. To support this approach, we construct DuraSet-440K, a large-scale corpus, and introduce multimodal reinforcement learning coupled with a modality-aware reward attribution mechanism for optimization. Experimental results demonstrate that the proposed framework achieves an optimal balance between translation quality and duration alignment on the CVSS-T benchmark, significantly outperforming existing open-source and commercial baseline systems.
📝 Abstract
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explicit temporal planning, making duration control an unresolved challenge. We introduce DuraS2ST, a duration-aligned reasoning framework that enables a single speech language model to first generate an explicit chain-of-thought (CoT) for planning target wording and phonetic length, and then synthesize the corresponding speech tokens. To support this paradigm, we construct DuraSet-440K, a high-quality duration-aligned CoT corpus for supervised initialization. We further optimize the model with multi-modal multi-dimensional reinforcement learning, using a Duration Margin Reward to balance translation quality and duration consistency, and Modality-Aware Reward Attribution to assign rewards to appropriate token spans. Experiments on CVSS-T show that DuraS2ST achieves a strong balance between translation quality and duration consistency, outperforming competitive open-source and commercial baselines. Project page: https://github.com/Mia11939/DuraS2ST.