Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出了一种基于语音活动信号的语音合成方法,用于自动唇同步配音,以实现与源视频精确匹配的目标。
📝 Abstract
Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.
Problem

Research questions and friction points this paper is trying to address.

speech synthesis
lip-synchronous dubbing
voice-activity signal
Innovation

Methods, ideas, or system contributions that make the work stand out.

voice-activity signal
lip-synchronous dubbing
speech synthesis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
A
Alejandro Pérez-González-de-Martos
AppTek GmbH, Germany
Florian Lux
Florian Lux
Speech Technology Scientist, AppTek
Speech SynthesisNatural Language ProcessingMachine LearningArtificial Intelligence
A
Angelina Elizarova
AppTek GmbH, Germany
M
Milana Shkhanukova
AppTek GmbH, Germany
A
Andreas Kellner
AppTek GmbH, Germany
Mattia Antonino Di Gangi
Mattia Antonino Di Gangi
AppTek GmbH, Germany