UltraDub: Towards Authentic Dubbing by Unifying Visually-Steered Flow Learning and Trajectory Guidance

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the vulnerability of temporal and speaker cues to interference in visual voice cloning, as well as lip-sync desynchronization caused by imbalanced inference guidance. To overcome these issues, this work proposes a unified framework integrating visual-guided flow learning with trajectory guidance. Core innovations include an MDR module that dynamically calibrates linguistic and style retrieval, and a training-free RTG mechanism that reinforces semantics at prediction midpoints while preserving temporal alignment. Furthermore, the approach combines multimodal conditional flow learning, independent temporal gating, and hierarchical correction techniques, alongside the construction of the DiverseDub benchmark dataset. Experimental results demonstrate that the proposed method achieves state-of-the-art performance across four datasets, significantly enhancing dubbing intelligibility, speaker consistency, and lip-sync accuracy.
📝 Abstract
Visual voice cloning requires intelligible, speaker-consistent speech synchronized with visible articulation. However, sequential multimodal conditioning can disrupt previously established temporal and speaker cues, while imbalanced inference guidance can improve linguistic accuracy at the expense of lip synchronization. In this paper, we propose UltraDub, a Unifying Visually-Steered Flow learning and trajectory Guidance Dubbing framework that leverages vision in two ways: as continuous motion for multimodal context aggregation, and as structural rhythm for trajectory rectification. Specifically, we introduce the Motion-guided Dual-context Retrieving (MDR) module, which continually recalibrates linguistic and speaker-style retrieval through shared lip-motion query residuals, utilizing independent time-conditioned gates to regulate their contributions. Furthermore, we propose Rhythm-anchored Trajectory Guidance (RTG), a training-free mechanism that evaluates hierarchical multimodal corrections at a visual-only predictive midpoint, safely strengthening semantic conditioning while better preserving temporal alignment. Finally, we construct DiverseDub, a multi-scenario benchmark to evaluate video dubbing in the wild. Extensive experiments demonstrate that UltraDub achieves state-of-the-art performance across four datasets.
Problem

Research questions and friction points this paper is trying to address.

Visual voice cloning
Video dubbing
Lip synchronization
Multimodal conditioning
Speaker consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Voice Cloning
Flow Matching
Trajectory Guidance
Multimodal Conditioning
Video Dubbing
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
G
Gaoxiang Cong
Institute of Computing Technology, Chinese Academy of Sciences
Liang Li
Liang Li
Institue of Computing Technology, CAS
Computer VisionImage UnderstandingMultimedia Content Analysis
J
Jianwei Wen
Institute of Computing Technology, Chinese Academy of Sciences
Z
Zhedong Zhang
Hangzhou Dianzi University
Z
Zheng-Jun Zha
University of Science and Technology of China
Qingming Huang
Qingming Huang
University of the Chinese Academy of Sciences
Multimedia Analysis and RetrievalImage and Video ProcessingPattern RecognitionComputer VisionVideo Coding