Watch Your Speech: Text-aware Video-to-Speech Synthesis with Textual Conditioning

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the phoneme ambiguity problem in video-to-speech synthesis caused by insufficient visual information. To mitigate such visual ambiguity, this work proposes the WYS framework, which introduces textual conditioning as an explicit linguistic cue. Methodologically, an attention-based embedding fusion module is designed to integrate text and video sequences, combined with a conditional flow matching objective to optimize generation quality. Experimental results demonstrate that the proposed approach establishes new state-of-the-art performance in audio-visual synchronization on the LRS2 and LRS3 datasets. It maintains a low word error rate while achieving subjective evaluation scores approaching human-level naturalness.
📝 Abstract
Video-to-speech synthesis aims to generate natural-sounding speech from silent talking-face videos while ensuring phonetic accuracy. A fundamental challenge in this task is the inherent one-to-many mapping problem, where visual dynamics often lack sufficient information to uniquely determine the corresponding utterance. To address this, we propose Watch Your Speech (WYS), a video-to-speech synthesis framework that incorporates textual conditioning as an explicit linguistic cue to mitigate visual ambiguity. Our framework features an attention-based embedding fusion module that synergistically integrates textual context with video sequences, coupled with a conditional flow matching objective for high-fidelity speech generation. Extensive experiments on the LRS2 and LRS3 datasets demonstrate that WYS achieves superior performance, establishing new state-of-the-art results in audio-visual synchronization (LSE-C/D) while maintaining highly competitive textual accuracy (WER). Subjective evaluations further confirm that our model generates speech with near-human naturalness, validating the effectiveness of textual conditioning in content-controlled video-to-speech synthesis. Project page: https://github.com/gunwoo5034/Watch-your-Speech
Problem

Research questions and friction points this paper is trying to address.

Video-to-speech synthesis
one-to-many mapping
visual ambiguity
phonetic accuracy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Video-to-Speech Synthesis
Textual Conditioning
Attention-based Embedding Fusion
Conditional Flow Matching
Audio-Visual Synchronization
🔎 Similar Papers
G
Gunwoo Lee
Department of Information and Telecommunication Engineering, Soongsil University, Seoul, Republic of Korea
Y
Yoori Oh
Graduate School of Data Science, Seoul National University, Seoul, Republic of Korea
Yoseob Han
Yoseob Han
Assistant Professor at Soongsil University, School of Electronic Engineering
Deep Learning (DL)Compressed Sensing (CS)Parallel ComputingComputed Tomography (CT)Magnetic Resonance Image (MRI)