🤖 AI Summary
This study addresses the scarcity of annotated data and limited multilingual coverage in audio-visual speech recognition by proposing a synthetic data augmentation framework based on audio-driven digital humans. The method leverages audio-visual multimodal fusion and head synthesis techniques to generate lip-synchronized videos, which serve as standalone or supplementary training resources. Furthermore, it demonstrates that synthetic visual data can replace real annotations to enable baseline training for zero-resource languages. Experimental results indicate that the proposed approach significantly reduces the word error rate by 16.2%, offering a scalable data solution for cross-lingual audio-visual speech recognition.
📝 Abstract
Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability of labeled audiovisual (AV) datasets. This work explores the use of synthetic visual data as a solution, using an audio-driven talking-head pipeline to generate lip-synchronized visual content from existing audio data. We evaluate the effectiveness of synthetic visual data both as an augmentation strategy and as a standalone training resource, applying our approach to Spanish and Catalan. Our results show that augmenting real AV data with synthetic samples yields relative Word Error Rate (WER) reductions of up to 16.2%, demonstrating the potential of this approach. Moreover, we demonstrate that synthetic data alone can serve as a baseline for AVSR training in languages lacking AV datasets. These findings provide evidence that synthetic visual data can serve as a scalable solution to AVSR data scarcity, enabling broader language coverage.