Tailored Design of Audio-Visual Speech Recognition Models using Branchformers

📅 2024-07-09
🏛️ arXiv.org
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
To address the high computational complexity and poor interpretability of cross-modal interactions in audio-visual speech recognition (AVSR) under noisy conditions, this paper introduces Branchformer—the first application of this architecture to AVSR—proposing a novel two-stage, customized unified encoder-decoder framework: “unimodal-first, then fusion.” By incorporating modality-specific branch scoring and layer-level structural pruning, the method achieves parameter-efficient and interpretable audio-visual joint modeling. Evaluated on multi-scenario English and Spanish benchmarks, it attains word error rates (WER) of 2.5% and 9.1%, respectively—significantly outperforming comparable large models while reducing parameter count substantially and achieving state-of-the-art performance. Key contributions include: (1) the pioneering adaptation of Branchformer to AVSR; (2) a new architectural paradigm that jointly optimizes model lightweighting and cross-modal interpretability; and (3) enhanced end-to-end cross-modal collaborative modeling capability.

Technology Category

Computer Vision: Multi-modal VisionMachine Learning: Multimodal LearningNatural Language Processing: Speech

Application Category

Search and Retrieval-Augmented AI: Multilingual and cross-lingual Web searchUser Modeling, Personalization and Recommendation: Explainable and interpretable methods for personalizationWeb Mining and Content Analysis: Mining multimedia, multimodal, multilingual, cross-lingual Web data
📝 Abstract
Recent advances in Audio-Visual Speech Recognition (AVSR) have led to unprecedented achievements in the field, improving the robustness of this type of system in adverse, noisy environments. In most cases, this task has been addressed through the design of models composed of two independent encoders, each dedicated to a specific modality. However, while recent works have explored unified audio-visual encoders, determining the optimal cross-modal architecture remains an ongoing challenge. Furthermore, such approaches often rely on models comprising vast amounts of parameters and high computational cost training processes. In this paper, we aim to bridge this research gap by introducing a novel audio-visual framework. Our proposed method constitutes, to the best of our knowledge, the first attempt to harness the flexibility and interpretability offered by encoder architectures, such as the Branchformer, in the design of parameter-efficient AVSR systems. To be more precise, the proposed framework consists of two steps: first, estimating audio- and video-only systems, and then designing a tailored audio-visual unified encoder based on the layer-level branch scores provided by the modality-specific models. Extensive experiments on English and Spanish AVSR benchmarks covering multiple data conditions and scenarios demonstrated the effectiveness of our proposed method. Even when trained on a moderate scale of data, our models achieve competitive word error rates (WER) of approximately 2.5% for English and surpass existing approaches for Spanish, establishing a new benchmark with an average WER of around 9.1%. These results reflect how our tailored AVSR system is able to reach state-of-the-art recognition rates while significantly reducing the model complexity w.r.t. the prevalent approach in the field. Code and pre-trained models are available at https://github.com/david-gimeno/tailored-avsr.
Problem

Research questions and friction points this paper is trying to address.

Optimize cross-modal architecture for AVSR.
Reduce model complexity and computational cost.
Achieve state-of-the-art recognition rates efficiently.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Branchformer-based AVSR system
Tailored audio-visual encoder design
Parameter-efficient cross-modal architecture
💼 Related Jobs
No related jobs found.
Universitat Polit`ecnica de Val`encia
D
David Gimeno-G'omez
Pattern Recognition and Human Language Technology research center, Universitat Polit`ecnica de Val`encia, Camino de Vera, s/n, 46022, Val`encia, Spain
C
Carlos-D. Mart'inez-Hinarejos
Pattern Recognition and Human Language Technology research center, Universitat Polit`ecnica de Val`encia, Camino de Vera, s/n, 46022, Val`encia, Spain