Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of fixed modulation intensity and interaction conflicts caused by head pose variations in visual speech recognition. To this end, we propose a dynamic residual FiLM framework that incorporates a deep residual weighting mechanism to adaptively regulate the modulation strength of multi-path FiLM modules by predicting input-dependent weights, thereby effectively mitigating feature interference induced by unweighted modulation. Evaluated on the LRS2 and LRS3 benchmark datasets, the proposed framework significantly reduces phoneme error rates to 15.74% and 23.91%, respectively, outperforming existing baseline models. These results validate the effectiveness of the dynamic conditional modulation strategy in complex pose-varying scenarios.
📝 Abstract
Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.
Problem

Research questions and friction points this paper is trying to address.

Visual Speech Recognition
Head-pose variation
Feature-wise Linear Modulation
Dynamic modulation
Feature interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

Visual Speech Recognition
Dynamic FiLM Modulation
Head-Pose Variation
Feature-wise Linear Modulation
Dynamic Residual Weighting
M
Matthew Kit Khinn Teng
Kyushu Institute of Technology, Japan
H
Haibo Zhang
Kyushu Institute of Technology, Japan
T
Takeshi Saitoh
Kyushu Institute of Technology, Japan