GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limited robustness of keyword recognition in noisy environments caused by the underutilization of visual streams. We propose GIVE, a fusion mechanism that injects lip movement evidence into audio queries via gated cross-attention and incorporates phoneme-level visual representations to enhance noise resilience. Our analysis reveals that visual robustness depends on two conditions: phoneme-level representations and evidence injection, with the latter outperforming feature rescaling. Evaluated on a trimodal Query-by-Example Keyword Spotting (QbyE-KWS) task, the proposed method reduces the Equal Error Rate (EER) by 72.9% at −10 dB for unseen keywords and achieves an average EER reduction of 62.8%, yielding an equivalent signal-to-noise ratio gain of 4–9.3 dB.
📝 Abstract
Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder's lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bearing encoder, injection yields an effective SNR gain of 4.0-9.3 dB over masking at -10 dB, whereas under a phoneme-poor one it nearly vanishes. Relative to the benchmark system, GIVE-KWS reduces unseen-keyword EER by 72.9% at -10 dB and 62.8% on average.
Problem

Research questions and friction points this paper is trying to address.

Query-by-Example Keyword Spotting
Noise-Robustness
Visual Speech
Multi-modal Fusion
Phonemic Information
Innovation

Methods, ideas, or system contributions that make the work stand out.

Query-by-Example Keyword Spotting
Gated Cross-Attention
Visual Evidence Injection
Noise-Robustness
Phoneme-bearing Visual Representation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
M
Ming-Hsiang Hu
Department of Computer Science and Information Engineering, National Taiwan Normal University
K
Kuan-Tang Huang
Department of Computer Science and Information Engineering, National Taiwan Normal University
Hung-Shin Lee
Hung-Shin Lee
North Co., Ltd., Taiwan
Speech Processing
Berlin Chen
Berlin Chen
Professor of Computer Science and Information Engineering, National Taiwan Normal University
speech and natural language processingcomputer-assisted language learningmachine learning