NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

📅 2026-09-21
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决工业环境中噪声对语音识别的影响,提出NAVIR系统,采用神经形态硬件上运行的视听融合方法,实现高效低能耗的语音识别。
📝 Abstract
Voice-controlled interaction in industrial settings is hampered by acoustic noise, which severely degrades audio-only speech recognition. Audio-visual speech recognition (AVSR) addresses this by fusing lip-motion cues with the audio stream, but state-of-the-art pipelines rely on three-dimensional convolutions, recurrent units, and attention modules that exceed the budget of typical edge devices. We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference. The pipeline factorises spatial and temporal encoding into separate AkidaNet-based modules: a per-frame visual encoder, a temporal video encoder, and a spectrogram audio encoder, fused by a lightweight predictor head and decoded by constrained beam search. Models are trained with connectionist temporal classification on noise-augmented audio and then fine-tuned with quantization-aware training. On the GRID benchmark, the quantized audio-visual model reaches 14.0% word error rate (WER) under noise on the unseen-speaker split and 3.3% WER on the overlapped-speaker split, against 22.5% and 11.8% for audio-only baselines, and it attains 98.6% command accuracy at 1.5% WER on a task-specific industrial-command corpus. Operation-count analysis indicates a 13-fold energy advantage of the spiking formulation over its artificial neural network counterpart at 27.6% mean firing rate. On-board measurements show roughly 5-fold lower energy per inference than a Raspberry Pi central processing unit on the lip-reading model, and over 100-fold lower than a laptop graphics processing unit, while sustaining 14.5 inferences per second. To the best of our knowledge, this is the first complete multimodal AVSR pipeline running on neuromorphic hardware of this class.
Problem

Research questions and friction points this paper is trying to address.

audio-visual speech recognition
industrial settings
edge devices
acoustic noise
lip-motion cues
Innovation

Methods, ideas, or system contributions that make the work stand out.

Neuromorphic Processor
Audio-Visual Speech Recognition (AVSR)
Edge Hardware
Energy Efficiency
Quantization-Aware Training
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Leonidas Delimpasis
Plaixus Ltd., Spyrou Patsi 62, 11855, Athens, Greece
Panagiota Moraiti
Panagiota Moraiti
Tech Hive Labs
RoboticsComputer VisionArtificial Intelligence
A
Antonis Porichis
AI Innovation Centre, University of Essex, Little Abington, CB21 6GP Cambridge, U.K.
P
Panos Chatzakos
Tech Hive Labs, 280 Kifisias Ave., 152 32 Halandri, Greece; AI Innovation Centre, University of Essex, Little Abington, CB21 6GP Cambridge, U.K.
M
Michail Karamousadakis
Plaixus Ltd., Spyrou Patsi 62, 11855, Athens, Greece