Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决设备上音频理解的资源限制问题,研究通过特定架构和三阶段训练方法构建了159.3M参数的音频-语言模型Mizar。
📝 Abstract
Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M parameters. We introduce a recipe that brings together architecture, data, and three-stage training to build Mizar, a 159.3M-parameter ALM. Its architecture connects a compact CED-Small audio encoder to SmolLM2-135M through a frequency-merging mapper. With supervision drawn from ReasonAQA, AudioMCQ, and AVQA, the model undergoes three training stages: audio-language alignment (Stage 1), audio-dependent fine-tuning (Stage 2), and post-training (Stage 3) aimed at strengthening weak skills while retaining learned capabilities. Across five random seeds, Mizar achieves mean accuracies of 52.92% on MMAU, 42.42% on MMAR, and 36.02% on ADQA-clean, surpassing the previous best-performing ALM below 200M parameters on all three benchmarks. It also supports local inference on a single CPU: on questions from the MMAU benchmark, the mean latency from opening the audio file to generating a complete answer is 1.09 seconds. Code and checkpoints are available at https://github.com/KaiyangLi1992/Mizar_159M.
Problem

Research questions and friction points this paper is trying to address.

Audio-Language Model
Limited Memory
Computation
Contextual Understanding
Auditory Events
Innovation

Methods, ideas, or system contributions that make the work stand out.

audio-language model
three-stage training
compact architecture
local inference
low latency
🔎 Similar Papers
No similar papers found.