ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对语音理解中的副语言和声学特征捕捉问题,通过构建包含22个副语言特性的框架及大规模数据集,并开发了两阶段训练的ParA-LLM模型来解决。
📝 Abstract
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristics and create a dataset of over 1.2M Audio-QA pairs. We develop ParA-LLM, trained with a two-stage curriculum: first on single-attribute questions to build foundational knowledge, then on multi-attribute questions for joint reasoning over speaker and acoustic characteristics. We also release ParA-Bench, a benchmark of 6,000 multiple-choice questions across speaker-speech, acoustic, and mixed categories, where frontier models like GPT-4o-Audio achieve only 36% accuracy. ParA-LLM surpasses state-of-the-art Audio LLMs like GPT-4o-Audio by 7.5% on ParA-Bench, with additional gains of 1.13% on MMAU-Pro Speech and 7.49% on MMAR Speech.
Problem

Research questions and friction points this paper is trying to address.

Paralinguistic
Acoustic Conditions
Audio LLMs
Innovation

Methods, ideas, or system contributions that make the work stand out.

ParA-LLM
paralinguistic characteristics
two-stage curriculum
ParA-Bench
🔎 Similar Papers