Score
Designs, trains, and optimizes compact (small, lightweight) audio language models (ALMs) by choosing efficient architectures and applying model-compression and runtime-optimization techniques so the audio sequence model fits strict size and latency constraints while maintaining task performance. Implements fine-tuning and evaluation workflows on target or low-resource audio datasets and analyzes trade-offs among accuracy, model size, and inference cost for deployment.
A systematic survey of Audio-Language Models (ALMs) for general-purpose audio tasks is currently lacking. Method: This paper introduces, for the first time, a six-dimensional comprehensive taxonomy—covering architectural design, pretraining paradigms, downstream adaptation strategies, benchmark datasets, evaluation protocols, and future challenges—and proposes a structured technical roadmap. Our methodology integrates multimodal representation learning, contrastive and generative pretraining, instruction tuning, multi-task collaborative optimization, and agent-based system design. Contribution/Results: The work fills a critical gap in the ALM literature by delivering the first authoritative, holistic survey; it provides researchers and practitioners with a rigorous technical reference and practical guidance, thereby significantly advancing human-like auditory modeling research and real-world applications.
为解决设备上音频理解的资源限制问题,研究通过特定架构和三阶段训练方法构建了159.3M参数的音频-语言模型Mizar。
This study addresses the challenge of efficiently allocating computational resources under a fixed budget to optimize speech model performance. The authors develop a unified framework to systematically investigate the joint impact of model scale, input duration, and representation resolution on automatic speech recognition (ASR) and speech emotion recognition (SER). Through large-scale scaling experiments, efficient LoRA-based fine-tuning, and multi-granularity computational cost modeling, they uncover nonlinear scaling laws across these dimensions and propose design principles for identifying optimal operating points that balance performance and efficiency. Key findings include diminishing returns from increasing model size, peak SER performance at approximately 4 seconds of audio input, and the ability to significantly reduce encoder resolution with less than 3% performance degradation while yielding substantial computational savings.
Existing audio-language models (ALMs) rely on discrete token sequences, constrained by low-bitrate, lossy audio codecs that compromise both audio fidelity and computational efficiency. To address this, we propose the Continuous Audio-Language Model (CALM), which abandons discrete tokenization and instead directly models continuous-time audio frames. CALM employs a large Transformer to encode contextual information, learns compact latent representations via a continuous-audio variational autoencoder (VAE), and utilizes a lightweight MLP decoder augmented with consistency modeling for efficient generation. By eliminating quantization-induced distortion, CALM achieves superior audio fidelity in speech and music synthesis while significantly reducing computational overhead and inference latency. Experiments demonstrate that CALM outperforms state-of-the-art discrete ALMs across multiple audio generation benchmarks, unifying high-quality output, low latency, and computational efficiency.
This study systematically evaluates 69 audio-language datasets available as of September 2024, revealing pervasive issues including acoustic class imbalance, multi-source duplication, linguistic homogeneity (dominant English bias), restricted accessibility, and latent societal biases. Methodologically, we innovatively integrate PCA-based cross-dataset embedding variance analysis, CLAP-guided detection of modality leakage, joint acoustic–textual distribution modeling, and open governance practices to quantitatively identify systemic biases—particularly in widely used sources such as YouTube and Freesound. As a key contribution, we release an open resource library comprising over two million samples and propose a comprehensive Audio-Language Modeling (ALM) data curation roadmap that explicitly balances diversity, robustness, and fairness. This work establishes an empirically grounded, reproducible methodology for dataset development, directly supporting improved generalization capabilities of multimodal models.
Existing Audio Large Language Models (AudioLLMs) lack a standardized, comprehensive benchmark for systematically evaluating instruction-following capabilities. Method: We introduce AudioBench—the first multidimensional evaluation benchmark specifically designed for AudioLLMs—covering three core task domains: speech understanding, acoustic scene understanding, and paralinguistic speech understanding. It integrates eight task categories across 26 datasets, including seven newly constructed ones. We formally define an instruction-following evaluation framework, propose a cross-modal instruction assessment protocol, and establish a unified metric system. Contribution/Results: We open-source the evaluation toolkit and a dynamic leaderboard. Comprehensive evaluation of five state-of-the-art AudioLLMs reveals significant capability imbalances across tasks. All data, code, and results are publicly released, establishing a new standard for AudioLLM capability assessment.
This study addresses the inherent lack of audio understanding in large language models (LLMs) and the textual performance degradation and prohibitive costs associated with fine-tuning. To overcome these limitations, we propose a symbiotic architecture that bypasses the LLM backbone by directly writing audio features into the key-value cache via an audio-conditioned vector injector. This approach endows frozen LLMs with audio processing capabilities without modifying any model weights, structurally preserving the original textual representation space to prevent catastrophic forgetting while significantly reducing training overhead. Experimental results demonstrate that the proposed architecture surpasses existing frozen-parameter methods across multiple audio tasks, approaching the performance of full fine-tuning while perfectly retaining the LLM's original textual proficiency.
This work addresses the challenge of slow convergence in multi-source heterogeneous Audio Question Answering (AudioQA) due to gradient conflicts arising from data heterogeneity during joint training. It introduces the first explicit modeling of inter-dataset heterogeneity by proposing a gradient affinity metric that eliminates the need for empirical transfer evaluation. Building upon this metric, the authors devise a Grouped Sequential Training (GST) strategy that leverages affinity-aware dataset grouping and a progressive scheduling protocol to enhance optimization efficiency while maintaining parallel training stability. The resulting model-agnostic framework achieves 30–40% faster convergence compared to standard parallel training across 14 diverse AudioQA datasets spanning speech, music, and environmental sounds, matching or exceeding the performance of full mixed-data training.
This work addresses the limitations of current audio language models trained on reasoning-based large language models, which often produce unnatural responses due to their reliance on self-generated textual targets, hindering effective audio-language alignment. To overcome this, we propose a self-rephrasing mechanism that reformulates the model’s auto-generated responses into an audio-understanding format compatible with reasoning architectures, augmented by a compressed multi-audio encoder to enhance representational capacity. Leveraging a large-scale multitask audio-language corpus comprising 6 million samples, we efficiently train a 4B-parameter model. Our approach achieves state-of-the-art open-source performance on MMAU-speech and MMSU benchmarks, demonstrating strong audio reasoning and text capabilities at low computational cost—outperforming not only models of comparable size but also most larger counterparts.
This work addresses the high computational and memory costs incurred by long audio prefixes during inference in audio language models, a challenge exacerbated by existing training-free compression methods that struggle to preserve both content relevance and local acoustic context. The authors propose Local Temporal Bisection Merging (LTBM), a novel approach that introduces explicit temporal window constraints in encoder space and performs content-aware local compression based on similarity between neighboring tokens. LTBM establishes temporal locality as an effective inductive bias for audio token compression—a principle validated for the first time. By designing a global merging variant to disentangle the effect of locality, the study further reveals the task dependency of compression strategies: under strong compression, LTBM significantly improves audio captioning performance on AudioCaps, Clotho, and MMAU datasets, whereas global matching proves more suitable for multiple-choice audio understanding tasks.
研究系统评估了连续和离散表示在语音、声音和音乐中的表现,揭示了语义约束在音频理解中的关键作用,并为未来LALMs的语义密度、保真度和效率平衡提供了指导。