train compact audio lm

Designs, trains, and optimizes compact (small, lightweight) audio language models (ALMs) by choosing efficient architectures and applying model-compression and runtime-optimization techniques so the audio sequence model fits strict size and latency constraints while maintaining task performance. Implements fine-tuning and evaluation workflows on target or low-resource audio datasets and analyzes trade-offs among accuracy, model size, and inference cost for deployment.

traincompactaudiolm

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.44
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This study addresses the challenge of efficiently allocating computational resources under a fixed budget to optimize speech model performance. The authors develop a unified framework to systematically investigate the joint impact of model scale, input duration, and representation resolution on automatic speech recognition (ASR) and speech emotion recognition (SER). Through large-scale scaling experiments, efficient LoRA-based fine-tuning, and multi-granularity computational cost modeling, they uncover nonlinear scaling laws across these dimensions and propose design principles for identifying optimal operating points that balance performance and efficiency. Key findings include diminishing returns from increasing model size, peak SER performance at approximately 4 seconds of audio input, and the ability to significantly reduce encoder resolution with less than 3% performance degradation while yielding substantial computational savings.

audio model scalingAutomatic Speech Recognitioncompute constraints

Continuous Audio Language Models

Sep 08, 2025
SR
Simon Rouard
🏛️ Kyutai | UMR STMS | IRCAM-CNRS Sorbonne Univ.

Existing audio-language models (ALMs) rely on discrete token sequences, constrained by low-bitrate, lossy audio codecs that compromise both audio fidelity and computational efficiency. To address this, we propose the Continuous Audio-Language Model (CALM), which abandons discrete tokenization and instead directly models continuous-time audio frames. CALM employs a large Transformer to encode contextual information, learns compact latent representations via a continuous-audio variational autoencoder (VAE), and utilizes a lightweight MLP decoder augmented with consistency modeling for efficient generation. By eliminating quantization-induced distortion, CALM achieves superior audio fidelity in speech and music synthesis while significantly reducing computational overhead and inference latency. Experiments demonstrate that CALM outperforms state-of-the-art discrete ALMs across multiple audio generation benchmarks, unifying high-quality output, low latency, and computational efficiency.

Addressing audio quality and computational cost trade-off in discrete audio tokensEnhancing efficiency and fidelity in speech and music generation modelsProposing continuous audio generation to avoid lossy compression limitations

Audio-Language Datasets of Scenes and Events: A Survey

Jul 09, 2024
GW
Gijs Wijngaard
🏛️ Maastricht University

This study systematically evaluates 69 audio-language datasets available as of September 2024, revealing pervasive issues including acoustic class imbalance, multi-source duplication, linguistic homogeneity (dominant English bias), restricted accessibility, and latent societal biases. Methodologically, we innovatively integrate PCA-based cross-dataset embedding variance analysis, CLAP-guided detection of modality leakage, joint acoustic–textual distribution modeling, and open governance practices to quantitatively identify systemic biases—particularly in widely used sources such as YouTube and Freesound. As a key contribution, we release an open resource library comprising over two million samples and propose a comprehensive Audio-Language Modeling (ALM) data curation roadmap that explicitly balances diversity, robustness, and fairness. This work establishes an empirically grounded, reproducible methodology for dataset development, directly supporting improved generalization capabilities of multimodal models.

Audio-lingual Model TrainingData Bias and LimitationsDataset Analysis

AudioBench: A Universal Benchmark for Audio Large Language Models

Jun 23, 2024
BW
Bin Wang
🏛️ Institute for Infocomm Research | A*STAR | Centre for Frontier AI Research

Existing Audio Large Language Models (AudioLLMs) lack a standardized, comprehensive benchmark for systematically evaluating instruction-following capabilities. Method: We introduce AudioBench—the first multidimensional evaluation benchmark specifically designed for AudioLLMs—covering three core task domains: speech understanding, acoustic scene understanding, and paralinguistic speech understanding. It integrates eight task categories across 26 datasets, including seven newly constructed ones. We formally define an instruction-following evaluation framework, propose a cross-modal instruction assessment protocol, and establish a unified metric system. Contribution/Results: We open-source the evaluation toolkit and a dynamic leaderboard. Comprehensive evaluation of five state-of-the-art AudioLLMs reveals significant capability imbalances across tasks. All data, code, and results are publicly released, establishing a new standard for AudioLLM capability assessment.

Lack of comprehensive benchmark for AudioLLMs' instruction followingNeed for evaluation metrics across diverse audio understanding tasksNo single model performs consistently well on all audio tasks

Latest Papers

What's happening recently
View more

This study addresses the inherent lack of audio understanding in large language models (LLMs) and the textual performance degradation and prohibitive costs associated with fine-tuning. To overcome these limitations, we propose a symbiotic architecture that bypasses the LLM backbone by directly writing audio features into the key-value cache via an audio-conditioned vector injector. This approach endows frozen LLMs with audio processing capabilities without modifying any model weights, structurally preserving the original textual representation space to prevent catastrophic forgetting while significantly reducing training overhead. Experimental results demonstrate that the proposed architecture surpasses existing frozen-parameter methods across multiple audio tasks, approaching the performance of full fine-tuning while perfectly retaining the LLM's original textual proficiency.

Audio Language ModelAudio UnderstandingFrozen LLM

This work addresses the challenge of slow convergence in multi-source heterogeneous Audio Question Answering (AudioQA) due to gradient conflicts arising from data heterogeneity during joint training. It introduces the first explicit modeling of inter-dataset heterogeneity by proposing a gradient affinity metric that eliminates the need for empirical transfer evaluation. Building upon this metric, the authors devise a Grouped Sequential Training (GST) strategy that leverages affinity-aware dataset grouping and a progressive scheduling protocol to enhance optimization efficiency while maintaining parallel training stability. The resulting model-agnostic framework achieves 30–40% faster convergence compared to standard parallel training across 14 diverse AudioQA datasets spanning speech, music, and environmental sounds, matching or exceeding the performance of full mixed-data training.

Audio Large Language Modelsconvergencedataset heterogeneity

This work addresses the limitations of current audio language models trained on reasoning-based large language models, which often produce unnatural responses due to their reliance on self-generated textual targets, hindering effective audio-language alignment. To overcome this, we propose a self-rephrasing mechanism that reformulates the model’s auto-generated responses into an audio-understanding format compatible with reasoning architectures, augmented by a compressed multi-audio encoder to enhance representational capacity. Leveraging a large-scale multitask audio-language corpus comprising 6 million samples, we efficiently train a 4B-parameter model. Our approach achieves state-of-the-art open-source performance on MMAU-speech and MMSU benchmarks, demonstrating strong audio reasoning and text capabilities at low computational cost—outperforming not only models of comparable size but also most larger counterparts.

audio-language alignmentauditory understandingchain-of-thought

This work addresses the high computational and memory costs incurred by long audio prefixes during inference in audio language models, a challenge exacerbated by existing training-free compression methods that struggle to preserve both content relevance and local acoustic context. The authors propose Local Temporal Bisection Merging (LTBM), a novel approach that introduces explicit temporal window constraints in encoder space and performs content-aware local compression based on similarity between neighboring tokens. LTBM establishes temporal locality as an effective inductive bias for audio token compression—a principle validated for the first time. By designing a global merging variant to disentangle the effect of locality, the study further reveals the task dependency of compression strategies: under strong compression, LTBM significantly improves audio captioning performance on AudioCaps, Clotho, and MMAU datasets, whereas global matching proves more suitable for multiple-choice audio understanding tasks.

audio-language modelscontext budgetinference efficiency

Hot Scholars

EE

Evangelos E. Papalexakis

Professor and Ross Family Chair, University of California Riverside
Data MiningTensor DecompositionGraph MiningSocial Media Mining