Score
Designs and implements trainable audio-side prompt modules and feature-wise modulation operators that adapt pretrained encoder representations by rescaling, shifting, or amplifying channel-wise features (FiLM-style adaptive feature scaling, frequency- and noise-aware modulation, stage-wise or decompose–enhance–reconstruct pipelines) and by fusing conditional audio and text prompts across modalities and utterances. These components are built and analyzed to refine task-specific acoustic representations, mitigate mismatch between pretrained encoders, and enable few‑shot or transfer adaptation.
This work addresses the limitation of existing audio language models in few-shot learning, which typically optimize only textual prompts while overlooking the potential of learnable prompts within the audio encoder. The study introduces trainable prompts directly into the audio encoder for the first time, employing a staged modulation mechanism to explicitly guide the learning of task-relevant acoustic features. These audio-side prompts are jointly optimized with textual prompts, forming a complementary bilateral prompting architecture. Implemented in a plug-and-play manner, the proposed method is seamlessly integrated into existing models and achieves substantial improvements in few-shot performance across eleven audio classification benchmarks, demonstrating the effectiveness and generalizability of audio-side prompting in modulating the representation space.
To address the limitation of single-audio-encoder architectures in speech large language models (LLMs)—which struggle to simultaneously optimize semantic understanding tasks (e.g., automatic speech recognition, audio captioning) and acoustic modeling tasks (e.g., speaker count verification)—this paper proposes Prompt-aware Multi-encoder Mixing (PaM). PaM employs a prompt-driven gating mechanism to dynamically route input audio to the most task-appropriate encoder and introduces a task-aware feature fusion strategy—replacing naive concatenation or averaging—to integrate heterogeneous encoder outputs. Unlike prior approaches, PaM enables a single unified speech LLM to achieve state-of-the-art performance across diverse downstream tasks—including ASR, speaker count verification, and audio captioning—outperforming all single-encoder baselines and conventional feature fusion methods. This work marks the first demonstration of a unified architecture attaining optimal performance across such heterogeneous speech tasks, thereby overcoming the fundamental constraints of the traditional single-encoder paradigm.
To address the lack of native audio understanding and generation capabilities in large language models (LLMs), this paper proposes an efficient audio discretization framework that compresses high-sample-rate continuous audio into ultra-low-bitrate (0.23 kbps) discrete acoustic tokens, jointly modeled with text tokens. Methodologically, it introduces the first audio representation learning paradigm integrating vector quantization (VQ) with conditional flow matching (CFM), and achieves evaluable audio understanding by fine-tuning a text-only LLM exclusively via LoRA. Experiments demonstrate that the proposed acoustic tokens outperform VQ-VAE on multiple acoustic event classification tasks and match state-of-the-art (SOTA) models in understanding performance. However, generative quality remains limited by dataset scale and evaluation protocols, highlighting key bottlenecks in current audio–language joint modeling.
This work addresses the challenge of enabling audio encoders to efficiently convey semantic information to large language models (LLMs), where cross-modal information utilization remains suboptimal. To this end, we propose a novel “probe-based understanding” paradigm: an LLM actively interprets audio representations via dedicated attention submodules, augmented by delayed audio fusion and complementary multi-encoder integration. Built upon the Pengi/LLaVA architecture, our model is trained end-to-end using mechanistic interpretability analysis and a three-stage unified training strategy on 5.6 million audio–text pairs. Experimental results demonstrate consistent improvements of 10–60% over strong baselines across diverse audio understanding tasks. Crucially, this study provides the first systematic empirical validation of LLMs’ capacity to actively probe and interpret audio representations—demonstrating both efficacy and interpretability. Our approach establishes a new framework for audio–LLM co-modeling, advancing multimodal foundation models beyond passive feature aggregation.
Existing audio language models employ dense, shared-parameter adapters to process heterogeneous audio modalities—such as speech, music, and environmental sounds—which often suffer from gradient conflicts that limit performance. To address this, this work proposes MoE-Adapter, the first audio adapter architecture based on sparse mixture-of-experts (MoE). It dynamically routes audio tokens to specialized experts via a gating mechanism to disentangle acoustic features, while retaining shared experts to preserve global contextual information. Under comparable computational costs, MoE-Adapter significantly outperforms dense baselines across both audio semantic understanding and paralinguistic tasks. The authors release code and models to support further research in this direction.
This work addresses the susceptibility of large audio language models to hallucinations caused by linguistic priors overpowering acoustic evidence. To mitigate this, the authors propose a task- and sample-adaptive perturbation selection mechanism within a contrastive decoding framework. Leveraging a structured audio perturbation bank spanning temporal, spectral, frequency, and amplitude domains, the method dynamically selects optimal negative-sample perturbation strategies and employs a lightweight selector for efficient routing. The approach yields a 4.3% absolute improvement in accuracy on existence tasks and significantly boosts performance on temporal tasks from 74.7% to 81.4%. Furthermore, the study validates the efficacy of binary-constrained prompts, underscoring the critical role of adaptive perturbation strategies in alleviating hallucinations in audio language models.
Existing diffusion models struggle to generate high-quality audio efficiently under low-frequency, highly compressed conditions due to implicit coupling in intermediate representations. This work proposes ReGen, a novel framework that jointly models data and representations through multi-vector fields and enhances the generalization of conditional flow matching via Generalized Flow Matching (GFM). ReGen introduces a hierarchical multi-prompt representation mechanism, enabling high-fidelity speech and audio synthesis at extremely low sampling rates (6.25–12.5 Hz). Built upon a diffusion Transformer, a neural audio codec, and a latent diffusion architecture, ReGenVoice requires only four GPUs for one day of training and achieves a real-time factor (RTF) of 0.08 during inference, significantly improving word error rate (WER) for intelligibility and speaker similarity (SIM).