audio prompt tuning

Designs and implements trainable audio-side prompt modules and feature-wise modulation operators that adapt pretrained encoder representations by rescaling, shifting, or amplifying channel-wise features (FiLM-style adaptive feature scaling, frequency- and noise-aware modulation, stage-wise or decompose–enhance–reconstruct pipelines) and by fusing conditional audio and text prompts across modalities and utterances. These components are built and analyzed to refine task-specific acoustic representations, mitigate mismatch between pretrained encoders, and enable few‑shot or transfer adaptation.

audioprompttuning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.51
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitation of existing audio language models in few-shot learning, which typically optimize only textual prompts while overlooking the potential of learnable prompts within the audio encoder. The study introduces trainable prompts directly into the audio encoder for the first time, employing a staged modulation mechanism to explicitly guide the learning of task-relevant acoustic features. These audio-side prompts are jointly optimized with textual prompts, forming a complementary bilateral prompting architecture. Implemented in a plug-and-play manner, the proposed method is seamlessly integrated into existing models and achieves substantial improvements in few-shot performance across eleven audio classification benchmarks, demonstrating the effectiveness and generalizability of audio-side prompting in modulating the representation space.

Acoustic FeaturesAudio EncoderAudio-Language Models

Enhancing Speech Large Language Models with Prompt-Aware Mixture of Audio Encoders

Feb 21, 2025
WS
Weiqiao Shan
🏛️ Northeastern University | Huawei | The Chinese University of Hong Kong | Harbin Engineering University | NiuTrans Research

To address the limitation of single-audio-encoder architectures in speech large language models (LLMs)—which struggle to simultaneously optimize semantic understanding tasks (e.g., automatic speech recognition, audio captioning) and acoustic modeling tasks (e.g., speaker count verification)—this paper proposes Prompt-aware Multi-encoder Mixing (PaM). PaM employs a prompt-driven gating mechanism to dynamically route input audio to the most task-appropriate encoder and introduces a task-aware feature fusion strategy—replacing naive concatenation or averaging—to integrate heterogeneous encoder outputs. Unlike prior approaches, PaM enables a single unified speech LLM to achieve state-of-the-art performance across diverse downstream tasks—including ASR, speaker count verification, and audio captioning—outperforming all single-encoder baselines and conventional feature fusion methods. This work marks the first demonstration of a unified architecture attaining optimal performance across such heterogeneous speech tasks, thereby overcoming the fundamental constraints of the traditional single-encoder paradigm.

Enhances task-specific audio feature extraction.Improves performance across diverse audio understanding tasks.Integrates multiple audio encoders with LLMs.

Make Some Noise: Towards LLM audio reasoning and generation using sound tokens

Mar 28, 2025
SM
Shivam Mehta
🏛️ KTH Royal Institute of Technology | Microsoft Research

To address the lack of native audio understanding and generation capabilities in large language models (LLMs), this paper proposes an efficient audio discretization framework that compresses high-sample-rate continuous audio into ultra-low-bitrate (0.23 kbps) discrete acoustic tokens, jointly modeled with text tokens. Methodologically, it introduces the first audio representation learning paradigm integrating vector quantization (VQ) with conditional flow matching (CFM), and achieves evaluable audio understanding by fine-tuning a text-only LLM exclusively via LoRA. Experiments demonstrate that the proposed acoustic tokens outperform VQ-VAE on multiple acoustic event classification tasks and match state-of-the-art (SOTA) models in understanding performance. However, generative quality remains limited by dataset scale and evaluation protocols, highlighting key bottlenecks in current audio–language joint modeling.

Achieving competitive audio comprehension with discrete tokensConverting audio into ultra-low bitrate discrete tokensIntegrating audio comprehension and generation into LLMs

PAL: Probing Audio Encoders via LLMs -- A Study of Information Transfer from Audio Encoders to LLMs

Jun 12, 2025
TA
Tony Alex
🏛️ University of Surrey | Mohamed bin Zayed University of AI

This work addresses the challenge of enabling audio encoders to efficiently convey semantic information to large language models (LLMs), where cross-modal information utilization remains suboptimal. To this end, we propose a novel “probe-based understanding” paradigm: an LLM actively interprets audio representations via dedicated attention submodules, augmented by delayed audio fusion and complementary multi-encoder integration. Built upon the Pengi/LLaVA architecture, our model is trained end-to-end using mechanistic interpretability analysis and a three-stage unified training strategy on 5.6 million audio–text pairs. Experimental results demonstrate consistent improvements of 10–60% over strong baselines across diverse audio understanding tasks. Crucially, this study provides the first systematic empirical validation of LLMs’ capacity to actively probe and interpret audio representations—demonstrating both efficacy and interpretability. Our approach establishes a new framework for audio–LLM co-modeling, advancing multimodal foundation models beyond passive feature aggregation.

Enhance cross-modal information transfer efficiencyOptimize architectural design for audio-LLM interactionStudy mechanisms of audio representation transfer to LLMs

Existing audio language models employ dense, shared-parameter adapters to process heterogeneous audio modalities—such as speech, music, and environmental sounds—which often suffer from gradient conflicts that limit performance. To address this, this work proposes MoE-Adapter, the first audio adapter architecture based on sparse mixture-of-experts (MoE). It dynamically routes audio tokens to specialized experts via a gating mechanism to disentangle acoustic features, while retaining shared experts to preserve global contextual information. Under comparable computational costs, MoE-Adapter significantly outperforms dense baselines across both audio semantic understanding and paralinguistic tasks. The authors release code and models to support further research in this direction.

audio modalitygradient conflictheterogeneous acoustic information

Latest Papers

What's happening recently
View more

This work addresses the susceptibility of large audio language models to hallucinations caused by linguistic priors overpowering acoustic evidence. To mitigate this, the authors propose a task- and sample-adaptive perturbation selection mechanism within a contrastive decoding framework. Leveraging a structured audio perturbation bank spanning temporal, spectral, frequency, and amplitude domains, the method dynamically selects optimal negative-sample perturbation strategies and employs a lightweight selector for efficient routing. The approach yields a 4.3% absolute improvement in accuracy on existence tasks and significantly boosts performance on temporal tasks from 74.7% to 81.4%. Furthermore, the study validates the efficacy of binary-constrained prompts, underscoring the critical role of adaptive perturbation strategies in alleviating hallucinations in audio language models.

acoustic evidenceaudio perturbationsaudio-language models

Existing diffusion models struggle to generate high-quality audio efficiently under low-frequency, highly compressed conditions due to implicit coupling in intermediate representations. This work proposes ReGen, a novel framework that jointly models data and representations through multi-vector fields and enhances the generalization of conditional flow matching via Generalized Flow Matching (GFM). ReGen introduces a hierarchical multi-prompt representation mechanism, enabling high-fidelity speech and audio synthesis at extremely low sampling rates (6.25–12.5 Hz). Built upon a diffusion Transformer, a neural audio codec, and a latent diffusion architecture, ReGenVoice requires only four GPUs for one day of training and achieves a real-time factor (RTF) of 0.08 during inference, significantly improving word error rate (WER) for intelligibility and speaker similarity (SIM).

generative capacitylatent entanglementlow-rate latent representation

Hot Scholars

CL

Chenglong Li

Professor, The University of Florida
Drug DesignDrug DiscoveryMolecular RecognitionMolecular Modeling
YY

Yinfeng Yu

Associate Professor, Xinjiang University
Embodied intelligence
LL

Liang Lin

Fellow of IEEE/IAPR, Professor of Computer Science, Sun Yat-sen University
Embodied AICausal Inference and LearningMultimodal Data Analysis