Reward-Tilted On-Policy Distillation for Acoustic Grounding in Audio-Language Models

πŸ“… 2026-09-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the tendency of audio language models to rely on textual shortcuts while neglecting acoustic evidence. To mitigate this, we propose Reward-Tilted Online Policy Distillation, a method that constructs a reward signal by contrasting the log-probabilities of teacher predictions with and without audio input, thereby explicitly disentangling acoustic support from linguistic predictability. The student model’s output distribution is subsequently reshaped via reverse Kullback-Leibler divergence optimization to reinforce its reliance on acoustic cues. Experimental results demonstrate that our approach significantly outperforms baselines across three benchmarks. Notably, a 3B-parameter model achieves 72.72% accuracy on MMAU, delivering performance comparable to substantially larger models.
πŸ“ Abstract
Audio-language models (ALMs) can exploit textual shortcuts to answer questions while overlooking acoustic evidence, weakening audio understanding. On-policy distillation (OPD) trains compact ALMs by supervising student-generated responses with teacher predictions, but does not explicitly distinguish acoustic support from linguistic predictability. We propose Reward-Tilted On-Policy Distillation (RT-OPD) to strengthen acoustic grounding. Given the same question and student-generated text, a frozen teacher predicts the next token with and without audio inputs. Their log-probability contrast defines a reward that reshapes the teacher distribution for reverse-KL distillation, emphasizing the additional evidence provided by audio. Across two compact students and three benchmarks, RT-OPD consistently outperforms Vanilla OPD. Experiments with silenced and replacement audio further suggest that RT-OPD strengthens the student's reliance on acoustic evidence. Our 3B model achieves 72.72% accuracy on MMAU, the highest among the compared 3B models and competitive with several 7B and 8B models. Code and model checkpoints are available at https://github.com/KaiyangLi1992/RT-OPD.
Problem

Research questions and friction points this paper is trying to address.

Audio-Language Models
Acoustic Grounding
Textual Shortcuts
On-Policy Distillation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Reward-Tilted On-Policy Distillation
Acoustic Grounding
Audio-Language Models
Reverse-KL Distillation
Log-probability Contrast
πŸ”Ž Similar Papers
No similar papers found.