Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering

📅 2025-03-14
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the severe underutilization of reinforcement learning (RL) in audio modalities. It introduces the first application of Group Relative Policy Optimization (GRPO) to the audio large language model Qwen2-Audio-7B-Instruct, systematically evaluating its efficacy on audio question answering (AQA). Methodologically, it adopts an RL-based post-training paradigm, achieving convergence with only 38K samples; ablation studies reveal that explicit reasoning yields marginal gains for AQA performance. On the MMAU Test-mini benchmark, the approach achieves 64.5% accuracy, establishing a new state-of-the-art at the time. This study fills a critical research gap in RL for audio understanding and reasoning. All models and code are publicly released on GitHub and Hugging Face, providing a reproducible baseline and a novel paradigm for multimodal RL research.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: Language Grounding & Multi-modal NLPMultiagent Systems: Multiagent Learning

Application Category

Search and Retrieval-Augmented AI: Retrieval-Augmented Generation (RAG) and multi-modal RAGEconomics, Online Markets and Human Computation: LLM based quality controls for crowd workResponsible Web: Machine-in-the-loop, human agency and autonomy
📝 Abstract
Recently, reinforcement learning (RL) has been shown to greatly enhance the reasoning capabilities of large language models (LLMs), and RL-based approaches have been progressively applied to visual multimodal tasks. However, the audio modality has largely been overlooked in these developments. Thus, we conduct a series of RL explorations in audio understanding and reasoning, specifically focusing on the audio question answering (AQA) task. We leverage the group relative policy optimization (GRPO) algorithm to Qwen2-Audio-7B-Instruct, and our experiments demonstrated state-of-the-art performance on the MMAU Test-mini benchmark, achieving an accuracy rate of 64.5%. The main findings in this technical report are as follows: 1) The GRPO algorithm can be effectively applied to large audio language models (LALMs), even when the model has only 8.2B parameters; 2) With only 38k post-training samples, RL significantly outperforms supervised fine-tuning (SFT), indicating that RL-based approaches can be effective without large datasets; 3) The explicit reasoning process has not shown significant benefits for AQA tasks, and how to efficiently utilize deep thinking remains an open question for further research; 4) LALMs still lag far behind humans auditory-language reasoning, suggesting that the RL-based approaches warrant further exploration. Our project is available at https://github.com/xiaomi/r1-aqa and https://huggingface.co/mispeech/r1-aqa.
Problem

Research questions and friction points this paper is trying to address.

Explores RL for audio understanding in AQA tasks.
Demonstrates RL outperforms SFT with fewer samples.
Highlights gap between LALMs and human auditory reasoning.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Applied GRPO algorithm to audio language models
Reinforcement learning outperforms supervised fine-tuning
Achieved 64.5% accuracy on MMAU Test-mini
🔎 Similar Papers
No similar papers found.