🤖 AI Summary
This work addresses the severe underutilization of reinforcement learning (RL) in audio modalities. It introduces the first application of Group Relative Policy Optimization (GRPO) to the audio large language model Qwen2-Audio-7B-Instruct, systematically evaluating its efficacy on audio question answering (AQA). Methodologically, it adopts an RL-based post-training paradigm, achieving convergence with only 38K samples; ablation studies reveal that explicit reasoning yields marginal gains for AQA performance. On the MMAU Test-mini benchmark, the approach achieves 64.5% accuracy, establishing a new state-of-the-art at the time. This study fills a critical research gap in RL for audio understanding and reasoning. All models and code are publicly released on GitHub and Hugging Face, providing a reproducible baseline and a novel paradigm for multimodal RL research.
📝 Abstract
Recently, reinforcement learning (RL) has been shown to greatly enhance the reasoning capabilities of large language models (LLMs), and RL-based approaches have been progressively applied to visual multimodal tasks. However, the audio modality has largely been overlooked in these developments. Thus, we conduct a series of RL explorations in audio understanding and reasoning, specifically focusing on the audio question answering (AQA) task. We leverage the group relative policy optimization (GRPO) algorithm to Qwen2-Audio-7B-Instruct, and our experiments demonstrated state-of-the-art performance on the MMAU Test-mini benchmark, achieving an accuracy rate of 64.5%. The main findings in this technical report are as follows: 1) The GRPO algorithm can be effectively applied to large audio language models (LALMs), even when the model has only 8.2B parameters; 2) With only 38k post-training samples, RL significantly outperforms supervised fine-tuning (SFT), indicating that RL-based approaches can be effective without large datasets; 3) The explicit reasoning process has not shown significant benefits for AQA tasks, and how to efficiently utilize deep thinking remains an open question for further research; 4) LALMs still lag far behind humans auditory-language reasoning, suggesting that the RL-based approaches warrant further exploration. Our project is available at https://github.com/xiaomi/r1-aqa and https://huggingface.co/mispeech/r1-aqa.