EvoAudio: Recursive Self-Improvement for Audio Understanding

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
EvoAudio通过递归自我改进系统解决音频理解问题,利用当前模型性能调整训练数据难度和焦点,结合强化学习提高模型表现。
📝 Abstract
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio, a recursive self-improvement system for audio understanding. To our knowledge, it is the first to evolve the model, waveforms, questions, and difficulty in one closed loop. EvoAudio uses the current model's performance to set the focus and difficulty of the next training data. A library of audio tools then constructs questions whose answers follow from how the audio was made, providing verifiable supervision without new human annotation. Reinforcement learning updates the model, and validation decides whether it enters the next evolution round. Across 13 rounds, EvoAudio improves five models with different audio encoders and language backbones on MMSU, MMAU-Pro, and MMAR. It achieves the highest average for every backbone, raising overall performance by up to 6.3 points. The improvement unfolds over successive rounds, with each stronger model starting the next round.
Problem

Research questions and friction points this paper is trying to address.

audio understanding
acoustic annotation
model improvement
Innovation

Methods, ideas, or system contributions that make the work stand out.

recursive self-improvement
audio understanding
reinforcement learning
evolutionary training
🔎 Similar Papers
No similar papers found.
Y
Yuxiang Wang
The Chinese University of Hong Kong, Shenzhen
S
Shengbo Cai
Tsinghua University
Y
Yingda Shen
The Chinese University of Hong Kong, Shenzhen
M
Ming-Hao Hsu
The Chinese University of Hong Kong, Shenzhen
Q
Qinke Ni
The Chinese University of Hong Kong, Shenzhen
L
Liqiang Zhang
Tencent Hunyuan
T
Teddy Sun
Tencent Hunyuan
S
Steve Yevs
Tencent Hunyuan
Zhizheng Wu
Zhizheng Wu
The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen), Mel Lab
Spoken Language ProcessingDeepFake detectionMusic Processing