Score
Designs and implements reinforcement-learning-based adaptation pipelines and policy-optimization methods that modify ASR model parameters or decoding policies to improve speech transcription; this includes defining reward functions, training/finetuning loops, and constrained policy updates for audio-language components. Practitioners build and evaluate RL agents or optimization procedures that adapt decoders or model behavior from limited synthetic or real data to handle distribution shifts such as mixed-language segments and boundary effects.
This work addresses the underexplored application of reinforcement learning (RL) in large language model (LLM)-driven automatic speech recognition (ASR) and text-to-speech (TTS) systems. We present the first systematic investigation into RL adaptation for audio-modal generative tasks, proposing a lightweight audio-aware RL framework that integrates Group Relative Policy Optimization (GRPO) with Differentiable Reward Optimization (DiffRO). To enhance reliability and efficiency, we design a rule-guided differentiable reward function and a low-resource-friendly data construction strategy. Experiments demonstrate substantial improvements in ASR word error rate reduction and TTS naturalness—measured by MOS and SIM scores—under stringent constraints of limited labeled data and few optimization steps. Our results validate the feasibility, effectiveness, and generalization potential of RL for LLM-based speech generation, establishing a foundation for efficient, reward-driven audio modeling without extensive supervision.
This work addresses the challenge of code-switched speech recognition in audio language models, where the absence of explicit language boundary modeling often leads to transcription errors and script contamination—i.e., inappropriate mixing of writing systems. To mitigate these issues, the authors propose a data-efficient reinforcement learning fine-tuning approach that integrates a verifiable reward mechanism combining word error rate and script fidelity. Building upon Qwen2-Audio, the method employs Group Relative Policy Optimization (GRPO) with LoRA adapters and introduces a two-stage draft-and-refine decoding strategy. Remarkably, using only 10% of synthetically generated code-switched data, the model matches the performance of full supervised fine-tuning across ten typologically diverse language pairs. The approach effectively eliminates translation errors, suppresses script pollution, and demonstrates strong zero-shot generalization to real human-recorded code-switched speech.
This work addresses the limited robustness of current automatic speech recognition (ASR) systems under distribution shifts—such as background noise or speaker accents—and the susceptibility of test-time adaptation methods to confirmation bias caused by high-confidence erroneous predictions. To mitigate these issues, the authors propose ASR-TRA, a novel framework that introduces causal intervention into test-time adaptation for the first time. ASR-TRA jointly optimizes the ASR model and a learnable decoder prompt during inference via reinforcement learning. It generates diverse transcription candidates through temperature-controlled stochastic decoding and leverages a reward signal based on audio-text semantic alignment to guide adaptation. This approach effectively alleviates confirmation bias, enhances adaptation stability and interpretability, and achieves significant performance gains over existing methods on noisy LibriSpeech and L2 Arctic accented speech benchmarks, all while maintaining low latency.
This work addresses the severe underutilization of reinforcement learning (RL) in audio modalities. It introduces the first application of Group Relative Policy Optimization (GRPO) to the audio large language model Qwen2-Audio-7B-Instruct, systematically evaluating its efficacy on audio question answering (AQA). Methodologically, it adopts an RL-based post-training paradigm, achieving convergence with only 38K samples; ablation studies reveal that explicit reasoning yields marginal gains for AQA performance. On the MMAU Test-mini benchmark, the approach achieves 64.5% accuracy, establishing a new state-of-the-art at the time. This study fills a critical research gap in RL for audio understanding and reasoning. All models and code are publicly released on GitHub and Hugging Face, providing a reproducible baseline and a novel paradigm for multimodal RL research.
ASR systems exhibit significant performance degradation in recognizing rare named entities and adapting to out-of-domain speech. To address this, we propose an unsupervised end-to-end domain adaptation method that integrates large language models (LLMs) as context-aware reward models within a reinforcement learning framework—specifically, Proximal Policy Optimization (PPO)—to generate label-free feedback signals for ASR model fine-tuning. Unlike conventional approaches relying on manual annotations or pseudo-labeling, our method leverages LLMs to dynamically assess the semantic plausibility of ASR hypotheses, thereby guiding policy optimization. On named entity recognition benchmarks, our approach reduces word error rate by 21% compared to standard self-training baselines, markedly improving cross-domain robustness. The core contribution lies in the native integration of LLMs into the ASR RL pipeline, enabling fully unsupervised, context-sensitive, and semantics-driven domain adaptation without any human supervision or external labeling.
This study addresses the challenge of adapting large language model–driven automatic speech recognition (ASR) systems in regulated domains such as banking, where access to real-world speech data is severely restricted due to privacy and compliance constraints. To overcome the acoustic distribution mismatch between synthetic and real speech, the authors propose the first application of the reward-free reinforcement learning algorithm GRPO (Group Relative Policy Optimization) for ASR adaptation using only synthetic data. By optimizing policies through rewards based on low word error rate (WER), GRPO reduces WER from 36.71% to 22.09%—a 40% relative improvement over supervised fine-tuning (SFT). Combining SFT with GRPO yields a further 45% reduction. The performance gain is attributed to behavioral policy optimization rather than changes in representation, thereby surpassing the limitations of conventional SFT.
This work addresses the limited exploration of flow-matching methods in reinforcement learning (RL) for text-to-speech (TTS) synthesis, where existing approaches predominantly focus on large language models. The paper proposes FlowTTS-GRPO, the first framework to apply online RL directly to end-to-end fine-tuning of open-source flow-matching TTS models. It reformulates the ordinary differential equation (ODE) trajectory as a stochastic differential equation (SDE) path and introduces a weighted multi-objective reward mechanism to replace conventional probabilistic guidance. Key findings reveal that omitting classifier-free guidance (CFG), emphasizing hard examples, and applying RL exclusively to the flow-matching component significantly enhance performance. Evaluated on CosyVoice 3.0 and F5-TTS, the method consistently improves speaker similarity and perceptual audio quality, with F5-TTS additionally achieving notable gains in intelligibility, outperforming baselines across both subjective and objective metrics.
This work addresses the challenge of insufficient supervision in action–environment interactions during reinforcement learning (RL) for language agents, as well as the reliance of existing world models on external simulators or added inference overhead. The authors propose PaW, a novel framework that unifies policy training and world modeling within a single on-policy learning loop, leveraging state transitions from policy-sampled trajectories as supervision signals for the world model—without incurring additional computational costs. Key innovations include action-entropy-based data selection, a noise-robust loss function, and a reward-adaptive dynamic loss balancing mechanism. Experiments demonstrate that PaW significantly outperforms strong RL baselines across three agent benchmark tasks, confirming that standard RL trajectories can serve as an effective source of supervision for world modeling.