๐ค AI Summary
This work addresses the challenge of code-switched speech recognition in audio language models, where the absence of explicit language boundary modeling often leads to transcription errors and script contaminationโi.e., inappropriate mixing of writing systems. To mitigate these issues, the authors propose a data-efficient reinforcement learning fine-tuning approach that integrates a verifiable reward mechanism combining word error rate and script fidelity. Building upon Qwen2-Audio, the method employs Group Relative Policy Optimization (GRPO) with LoRA adapters and introduces a two-stage draft-and-refine decoding strategy. Remarkably, using only 10% of synthetically generated code-switched data, the model matches the performance of full supervised fine-tuning across ten typologically diverse language pairs. The approach effectively eliminates translation errors, suppresses script pollution, and demonstrates strong zero-shot generalization to real human-recorded code-switched speech.
๐ Abstract
Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable rewards recipe for data-efficient adaptation of audio-language models to code-switched ASR using group relative policy optimization, combining an error rate reward with a script fidelity reward that penalizes wrong writing systems and a two-pass draft-and-refinement procedure. Using Qwen2-Audio as a reproducible testbed across 10 language pairs, training on only TTS code-switched speech, we show that RLVR with 10% of the data matches LoRA supervised fine-tuning trained on the full dataset, with the largest gains on typologically distant pairs. The error rate reward eliminates translation errors while the script fidelity reward separately reduces script contamination without degradation. These gains transfer zero-shot to a human-recorded code-switching corpus.