Reinforcement Learning for Data-Efficient Code-Switched ASR

๐Ÿ“… 2026-07-02
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the challenge of code-switched speech recognition in audio language models, where the absence of explicit language boundary modeling often leads to transcription errors and script contaminationโ€”i.e., inappropriate mixing of writing systems. To mitigate these issues, the authors propose a data-efficient reinforcement learning fine-tuning approach that integrates a verifiable reward mechanism combining word error rate and script fidelity. Building upon Qwen2-Audio, the method employs Group Relative Policy Optimization (GRPO) with LoRA adapters and introduces a two-stage draft-and-refine decoding strategy. Remarkably, using only 10% of synthetically generated code-switched data, the model matches the performance of full supervised fine-tuning across ten typologically diverse language pairs. The approach effectively eliminates translation errors, suppresses script pollution, and demonstrates strong zero-shot generalization to real human-recorded code-switched speech.
๐Ÿ“ Abstract
Audio-language models can be prompted for code-switched speech, but their decoding is not optimized for code-switching and often fails at language boundaries. We propose a practical reinforcement learning with verifiable rewards recipe for data-efficient adaptation of audio-language models to code-switched ASR using group relative policy optimization, combining an error rate reward with a script fidelity reward that penalizes wrong writing systems and a two-pass draft-and-refinement procedure. Using Qwen2-Audio as a reproducible testbed across 10 language pairs, training on only TTS code-switched speech, we show that RLVR with 10% of the data matches LoRA supervised fine-tuning trained on the full dataset, with the largest gains on typologically distant pairs. The error rate reward eliminates translation errors while the script fidelity reward separately reduces script contamination without degradation. These gains transfer zero-shot to a human-recorded code-switching corpus.
Problem

Research questions and friction points this paper is trying to address.

code-switched ASR
audio-language models
language boundaries
speech recognition
data efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

reinforcement learning
code-switched ASR
script fidelity reward
data-efficient adaptation
audio-language models
Z
Ziwei Ye
Independent Researcher
P
Peter Vickers
Spotify Canada