Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
为解决语音翻译中训练与推理的不匹配问题,本文提出了一种基于组相对策略优化的联合识别和翻译微调方法,有效提升了翻译质量和识别准确率。
📝 Abstract
In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.
Problem

Research questions and friction points this paper is trying to address.

Speech Translation
Transcription
Chain-of-Thought (CoT)
Training-Inference Mismatch
Supervised Fine-Tuning (SFT)
Innovation

Methods, ideas, or system contributions that make the work stand out.

Group Relative Policy Optimization (GRPO)
Joint Recognition and Translation Fine-Tuning
Token Advantage Strategies
🔎 Similar Papers
No similar papers found.
Y
Yanghe Dong
Independent Researcher
W
Wanting Huang
Department of Computer Science, University of Iowa, USA
Weiran Wang
Weiran Wang
University of Iowa
Machine learningspeech processing