TS-OPD: Reconciling ASR and QA in Speech Language Models via Task-Specific On-Policy Distillation

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the significant degradation in instruction-following and question-answering (QA) capabilities observed when speech language models are specialized for automatic speech recognition (ASR). To mitigate this issue, we propose a task-specific online distillation framework that introduces a novel dual-teacher complementary mechanism. By leveraging both pre- and post-specialization models as complementary teachers, our approach generates independent supervision signals through task-conditioned trajectory generation, effectively decoupling the optimization conflicts between ASR and QA objectives. Experimental results demonstrate that the proposed method substantially improves ASR accuracy while effectively preserving QA proficiency. Furthermore, the framework exhibits robustness to variations in the balancing coefficient and yields consistent performance gains as the data scale increases.
📝 Abstract
Speech Language Models (SLMs) inherit strong instruction-following capabilities from pretrained language models, yet ASR specialization can substantially degrade them. To address this ASR--QA trade-off, we propose Task-Specific On-Policy Distillation (TS-OPD), which leverages models before and after ASR specialization as complementary QA and ASR teachers. The student generates separate task-conditioned trajectories for ASR and QA, each supervised only by its corresponding teacher, thereby reducing direct competition between the two supervision signals. Experiments on basic ASR, contextual ASR, and QA demonstrate that TS-OPD improves recognition while preserving QA capability. Moreover, TS-OPD remains robust across different balancing coefficients and continues to benefit from increased distillation data.
Problem

Research questions and friction points this paper is trying to address.

Speech Language Models
Automatic Speech Recognition
Question Answering
Trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Task-Specific On-Policy Distillation
Speech Language Models
ASR-QA Trade-off
Knowledge Distillation
Multi-task Learning
🔎 Similar Papers
No similar papers found.