SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issue in online knowledge distillation where weak student models suffer from unrepresentative teacher supervision signals due to state shift. To tackle this, we propose the SAKI framework, which introduces a novel maximal coupling-based supervision allocation mechanism that rigorously bounds correction probabilities via total variation distance for adaptive conflict resolution. Furthermore, it incorporates guided generation and routing mechanisms constrained by KL divergence to apply direct supervision at corrected positions, alongside an engine-resident speculative verifier to optimize training efficiency. Empirical evaluations demonstrate that our approach substantially enhances the performance of small-scale models across seven mathematical reasoning benchmarks while achieving a 4.22× improvement in sampling throughput.
📝 Abstract
On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and upper-bounds intervention and specialized-supervision frequency. We further implement an engine-resident speculative verifier that preserves the exact-q trajectory distribution and coupling semantics while improving matched-workload rollout throughput by 4.22x. Across seven mathematical reasoning benchmarks, SAKI improves the matched teacher-guided baseline in Mean@8 and Pass@8 for both 1.7B and 0.6B students. Placement controls and fixed-prefix analysis further support correction-triggered routing as a conflict-adaptive supervision signal.
Problem

Research questions and friction points this paper is trying to address.

On-policy distillation
state mismatch
weak student models
supervision misalignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

On-Policy Distillation
Maximal Coupling
KL-Constrained Interpolation
Token-Level Supervision Routing
Speculative Verifier
🔎 Similar Papers
2024-07-21arXiv.orgCitations: 1