🤖 AI Summary
This study addresses the challenge in online policy distillation where teacher intervention alters the student's distribution, making it difficult for fixed-intensity strategies to balance trajectory quality against off-policy deviation. To overcome this limitation, this work proposes MAESTRO, a framework that adaptively regulates both the timing of teacher takeover and the generation duration through a segment-aggregated policy divergence scoring mechanism, thereby achieving joint dynamic adaptation of intervention depth and length. Empirical evaluations demonstrate that MAESTRO attains state-of-the-art average accuracy across eight mathematical reasoning benchmarks. Furthermore, it reduces training response lengths by 67.3% compared to standard methods, substantially improving knowledge transfer efficiency.
📝 Abstract
On-policy distillation (OPD) trains a student on its own reasoning trajectories using feedback from a stronger teacher. Teacher interventions can improve these trajectories, but also change the distribution on which the student learns. Our controlled studies show that rollout quality alone is an incomplete criterion for allocating teacher guidance. Deeper intervention yields diminishing gains in rollout accuracy while increasing off-policy load. In a training probe with a restricted rollout horizon, peak student accuracy and performance retention favor different intervention strengths. The preferred intervention depth and placement also vary across benchmarks. These findings motivate MAESTRO, which uses local policy disagreement to jointly adapt when the teacher takes over and how long it generates. Its {policy disagreement score} combines teacher-weighted candidate coverage with local distribution similarity and is aggregated within reasoning paragraphs. Across eight mathematical reasoning benchmarks, MAESTRO achieves the highest macro-average accuracy among the compared methods for both 0.6B and 1.7B Qwen3 students, with the 1.7B student leading on every benchmark. MAESTRO also reduces average training response length by 67.3\% relative to standard OPD. The code is available at https://github.com/yhao-wang/MAESTRO.