🤖 AI Summary
This study addresses the issue that multi-turn agent self-distillation often induces hallucinated confidence due to information loss, resulting in performance inferior to the base model. To overcome this, we propose Privileged Self-Practice (PSP), whose core innovation lies in shifting privileged information from the loss function to the sampling stage. Specifically, PSP injects privileged information into instructions and performs online policy resampling based on the GRPO objective, combined with an analyzer model to guide training, thereby effectively mitigating spurious confidence. Experimental results demonstrate that PSP comprehensively surpasses existing baselines on the AppWorld and SWE-bench benchmarks, achieving up to a 65% improvement in task completion rate.
📝 Abstract
On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain RL, in the worst case below the untrained base model. Therefore, we propose Privileged Self-Practice (PSP), which keeps the PI and moves it from the loss to the sampler. When the student's rollouts on a task mostly fail, we inject a short per-task instruction written by an analyzer model, sample the task again with the instruction in context, and train on the result with an unchanged GRPO objective. The privileged information stays in the prompt and never enters the loss. Across AppWorld and SWE-bench Verified, with three different student models, PSP obtains the best average score in every setting and is the only method that consistently outperforms plain GRPO, improving task-goal completion by up to 65% on AppWorld and the resolved rate by up to 61% on SWE-bench Verified.