AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation

πŸ“… 2026-09-29
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the challenge in parameter-sharing architectures where ineffective corrections from advisor models interfere with learning, thereby constraining their capacity to guide frozen large language models. To overcome this, we propose AdviSD, a novel framework that introduces the first decision selection mechanism based on advice discrepancy magnitude to eliminate conflicts among correction objectives. Furthermore, it integrates outcome-oriented reinforcement learning with multi-turn selective self-distillation, driving training by evaluating the impact of advice to filter effective feedback. Extensive experiments demonstrate that AdviSD significantly outperforms baseline methods on the BFCL-v3 and EnvScaler benchmarks. Additionally, the framework exhibits strong cross-domain generalization and transferability across diverse model families, highlighting its robustness and broad applicability in enhancing advisor-guided learning for frozen large models.
πŸ“ Abstract
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
Problem

Research questions and friction points this paper is trying to address.

advisor model
self-distillation
frontier LLMs
feedback learning
correction selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Self-Distillation
Reinforcement Learning
LLM Advisor
Selective Supervision
Cross-Model Transfer
πŸ”Ž Similar Papers
No similar papers found.